How I used Jev as a probability gate to crack a platinum puzzle

839 attempts. Zero finders. One probability gate. How I used Jev to crack a platinum secret.

A secret on a puzzle site that nobody had ever found. I spent 3 days on it with a crew of AI agents, 839 real attempts and a probability gate built on Jev, and I got there first.

In this post:

Stuck

The site hides a few hundred secrets. Each one has an exact rule on the server, and you either meet it or you don't. They come in tiers, and platinum is the hardest. Mine was platinum, with zero finders and a cash bounty for whoever got there first.

The secret lives in one of the site's mini-games. You move elements around in the browser, save, and the server checks the result against its rule. Every attempt means actually moving things on a real page.

I had very little to go on. One known secret was the parent of mine. There was a locked section I couldn't open. And a couple of thin clues.

That's thin for a rule that has to match exactly.

One rule: no shortcuts

I ran a small crew. Claude Code orchestrated. Grok CLI and Devin CLI generated ideas and built attempts. Later, TypeSafe's Jev would score them.

Before the first build I set one rule for all of them. No API shortcuts. Nobody POSTs to the site directly. Every build gets made with real mouse drags in a real browser, and the site's own JS does the saving.

Part of that is about the game. Fuzzing someone's endpoint isn't playing.

The other part turned out to matter more. A real build takes real seconds. When every attempt costs time, dumb attempts hurt.

The tool that made real clicks cheap

A no-shortcuts rule sounds noble until you have to script it. Every agent in the crew went through the same tool: dev-browser, my extended fork of SawyerHood/dev-browser. It's a browser-automation plugin for Claude Code, open source at https://github.com/MarcinDudekDev/dev-browser.

What I added on top of the original:

In this hunt it carried thousands of real mouse drags from Claude, Grok and Devin, all through one tool. The session was bridged from my real Brave browser with a cookie bridge, so no agent ever touched a password. And a small script read the site's own near-miss signal after each save. Remember that script. It comes back later.

If every build had needed a hand-written Playwright setup, I'd have given in and POSTed by day two. dev-browser made "real clicks only" practical, and that constraint is what pushed me to build a filter.

Brute force

The first day was pure volume. Every element alone, everything at once, groups by category, pop culture, drawn faces. Grok ran a couple of rounds of its own. Nothing.

Then Devin started an exhaustive sweep over every subset of a small group of elements. Hundreds of builds queued up. I watched a few go by and killed it.

My exact words were "they're so dumb".

A sweep has no idea what a puzzle designer finds funny. It only knows what's left to enumerate. With real clicks, a blind sweep is an afternoon gone.

The second day was quiet. I was running out of ideas that weren't more of the same.

Day Real saves What happened
Day 1 610 brute force: singles, everything at once, categories, pop culture, drawn faces
Day 2 46 running out of non-repeating ideas
Day 3 183 Jev gate v1, then v2 + Devin loop, near-miss, solved

I asked for a judge

At that point I didn't want more builds. I wanted a way to decide which builds were worth making. So I stopped everything and asked for a probability gate with Jev before anyone built anything else.

Jev is TypeSafe's "System One" model. You ask it a narrow question and it returns a typed judgment: a probability, a score, or a pick from a list. No paragraph to parse. That makes it cheap to call many times and easy to combine.

The first version asked Jev 2 questions per idea and scored a few dozen ideas. The best one got 0.26.

I built the top ten anyway. Nothing.

The judge was useless

Worse than the low top score, everything else sat packed right under it. The scores were flat.

A judge that gives everything the same score can't rank anything, however high or low that score is.

I wanted to know why before trusting it again. Did it use the thing I had raised earlier, that elements keep their positions and the layout might count?

It hadn't. Worse, the prompt told Jev that position probably doesn't matter. Half of the gate rested on the opposite of something I had already flagged. No wonder it couldn't tell one idea from another.

Rebuilding it

So I specified the fix myself. My words were "many more gates and things we know... a composite score with ALL of what we know". Every rule, every clue, the locked section, the layout, the full list of what had already failed, on every call.

The second version asked Jev 12 narrow questions per idea instead of 2 broad ones:

Each answer became a number, and code combined them with fixed weights. Because the weights live in code, I can change them and re-rank every idea without calling the model again.

On top of Jev I added plain code checks. Repeats of earlier builds get dropped or pushed down, and ideas that ignore the parent secret get a small penalty.

The table below has the before and after:

v1 v2
Questions per idea 2 12 (+ code features)
Ideas scored 49 393
Best score 0.26 0.65
Mean score 0.12 ~0.39
Spread (std dev) 0.042 0.101
Built, result top 10, nothing top 20, first near-miss in 3 days

The best score went up, but I didn't look at that first. I looked at the spread. It more than doubled. For the first time the judge could say "this one, not that one".

Then I froze it. Nobody could change the questions or the weights after that.

This is the part I'd insist on again. If the generator can tune the judge, it will keep tuning until the judge agrees with it. With a frozen judge, a score in the last round means the same thing as a score in the first.

The agent brainstorms, the judge ranks

The loop was simple. Devin brainstorms a batch of ideas per round. The frozen judge ranks them. Devin reads the per-question scores, sees what worked, and pushes further in that direction. Only the very top of the list ever gets built with real clicks.

Round Ideas Best Mean
1 64 0.645 0.427
2 60 0.650 0.434
3 60 0.578 0.388
4 58 0.613 0.397
5 49 0.530 0.333
6 35 0.539 0.368

Look at the shape. The first two rounds were the best. After that Devin was running dry. With a frozen judge you can watch that happen, and it's your cue to stop generating and start building.

Here's where every idea landed:

Bin Ideas
0.15 8
0.20 26
0.25 43
0.30 59
0.35 73
0.40 65
0.45 64
0.50 27
0.55 22 ← the idea that led to the answer (0.59, rank 12 of 393)
0.60 5
0.65 1

The odd one

The top of the list was crowded with one theme. The one everybody found obvious, me included.

And then, at rank 12 of 393, sat an idea that didn't look like the others.

Jev loved it on two of the clue questions and liked it on two more. It was weak on the obvious theme. And when asked straight out whether this was the answer, Jev said probably not.

So Jev didn't know. What it did was keep an odd idea near the top, inside the list I was going to build. The old 2-question gate would have left it in the flat middle with everything else.

One more honest detail. That idea didn't come from me, and it didn't come from Claude. Devin wrote it, in the very first round. I think that's the best argument in this post for keeping the generator and the judge separate.

The generator's job is to be weird. The judge's job is to not bury the weird ones.

The near-miss

I built the top of the list. The locked section stayed locked.

But the site has a second feedback channel. The page carries a token with server-side marks, and some of them record when you came close to a secret. That's what my small script was reading after every save.

After that batch, a new mark showed up for this secret. The first near-miss in 3 days. Its timestamp sat a fraction of a second before the save of the rank-12 build.

I re-saved it to read the hint that comes with a near-miss. It told me I was on the right track and roughly which way to adjust. A couple of builds later the secret was mine, about 6 minutes after the near-miss. First finder.

I was stoked. Three days of dead ends, and it closed in 6 minutes once the right build existed.

Here's the whole effort in one table:

What Number
Real saves 839
Elements moved by real mouse drags 3,036
Distinct combinations 755
Jev judgments 4,814 (49 ideas x 2 questions + 393 ideas x 12)
Builds chosen by the gate 18 + 10
First near-miss to solved ~6 minutes

What Jev is good for

Jev's job here was cheap, repeatable triage over a search space too big to try by hand. Pair a creative generator with a cheap frozen judge, keep the weights in code, and do the expensive thing only at the end. That fits a lot of normal engineering work:

In every one of these the judge doesn't have to know the answer. It only has to keep the good odd ideas from getting buried before the expensive step.

What I learned

I love solving problems. It's what I do for a living, and it's what I love in games. This one came from an agent's idea that a model ranked 12th without betting on it. My part was building the system that kept it on the list, and knowing when to stop sweeping and start judging.

We're living in amazing times. AI has become an amazing tool for games and for the people who play them. It let me play at a scale I could never reach alone, and the win still felt like mine.

Want to start your own hunt? Here's my invite: thesecretsproject.com/invite/@MythThrazz.