A secret on a puzzle site that nobody had ever found. I spent 3 days on it with a crew of AI agents, 839 real attempts and a probability gate built on Jev, and I got there first.
In this post:
- Stuck
- One rule: no shortcuts
- The tool that made real clicks cheap
- Brute force
- I asked for a judge
- The judge was useless
- Rebuilding it
- The agent brainstorms, the judge ranks
- The odd one
- The near-miss
- What Jev is good for
- What I learned
Stuck
The site hides a few hundred secrets. Each one has an exact rule on the server, and you either meet it or you don't. They come in tiers, and platinum is the hardest. Mine was platinum, with zero finders and a cash bounty for whoever got there first.
The secret lives in one of the site's mini-games. You move elements around in the browser, save, and the server checks the result against its rule. Every attempt means actually moving things on a real page.
I had very little to go on. One known secret was the parent of mine. There was a locked section I couldn't open. And a couple of thin clues.
That's thin for a rule that has to match exactly.
One rule: no shortcuts
I ran a small crew. Claude Code orchestrated. Grok CLI and Devin CLI generated ideas and built attempts. Later, TypeSafe's Jev would score them.
Before the first build I set one rule for all of them. No API shortcuts. Nobody POSTs to the site directly. Every build gets made with real mouse drags in a real browser, and the site's own JS does the saving.
Part of that is about the game. Fuzzing someone's endpoint isn't playing.
The other part turned out to matter more. A real build takes real seconds. When every attempt costs time, dumb attempts hurt.
The tool that made real clicks cheap
A no-shortcuts rule sounds noble until you have to script it. Every agent in the crew went through the same tool: dev-browser, my extended fork of SawyerHood/dev-browser. It's a browser-automation plugin for Claude Code, open source at https://github.com/MarcinDudekDev/dev-browser.
What I added on top of the original:
- a few dozen CLI commands
- dev, stealth and user browser modes
- per-project page isolation
- crash auto-retry
- YAML scenarios
- an optional faster JS runtime that shaves time off every start
clientandpageauto-injected into every script, so no setup code
In this hunt it carried thousands of real mouse drags from Claude, Grok and Devin, all through one tool. The session was bridged from my real Brave browser with a cookie bridge, so no agent ever touched a password. And a small script read the site's own near-miss signal after each save. Remember that script. It comes back later.
If every build had needed a hand-written Playwright setup, I'd have given in and POSTed by day two. dev-browser made "real clicks only" practical, and that constraint is what pushed me to build a filter.
Brute force
The first day was pure volume. Every element alone, everything at once, groups by category, pop culture, drawn faces. Grok ran a couple of rounds of its own. Nothing.
Then Devin started an exhaustive sweep over every subset of a small group of elements. Hundreds of builds queued up. I watched a few go by and killed it.
My exact words were "they're so dumb".
A sweep has no idea what a puzzle designer finds funny. It only knows what's left to enumerate. With real clicks, a blind sweep is an afternoon gone.
The second day was quiet. I was running out of ideas that weren't more of the same.
| Day | Real saves | What happened |
|---|---|---|
| Day 1 | 610 | brute force: singles, everything at once, categories, pop culture, drawn faces |
| Day 2 | 46 | running out of non-repeating ideas |
| Day 3 | 183 | Jev gate v1, then v2 + Devin loop, near-miss, solved |
I asked for a judge
At that point I didn't want more builds. I wanted a way to decide which builds were worth making. So I stopped everything and asked for a probability gate with Jev before anyone built anything else.
Jev is TypeSafe's "System One" model. You ask it a narrow question and it returns a typed judgment: a probability, a score, or a pick from a list. No paragraph to parse. That makes it cheap to call many times and easy to combine.
The first version asked Jev 2 questions per idea and scored a few dozen ideas. The best one got 0.26.
I built the top ten anyway. Nothing.
The judge was useless
Worse than the low top score, everything else sat packed right under it. The scores were flat.
A judge that gives everything the same score can't rank anything, however high or low that score is.
I wanted to know why before trusting it again. Did it use the thing I had raised earlier, that elements keep their positions and the layout might count?
It hadn't. Worse, the prompt told Jev that position probably doesn't matter. Half of the gate rested on the opposite of something I had already flagged. No wonder it couldn't tell one idea from another.
Rebuilding it
So I specified the fix myself. My words were "many more gates and things we know... a composite score with ALL of what we know". Every rule, every clue, the locked section, the layout, the full list of what had already failed, on every call.
The second version asked Jev 12 narrow questions per idea instead of 2 broad ones:
- fit with the obvious theme
- fit with the second clue
- fit with the parent secret
- do the counts mean something?
- does it match the designer's usual style?
- is it fair but hard, the way a platinum rule should be?
- novelty against what was already tried
- is it a real thing people would recognize?
- a second, narrower theme check
- does it use elements across categories?
- does layout play a role?
- overall: is this the answer?
Each answer became a number, and code combined them with fixed weights. Because the weights live in code, I can change them and re-rank every idea without calling the model again.
On top of Jev I added plain code checks. Repeats of earlier builds get dropped or pushed down, and ideas that ignore the parent secret get a small penalty.
The table below has the before and after:
| v1 | v2 | |
|---|---|---|
| Questions per idea | 2 | 12 (+ code features) |
| Ideas scored | 49 | 393 |
| Best score | 0.26 | 0.65 |
| Mean score | 0.12 | ~0.39 |
| Spread (std dev) | 0.042 | 0.101 |
| Built, result | top 10, nothing | top 20, first near-miss in 3 days |
The best score went up, but I didn't look at that first. I looked at the spread. It more than doubled. For the first time the judge could say "this one, not that one".
Then I froze it. Nobody could change the questions or the weights after that.
This is the part I'd insist on again. If the generator can tune the judge, it will keep tuning until the judge agrees with it. With a frozen judge, a score in the last round means the same thing as a score in the first.
The agent brainstorms, the judge ranks
The loop was simple. Devin brainstorms a batch of ideas per round. The frozen judge ranks them. Devin reads the per-question scores, sees what worked, and pushes further in that direction. Only the very top of the list ever gets built with real clicks.
| Round | Ideas | Best | Mean |
|---|---|---|---|
| 1 | 64 | 0.645 | 0.427 |
| 2 | 60 | 0.650 | 0.434 |
| 3 | 60 | 0.578 | 0.388 |
| 4 | 58 | 0.613 | 0.397 |
| 5 | 49 | 0.530 | 0.333 |
| 6 | 35 | 0.539 | 0.368 |
Look at the shape. The first two rounds were the best. After that Devin was running dry. With a frozen judge you can watch that happen, and it's your cue to stop generating and start building.
Here's where every idea landed:
| Bin | Ideas |
|---|---|
| 0.15 | 8 |
| 0.20 | 26 |
| 0.25 | 43 |
| 0.30 | 59 |
| 0.35 | 73 |
| 0.40 | 65 |
| 0.45 | 64 |
| 0.50 | 27 |
| 0.55 | 22 ← the idea that led to the answer (0.59, rank 12 of 393) |
| 0.60 | 5 |
| 0.65 | 1 |
The odd one
The top of the list was crowded with one theme. The one everybody found obvious, me included.
And then, at rank 12 of 393, sat an idea that didn't look like the others.
Jev loved it on two of the clue questions and liked it on two more. It was weak on the obvious theme. And when asked straight out whether this was the answer, Jev said probably not.
So Jev didn't know. What it did was keep an odd idea near the top, inside the list I was going to build. The old 2-question gate would have left it in the flat middle with everything else.
One more honest detail. That idea didn't come from me, and it didn't come from Claude. Devin wrote it, in the very first round. I think that's the best argument in this post for keeping the generator and the judge separate.
The generator's job is to be weird. The judge's job is to not bury the weird ones.
The near-miss
I built the top of the list. The locked section stayed locked.
But the site has a second feedback channel. The page carries a token with server-side marks, and some of them record when you came close to a secret. That's what my small script was reading after every save.
After that batch, a new mark showed up for this secret. The first near-miss in 3 days. Its timestamp sat a fraction of a second before the save of the rank-12 build.
I re-saved it to read the hint that comes with a near-miss. It told me I was on the right track and roughly which way to adjust. A couple of builds later the secret was mine, about 6 minutes after the near-miss. First finder.
I was stoked. Three days of dead ends, and it closed in 6 minutes once the right build existed.
Here's the whole effort in one table:
| What | Number |
|---|---|
| Real saves | 839 |
| Elements moved by real mouse drags | 3,036 |
| Distinct combinations | 755 |
| Jev judgments | 4,814 (49 ideas x 2 questions + 393 ideas x 12) |
| Builds chosen by the gate | 18 + 10 |
| First near-miss to solved | ~6 minutes |
What Jev is good for
Jev's job here was cheap, repeatable triage over a search space too big to try by hand. Pair a creative generator with a cheap frozen judge, keep the weights in code, and do the expensive thing only at the end. That fits a lot of normal engineering work:
- Choosing what to test first. Ask narrow questions per test about the diff, like "does this change touch the code path this test covers?", and run the top slice first.
- Triage of bug reports, support tickets or sales leads. Score each one on severity, customer impact and "is this a duplicate of something open". When priorities shift, change the weights and re-rank without calling the model again.
- Ad and headline variants. An LLM writes hundreds. A judge scores clarity, claim accuracy and brand fit. Only a handful go into a paid A/B test.
- This post's own title. I ran it through the same setup. A couple of dozen candidates, each judged on curiosity, clarity, specificity, honesty, the Jev angle and a hard no-spoiler gate, among others. The punchy one-liner "Stop brute-forcing with AI agents. Build a judge." ended up near the bottom, sunk by a specificity score of almost nothing.
In every one of these the judge doesn't have to know the answer. It only has to keep the good odd ideas from getting buried before the expensive step.
What I learned
- Cost shapes strategy. When each attempt costs real time, you start asking which attempts deserve to exist.
- Brute force has a smell. When attempts stop teaching you anything, stop.
- A judge only knows what you feed it. The same model went from useless to useful once it got narrow questions and every known fact.
- Check the spread before the top score. A judge that gives everything the same number can't rank anything.
- Put the weights in code. You get a ranking you can read, question and re-weight for free.
- Keep the generator and the judge apart, and freeze the judge. If one can tune the other, they end up agreeing instead of finding anything.
- The model doesn't need to know the answer. It needs to keep the odd good ideas alive until the expensive step.
- Listen to the system you're poking. One near-miss signal was worth more than hundreds of blind attempts.
I love solving problems. It's what I do for a living, and it's what I love in games. This one came from an agent's idea that a model ranked 12th without betting on it. My part was building the system that kept it on the list, and knowing when to stop sweeping and start judging.
We're living in amazing times. AI has become an amazing tool for games and for the people who play them. It let me play at a scale I could never reach alone, and the win still felt like mine.
Want to start your own hunt? Here's my invite: thesecretsproject.com/invite/@MythThrazz.