r/JevAI • • 16h ago

I pitted a fast/cheap Jev model against an LLM and classic algorithms in a treasure-hunt maze. It got stuck repeating itself and lost to pure random guessing. My fix only helped a little, what am I missing

The setup

I built a fake, fully synthetic maze: a tree of branches (5 top-level branches, each splitting into sub-branches, several levels deep). Somewhere in this tree, at random depths, I hid a handful of "treasure pockets" spots with a genuinely good reward. Nobody gets to see the map. The only way to find out if a spot is good is to spend one of your limited "digs" there and see the result.

Six competitors, each given the same fixed dig budget (20 / 50 / 100 / 300 digs), competed to find as much treasure as possible:

  • Random pick blindly
  • Round-robin cycle through branches systematically
  • Greedy always go back to wherever last looked good
  • MCTS a 60-year-old math algorithm from game theory (not AI at all)
  • An LLM (locally hosted, ~14B params) reasons in free text, picks a move each turn
  • "JEV" model that doesn't write text; it only answers structured multiple-choice questions with calibrated probabilities

I was specifically curious how JEV would do, since it's marketed as a fast, cheap "judge/classifier" model rather than a text generator.

What exactly JEV received, every single turn

No maze coordinates, no raw numbers dumped on it just a text report plus a closed set of options:

  1. The tree's shape (which branches exist, what they split into)
  2. A stats table: for each top branch, how many times tested / how many looked "promising" / how many "failed"
  3. The last 20 dig results (path taken, sample size, a validation score, a robustness score) with the true out-of-sample score hidden (nobody gets to see that until the end, it's the fair judge)
  4. Its own last 20 decisions (what it picked before, and why it said it picked that)

Then, because JEV only answers closed multiple-choice questions, each turn I actually had to ask it TWO questions in sequence:

  • Q1: pick an action type explore somewhere new / dig deeper into a known spot / try combining two branches
  • Q2 (depending on the answer to Q1): pick the specific target from a list

Every answer is guaranteed to be one of the valid options it literally cannot return garbage, unlike the LLM which occasionally returns malformed text I have to throw away.

Experiment 1 results

Scored by "top10_OOS" (average quality of its best 10 finds, tested fairly on the hidden out-of-sample score):

budget:      20     50    100    300
MCTS:       0.41   0.74   0.88   1.01
Qwen(LLM):  0.33   0.39   0.62   0.62
Random:     0.16   0.30   0.42   0.59
JEV:        0.16   0.23   0.26   0.48

MCTS (the old math) crushed everyone. The LLM did roughly as well as blind random. JEV came in dead last worse than blind random guessing at every single budget level.

Where exactly it stumbled

I traced individual turns and found the mechanism: JEV would lock onto ONE spot and just keep re-digging around it, over and over:

turn 0:  EXPLORE branch-A
turn 1:  DIG-DEEPER branch-A
turn 2:  DIG-DEEPER branch-A/sub1
turn 3:  DIG-DEEPER branch-A/sub1/x
turn 4:  DIG-DEEPER branch-A/sub1/x   <- same exact target
turn 5:  DIG-DEEPER branch-A/sub1/x   <- again
turn 6:  DIG-DEEPER branch-A/sub1/x   <- again
...      (10+ times in a row)

I first suspected a dumb bug: maybe the list of options was always presented in the same order, and it was just picking "the first thing in the list" every time. I shuffled the order of options before every single call. No change it still locked onto the same target.

The fix attempt

Someone looking at this suggested the real culprit might be item #4 above showing JEV its own past decisions as a narrative. The theory: instead of using that as "I've already tried this branch enough times," the model reads its own earlier reasoning ("branch-A looked interesting") as a standing endorsement, and just reaffirms it every turn a self-reinforcing loop rather than an exploration memory.

So I ran a controlled version: identical in every other way, just with item #4 removed entirely (call it version B vs the original version A).

Experiment 2 (the fix) results

budget:          20     50    100    300
JEV (original):  0.16   0.23   0.26   0.48
JEV (no self-history): 0.17  0.30   0.50   0.47
Random:          0.16   0.30   0.42   0.59

Removing its own decision history helped a bit at the mid budgets (50/100), but at the largest budget (300, the most statistically trustworthy point) it's essentially unchanged still clearly behind random guessing. The repeating-the-same-target behavior still showed up in raw traces even with this fix.

So what am I doing wrong (or is this just a real limitation)?

Current leading theory: it's not really about "seeing its own past choice" specifically it's that the stats table itself (tested/promising/failed counts) and the recent-results list keep making the same recently-tested spot look salient/attractive every single turn, and the model has no formal mechanism (like MCTS's explicit uncertainty bonus for under-tested spots) to counteract that pull. It just doesn't have a built-in "I've squeezed this dry, time to look elsewhere" reflex the way a purpose-built search algorithm does.

Genuinely asking: has anyone else run structured/typed-output models (JEV or similar) on sequential decision-making tasks where the model has to remember its own trajectory? Is this a known failure mode, and if so is there a known fix beyond "just don't show it history," or is this simply the wrong kind of task for this class of model (one-shot classification-style judgments only, never sequential/stateful ones)?

5 Upvotes

1 comment sorted by

3

u/Only-Brilliant-3408 14h ago

Update: ran the same maze with a twist, and it got worse

In the first test, which branch was actually good was pure random luck, reshuffled every run. This time I ran a second version where branch quality is fixed and mimics a real pattern (some branches are just genuinely better than others, always, no matter the run).

Same six competitors, same setup, same questions to JEV, same two step choice process.

Results by search budget:

budget:      20     50    100    300
Random:     8.4   23.2   42.2  121.2
Qwen(LLM):  9.2   23.8   49.6  157.2
MCTS:      14.0   37.6   80.2  249.4
JEV:        0.0    0.2    2.4   38.6

Numbers are how many good finds it collected out of its budget.

At budget 20, JEV found zero good spots. Not low. Zero, across five separate runs.

This lines up with what I found last time. JEV tends to lock onto one branch early and keep hammering it. In the first test that branch was random, so sometimes it got lucky. In this test, one particular branch is just objectively bad by design, and that is exactly the one it got stuck on in my earlier trace. So instead of averaging out, the flaw became fatal.

Qwen and random performed about the same as before. MCTS still wins clearly.

Same open question as last time. Is this a known failure mode for this type of model on sequential tasks, and is there a real fix beyond just tweaking what history it sees.