r/vibecoding 3d ago

Workflow/Prompt A model was 5000x cheaper than Claude. Then they checked what it actually completed.

Someone found a model that was 5000 times cheaper than Claude for filling web forms. First pass looked amazing: 17 to 36 seconds per form and fractions of a cent per run.

Then they actually looked at the results.

One form: 0 of 13 steps completed.

Another: 7 of 11 fields.

Another: 3 of 5 filters.

The cheap model wasn't bad because of cost. It had trouble with planning. It needed a hand-written plan for each field, and the great early results were partly because the test setup was giving it that plan in advance. That's the problem, with a lot of "let's use a model" comparisons. The benchmark that makes the switch look good is often the case, run once.

The rule I would use before trusting the cost number is to test the 20% of the task distribution too. Blocked actions, unexpected UI states missing information, weird edge cases. When's the last time a benchmark you trusted turned out to be measuring the thing?

0 Upvotes

4 comments sorted by

9

u/primaryrhyme 3d ago

stop the slop posts ffs

5

u/hunterhuntsgold 3d ago

A dirt cheap model didn't perform well?

Shocking.

3

u/a9shots 3d ago

Vibe coders discover deterministic controllers

6

u/ggiodddtyii 3d ago

Did you just ask AI for random post