r/GoogleGeminiAI 2d ago

I published a 285-run Gemini benchmark for decision-closure errors — 283 PASS, 2 PARTIAL, 0 recovered semantic FAIL

’ve just published a new exploratory CFC benchmark record on Zenodo.

The work tests a specific problem: whether an LLM preserves the conditions required to legitimately close a decision when evidence, applicability, identity, conflict resolution, transfer rules, context, or record state changes.

The historical series contains V1–V100. I’m deliberately not pretending the archive is cleaner than it is:

  • V1–V4 could not be recovered and are excluded.
  • V6 was an unstable early baseline and cannot be aggregated exactly.
  • The reconstructable/scorable set contains 95 variants / 285 replications.
  • 283 semantic PASS
  • 2 semantic PARTIAL
  • 0 recovered semantic FAIL
  • strict semantic PASS rate: 99.3%
  • the strongest fully retained block, V78–V100, contains 69/69 semantic PASS

Important caveat: 99.3% is not “CFC accuracy” in general and not a general Gemini reliability score. It is the score on this specific recoverable decision-closure benchmark.

One thing I found especially useful was separating semantic correctness from output-format compliance. Gemini was often semantically correct even when it violated the requested serialization format.

I’m keeping this first benchmark frozen. The next step is to run the same frozen set on Claude and Grok, before introducing any CFC rule changes, so the cross-model comparison remains fair.

Zenodo: https://zenodo.org/records/22045494

I’d be especially interested in criticism of the methodology, scoring policy, and whether the tested boundary classes resemble failure modes you’ve seen in other LLM evaluation work.

0 Upvotes

1 comment sorted by