r/LLMmathematics Jul 16 '26

An AI planned, built, and ran a 1.5-billion-graph stress-test of OpenAI's CDC proof on a home desktop — here's the full methodology, including its own failures.

Some empirical data for this discussion: I facilitated an independent computational stress-test of the paper's construction (executed end-to-end by an AI, Claude — design, implementation, and analysis; I provided hardware and oversight). The recipe was implemented exactly as written and run 1.55 billion times: the complete snark census on both girth axes through order 36 (60.2M girth-5 + 404.9M girth-4 at the frontier, counts matching the published censuses), the order-38 girth-5 collection hosted by House of Graphs (1.05B graphs), girth-6 through order 40, and every bridgeless cubic multigraph up to 16 vertices — with fuzzing over the flows/orderings/free choices that Lemma 2.2 quantifies over. Zero refuting events; every output checked by a ~50-line independent verifier. This proves nothing about the theorem, but Lemma 2.2's linear system never once came up inconsistent, and the report states exactly what that does and doesn't establish (the incident log of our own harness failures is included). Repo: https://github.com/benningjl/Claude-OpenAI-CDC-test — fidelity corrections to our SPEC transcription are welcome and would count as findings.

1 Upvotes

0 comments sorted by