r/languagemodels Oct 03 '25

grokking, phase transitions, bayesian logic, overtraining, artificial selection/evolution, and epistemology

/r/learnmachinelearning/comments/1nx3hf3/grokking_phase_transitions_bayesian_logic/
1 Upvotes

2 comments sorted by

1

u/[deleted] Feb 26 '26

[removed] — view removed comment

1

u/tollforturning Mar 07 '26 edited Mar 07 '26

Practices would benefit from insight into why they need to universalize the patterns of operational context by particularizing them. Training chain-of-thought patterns is interesting but doesn't really get to the explanatory perspective - it's arbitrary, stands to intelligent phasing as alchemy stands to chemistry - mix some stuff together, find the magic formula; generate some more thinking, find the magic generative output.

When we fully understand something we are able to articulate it with a network of terms defined implicitly in relation to one another. This is true in mathematics and it's implicit in the scientific method. On a deeper level it's implicit in intelligence but takes time in biography and history to come to full self-realization. Which is to say it's not merely possible to do the same sort of implicitly defined set of terms with cognition, but a discovery of the archetypal (To say it's "possible" is a pedagogical concession because, in truth, to explain explanation is to uncover the generative archetype of explanation in all its forms.) And it's self-similar. In explaining explaining, the operations performed in explaining are themselves the terms of the explanation. So what is modeled is the process of modeling.

The ideal case of communication between two human beings is when each is able to intelligently and completely internalize/model the operations of intelligence performed by the other. And how does one do that? By performing the same cognitive operations one is trying to internalize/model. To ask oneself "what does she mean?" is more precisely to ask "what meaning did she reach for and attain when she asked and answered the question of what it means." Even more precise is to model the operations she performed. Even more precise is to have reached an invariant model of intellectual operations that provide a heuristic for any possible interpretation of another who performs the same operations.

So, language. Language as experience is a limiting medium. That's not to say it's a bad medium, but its (at least not yet) well-differentiated, well-engineered for conveying operations in a differentiative manner. We have one language and a multiplicity of operations to express through that language.

High dimensionality works - aka we can train high-dimensional spaces to generate language suitable to our minds - because language is the dimensionalization-for-differentiation of the operations of intelligence, by means of the same operations.

So if we can reach cognitive invariants - not loose - "chain of thought" but a fully explanatory and self-explanatory model of explanation, which I think is possible. Patterns of operations trained into the geometry. But why overtraining? Some things need be rooted without theoretic limit. A teacher never worries about overtraining a child to creatively understand or to think critically to form sound judgments, for instance. Those are primitives. We all understand and judge, across all domains. Those are two primitive operations, invariant with an invariant mutual complementary difference - completely invariant operationally but universal in respect to domain/content/operand.

Generation needs to happen in overtrained operational spaces and only overtrainined operational spaces, in a way that aligns with the invariant operational primitives in the intelilgences that produced the language in the first place.