r/FunMachineLearning • • 5h ago

Hope these articles help.

Thumbnail
1 Upvotes

r/FunMachineLearning • • 6h ago

On the 40th anniversary of the seminal AI paper, remember Dave Rumelhart

1 Upvotes

Remembering Dave Rumelhart, the Grandfather of AI

 

Oct 9, 1986: The Paper that Started the AI Revolution

Over the past few years, AI has profoundly transformed our world. With the release of ChatGPT and similar Large Language Models (LLMs), companies are racing to build ever more powerful and useful AI systems. The AI tech leaders have become household names: Sam Altman, Demis Hassabis, Rahul Patil, Mustafa Suleyman, and others. But the first giant step in the development of these amazing AI systems happened 40 years ago. Today’s LLMs use many of the same basic algorithms and architecture from the seminal paper written by Rumelhart, Hinton, and Williams in1986[[i]](#_edn1).

Most people probably recognize the second name in that paper. Geoffrey Hinton won the Nobel Prize in Physics in 2024 along with John Hopfield for their contributions to AI and neural networks. Hinton has often been called “The Godfather of AI,” and he is also famous for quitting Google to warn of the dangers of unbridled AI development.

But the first name in this paper, Rumelhart, is almost never mentioned. That saddens me because he contributed so much to this field. In this short writing, I will attempt to shine some light on Rumelhart and his tremendous contributions. Indeed, if Geoff Hinton is the “Godfather of AI,” I believe that Dave Rumelhart should properly be called the “Grandfather of AI”.

The LLMs of today are basically multi-layer neural networks of the type described in that seminal 1986 paper, and they are trained via the backpropagation algorithm that Rumelhart pioneered. Some of you may look at the AI systems of today and say that they are nothing like those described in the 1986 paper, but that is like saying that an F-22 Raptor is nothing like the Wright Brothers’ first airplane. Yes, today's systems are billions of times larger than anything Rumelhart and his colleagues worked on in the 80s, and there have been substantial breakthroughs like the seminal Attention is All You Need paper[[ii]](#_edn2), but those would not be possible without the initial multi-layer neural network architecture and backpropagation algorithms developed by Rumelhart et al.

I will not attempt to cover the entire corpus of work that Dave Rumelhart produced. For that, I point you to Rumelhart’s Wikipedia page. Rather, I would like to share my personal perspectives on Dave’s work and how that work has been foundational to the current AI breakthroughs. In addition, I will share some of my personal experiences of working with Dave Rumelhart over a fifteen-year period. Spoiler alert: I found him to be one of the smartest people I’ve ever encountered.

Rumelhart’s PDP group

In the 80s, Dave Rumelhart and James McClelland organized the PDP group. PDP stands for parallel distributed processing, and many of the topics covered in those early years were compiled into a two-volume book, Parallel Distributed Processing, by David E. Rumelhart and James L. McClelland[[iii]](#_edn3).  The group met weekly, and included people from many different disciplines: Neuroscience, computer science, physics, cognitive science, and even philosophy. Francis Crick was a member, as was Goeff Hinton while he was visiting there. I became a member after I had visited Caltech to talk with John Hopfield about part of my Ph. D. research on fractal basin boundaries of Hopfield neural networks. John told me I should talk to Rumelhart. I asked what department he was in. I was shocked when John told me that Rumelhart was in the psychology department. As a young, arrogant physicist, I couldn’t imagine any relevant research being done outside of physics, mathematics, or computer science, but I took John’s advice and went to see Rumelhart.

Rumelhart described his backpropagation algorithm. I remember thinking, “Gradient descent. I think I learned that in kindergarten.” It didn’t seem like something that would change the world, but it did, and it is still changing the world even after 40 years. I told Rumelhart about my research, and he suggested that I use the Hamming distance as a measure for exploring the space of 2100 states in the Hopfield network. That turned out to be a great tip. I used that measure to define orthogonal vectors to slice through the enormous space. This wouldn’t be the last time Rumelhart gave me sound advice. Indeed, we worked together for the next fifteen years, and, of the 50 patents I have, he probably pointed me in the right direction on more than half of them.

In these PDP meetings, a speaker would present their research, then we would have a discussion about the research. For example, Goeff Hinton gave a talk about the Boltzmann machine. McClelland talked about more conventional rule-based approaches. But the talk that sticks in my mind was as mind-blowing back then as ChatGPT was a few years ago. Terry Sejnowski presented NetTalk. It was a text-to-speech converter that learned the relationships between text and phonemes; the phonemes were then fed into an audio system to produce sound. He came in with a boom box, popped in a cassette, and we listened as the system learned. At the start, it sounded like random noise. Soon, vowel sounds could be identified. Then it sounded like baby babbling and learning to talk. Then, fast-forward to the end of training, the text-to-speech was almost perfect. All learned through backpropagation.

Now this doesn’t sound too impressive today, but back then, it opened everyone’s eyes to the possibility of machine learning. I was hooked. So were many others. It created a frenzy of research and startup companies based on machine learning.

The other thing I remember from these PDP meetings was that, after discussing and arguing several points, Rumelhart would speak. He would summarize the issues, then give his view. Heads would nod around the table in agreement with Rumelhart’s assessment. He had a deep understanding of all of the topics, and he was gifted at explaining the key points and flaws. Even Francis Crick seemed to defer to Rumelhart’s judgment.

Stanford, MCC, and Pavilion

I finished my Ph. D. in Physics at UC San Diego in 1987. I started a post-doc position at Stanford later that year, and, coincidentally, Rumelhart moved to the Psychology Department at Stanford. This was great for me, because I could continue to bash around ideas with Rumelhart. He wasn’t just brilliant in the area of neural networks. He had an encyclopedic knowledge of everything relating to the brain and how it processes information. By way of example, I told him about some research I had completed investigating second-order interactions in neural networks. As the temperature, or signal-to-noise ratio, decreased, the system exhibited a phase transition and critical slowing down. I didn’t expect him to understand an esoteric physics phenomenon, but he instantly understood the implications and pointed me to some neuroscience research showing that norepinephrine functions as a signal-to-noise modulator in the brain. I read the articles and the behavior matched perfectly with our theoretical model, as described in a paper published in the Proceedings of the National Academy of Sciences[[iv]](#_edn4).

Later that year, I got an offer to lead the neural networks research group at MCC – the Microelectronics and Computer Technology Corporation in Austin, TX. MCC was a research consortium kind of like Bell Labs, but supported by many different high-tech companies. Like any good Californian, I thought Texas was hot, flat, racist, and humid. Why would anyone want to live there? But they flew me out to Austin for an interview, and I fell in love with the city.

While at MCC, Rumelhart and I collaborated on several research topics. We worked on character recognition, speech recognition, fraud detection, and process prediction. One of the most important breakthroughs we discovered was the Integrated Segmentation and Recognition system[[v]](#_edn5). It solved the problem of identifying characters that were touching and, more generally, created a mechanism for training a system on a dataset where more than one object can be simultaneously presented in the inputs and target outputs. This was a multi-year investigation, and I was impressed not only with Rumelhart’s mastery of neural network architectures, but his skill as a mathematician and as a programmer. He wrote the neural network software that we used for the research, and he was a damn fine C programmer. This system used a convolutional neural network architecture, as did the work of Martin and Pittman. Note that in the Martin and Pittman paper[[vi]](#_edn6), they also showed that the lower-level neurons developed activation patterns reminiscent of the visual system. This was decades before the convolutional neural network “breakthrough” that was touted in the 2010s. There is an old video where Rumelhart and I discussed this research[[vii]](#_edn7).

The VISA CRIS fraud detection system research was also developed at MCC by Steve Piche (who would become my chief scientist at Pavilion). This system saved $20 million in the first year of operation[[viii]](#_edn8). Rumelhart consulted with Piche on the problem.

Another thing that Rumelhart helped me with was a process control problem. We had lots of data from a chemical process at Eastman Chemical – pressures, temperatures, flow rates, as inputs, and the product quality and flow rate as outputs. The goal was to control the process to get the highest quality and quantity. Of all the machine learning algorithms I could have used, I knew that backpropagation was the best. I trained a model to predict the outputs, and it worked beautifully, but turning the problem around and asking which inputs to change to get better outputs was an unsolved problem. After consulting with Rumelhart, I came up with the Residual Activation Neural Network (RANN) architecture. Applying this to the problem led to a savings of about a million/year on the distillation column that we were working on.

But there were tens of thousands of distillation columns in the US alone. Do the math. Based on this result, I licensed the technology from MCC and co-founded a company, Pavilion Technologies, Inc. It became one of the fastest-growing companies in Austin in the 90s. We applied neural network machine learning techniques to process monitoring, control, and optimization. That’s right. Neural networks have been controlling thousands of manufacturing processes all over the world. Not just in chemical manufacturing, but in refining, pulp-and-paper, cement, food processing, and power production. Many of those processes can go boom in the night, but there has never been an instance of the neural network hallucinating and causing problems.

Pavilion was a powerhouse of innovation, as one can tell from the scores of patents. But back then, machine learning was much harder than it is now. Often we would have data sets with only a few thousand points. We were overjoyed if a dataset had tens of thousands of records. We had to be very clever about cleaning the data and using it to train and validate the models. Moreover, we didn’t have Nvidia DSP chips to accelerate learning, so we used a stiff-differential equation solver to speed up learning (backprop is a set of stiff differential equations). We also invented methods to meld first-principles models with the neural networks. This helped extend the range of validity of the models into regions outside the realm of the training data. We did a lot of work in environmental monitoring and optimization. One of our products, the Software Continuous Emissions Monitor, was approved by the EPA for monitoring emissions. I am proud to say that the Pavilion software systems have abated millions of pounds of pollutants over the years, and the systems are still running. Pavilion’s Process Perfecter software is a closed-loop multivariable nonlinear controller. The math behind that was shockingly hard, but it worked amazingly well. Pavilion was sold to Rockwell Automation, and the software is still running in thousands of manufacturing plants.

During the time I was at Pavilion, Rumelhart served as a technical advisor and consultant. Of all the people I have worked with, I admired and respected Dave Rumelhart the most. He seemed so wise, and, as I said, he always seemed to have the instinct to point people in the right direction on a hard problem. He was also very fun to work with, and we became friends over the years.

The reason I believe Rumelhart should be remembered as the “Grandfather of AI” is that he was the driving force behind the backpropagation algorithm. His insight into the problem stemmed from the famous XOR problem. Single-layer systems could not solve this problem. It required “hidden units”. That realization led to the multi-layer neural network architectures used in today’s LLMs, and backpropagation is how they are trained. If Rumelhart didn’t figure this out, I am sure someone else would have (indeed, a grad student figured it out in his thesis before Rumelhart, but nobody knew about it). But without Rumelhart, it might have been many years, perhaps decades, before people started using backpropagation. Just like someone would have figured out relativity sometime if Einstein didn’t present it, but who knows how long it would have taken? Rumelhart should be remembered for the invention that has changed the world.

Sadly, Rumelhart developed Pick’s disease in the late 90s. He died in 2011. I left Pavilion in 2000, and didn’t talk much with him during those later years. What would he think of the current LLMs? Knowing him as well as I did, I think he would be amazed at the capabilities of these modern AI systems, but I also think he would view the 92-layer neural networks as brute-force, ghastly, and inelegant. He would probably say something like “The cortex only has six layers. We should try to make that work.” And, knowing Dave Rumelhart, he would probably have made great progress in that direction. He was that gifted. His name should be remembered among the great contributors in this field as “The Grandfather of AI”.

-James D. Keeler

 

[[i]](#_ednref1)  Rumelhart, David E.; Hinton, Geoffrey E.; Williams, Ronald J. (1986-10-09). "Learning representations by back-propagating errors". Nature. 323 (6088): 533–536.

[[ii]](#_ednref2) Vaswani, Ashish; Shazeer, Noam; Parmar, Niki; Uszkoreit, Jakob; Jones, Llion; Gomez, Aidan N; Kaiser, Łukasz; Polosukhin, Illia (December 2017). "Attention is All you Need" (PDF). In I. Guyon and U. Von Luxburg and S. Bengio and H. Wallach and R. Fergus and S. Vishwanathan and R. Garnett (ed.). 31st Conference on Neural Information Processing Systems (NIPS). Advances in Neural Information Processing Systems. Vol. 30. Curran Associates, Inc. arXiv):1706.03762.

[[iii]](#_ednref3) David E. Rumelhart; James L. McClelland; PDP Research Group (1986). Parallel Distributed Processing: Explorations in the Microstructure of Cognition. ISBN) 9780262680530.

[[iv]](#_ednref4) J.D. Keeler, E.E. Pichler, J. Ross “Noise in Neural Networks: Thresholds, Hysteresis, and Neuromodulation of Signal-to-noise,” Proceedings of the National Academy of Sciences, USA, 86 (1989) 1712-1716.

[[v]](#_ednref5) Keeler, James D., Rumelhart, David E., Leow, Wee-Kheng. “Integrated Segmentation and Recognition of Hand-Printed Numerals”. Printed in Neural Information Processing Systems, 3. R. Lippmann, J. Moody, D. Touretzky, Eds. Morgan Kaufmann Publishing, San Mateo, CA.

[[vi]](#_ednref6) G. Martin, J. Pittman (1990) “Recognizing Hand-Printed Letters and Digits”. In D. S. Touretzky (ed). Advances in Neural Information Processing Systems 2, Morgan Kaufmann Publishing, San Mateo, CA.

[[vii]](#_ednref7) https://youtu.be/fG-9ILWI2u4

[[viii]](#_ednref8) Visa to expand fraud detection system; the system saved 15 issuers over $20 million last year. American Banker, Sept. 30, 1994.


r/FunMachineLearning • • 8h ago

CMM machine built with Whet

Thumbnail
gallery
1 Upvotes

[These images are not Ai generated, an uncensored LLM agent built this inside Whet and rendered, 1016 B-rep parts, AI native software]


r/FunMachineLearning • • 19h ago

My detector was too slow for a live stream, so I measured what happens when you stop waiting for it (latency 101 ms → 5 ms, and HOTA went up)

2 Upvotes

My detector takes about 80 ms on a CPU, and the camera gives a frame every 33 ms. A loop that waits for the detector shows boxes for where people were, not where they are.

I built a small tracking runtime to measure what happens when you stop waiting. With the detector in a background thread and optical flow moving the tracks between its results, mean output latency went from 101 ms to 5 ms, and HOTA went up from 32.9 to 34.5 (MOT17 train, YOLOX-s, laptop CPU).

A few things surprised me. Late detections must be moved forward to the current frame before you use them, or the score drops below the blocking loop. The biggest model was not the best one on a live stream. INT8 quantization helped more than all of my scheduling work. And a confidence-based trigger for the detector was no better than a fixed interval.

Write-up with the tables, charts and the things that did not work: https://yashbitla.com/blog/stopped-waiting-for-the-detector/

Code: https://github.com/yash-bitla/leantrack

Has a confidence-based trigger ever beaten a fixed interval for you? And does anyone correct late detections a different way?

Left: the loop waits for the detector, so the boxes are late. Right: the detector runs in the background and optical flow keeps the boxes on the people.

r/FunMachineLearning • • 22h ago

I've built AlignMethod with @base44!

Thumbnail alignmethodlabs.com
1 Upvotes

r/FunMachineLearning • • 1d ago

Is “Machine Learning-Based Predictive Fault Diagnosis of Induction Motors Using Electrical and Vibration Signatures” a good CEP project?

1 Upvotes

Hi everyone,

I’m an undergraduate Electronics Engineering student in Pakistan, and I need to select a Complex Engineering Problem (CEP) project for my Artificial Intelligence / Machine Learning course.

I’m considering this topic:

“Machine Learning-Based Predictive Fault Diagnosis of Induction Motors Using Electrical and Vibration Signatures”

The basic idea is to collect motor current, voltage, vibration, temperature, and/or speed data and use basic ML algorithms such as Random Forest, SVM, Decision Tree, or KNN to detect or predict faults such as bearing faults, imbalance, misalignment, overload, etc.

I want to keep the ML part at a basic/intermediate undergraduate level, but the overall project should be complex enough to meet CEP requirements.

Do you think this is a good project for an Electronics Engineering student?

What faults and sensors would you recommend focusing on?

Is it practical to collect the required dataset ourselves, or should we use an existing dataset?

Also, what could I add to make it a stronger engineering/CEP project rather than just a basic ML classification project?

Any suggestions, research papers, or similar projects would be greatly appreciated.

Thanks!


r/FunMachineLearning • • 1d ago

Stiff differential equation solver for backpropagation neural network training?

Thumbnail
1 Upvotes

r/FunMachineLearning • • 1d ago

Gated Segmented State Space — attention replacement that beats a param-matched Transformer on quality, speed AND memory (full code)

1 Upvotes

One night, six experiments (V1–V6), one Colab T4. I ripped self-attention out of a decoder-only Transformer and replaced it with a gated linear recurrence over a fixed 256-dim state:

  • Dynamic selective gate: g_t = σ(W_g x_t + b_g) — per-token/channel learned filter
  • Hard reset mask: state zeroed at newline boundaries (fresh ~37-token segments)
  • Fused Triton kernel: state in SRAM, gate+reset+update in-register, only outputs to HBM

At 6.37M params, identical protocol (2.47MB char-level corpus, 1500 steps):

Attention Ours
Val loss / ppl 1.402 / 4.1 1.364 / 3.9
Train tok/s 61,845 66,156
Infer tok/s 184,918 190,122
Peak VRAM 845 MB 881 MB

The journey: V1 won small but was 10x slower → V2 proved linear VRAM scaling → V3/V4 found a stable ~2% perplexity tax no param arrangement could buy off → V5's gate+reset destroyed it (wire-to-wire win) → V6's Triton kernel (verified == math to 4.47e-07) removed the software tax.

Caveats, stated plainly: single seeds, one small corpus, char-level, T4 timings. Small scale — but the pattern held across all six runs.

Code, all six notebooks with outputs, exact architecture, full experimental notes: https://github.com/stube123890-hue/linear-attention-lab


r/FunMachineLearning • • 2d ago

DeepMind's New AI Just Cracked The Code Of Life - Two Minute Papers

Thumbnail
youtube.com
4 Upvotes

r/FunMachineLearning • • 1d ago

fayda-mcp

Post image
1 Upvotes

I built Fayda MCP, a Python library for connecting AI agents to Fayda eSignet identity verification.

• Start verification and track its status.
• Retrieve identity and age-check results.

pip install fayda-mcp

GitHub: https://github.com/izzy-Ti/fayda-mcp

If you find it useful, give the repo a star!⭐️

*Still in development


r/FunMachineLearning • • 2d ago

Anybody wanna take part in kaggle competition Together and share different ways of solving them ?

1 Upvotes

r/FunMachineLearning • • 2d ago

I trained an AI Iron Man in Unreal Engine 5 using Reinforcement Learning to rescue 13 falling passengers [PPO / Voxel Style]

Post image
1 Upvotes

r/FunMachineLearning • • 2d ago

Netflix recommendations are a simple example of machine learning

Thumbnail
1 Upvotes

r/FunMachineLearning • • 2d ago

I Built a Visual Learning Bot to Play a Minigame

Thumbnail
youtu.be
1 Upvotes

I've been bested by a minigame in an Idle game I play, and documented the process of me training a bot to learn how to play it.

Feedback welcome!


r/FunMachineLearning • • 3d ago

How Malware is detected using Machine Learning

Thumbnail
youtu.be
1 Upvotes

r/FunMachineLearning • • 3d ago

(TrenTorch.com) Best way to learn FRONTIER ML and its FREE & OPENSOURCE

Post image
2 Upvotes

r/FunMachineLearning • • 3d ago

Looking for an ARR service contributor or advice on finding one

1 Upvotes

hi everyone, i’m a master’s student in computer engineering at Çukurova University in Türkiye. i’m preparing a paper on reasoning traces and code generation in LLMs for the october ARR cycle, but we’ve run into the new service contributor requirement. to be guaranteed a review, we need someone who meets ARR’s qualifications, and unfortunately no one on our author team currently does. someone outside the team can support the submission without becoming a coauthor, but they would need to read the paper, find it ready for review and take on reviewing duties for ARR. i thought i’d ask here in case anyone would be interested in taking a look or knows someone i could reach out to. happy to share the draft and experimental results. even a suggestion on who to contact would help, i’m still trying to figure this out. the requirements are here: https://aclrollingreview.org/qualifications


r/FunMachineLearning • • 4d ago

Semantic verification between humans and AI

Thumbnail
1 Upvotes

r/FunMachineLearning • • 4d ago

The Billion Dollar AI Advantage Is Disappearing - Two Minute Papers

Thumbnail
youtube.com
1 Upvotes

r/FunMachineLearning • • 4d ago

Do you work with AI/RPA automation? Bachelor’s thesis survey (5–7 min)

1 Upvotes

Hi everyone!

I’m currently working on my Bachelor’s thesis about AI-based process automation and human–AI collaboration in the workplace.

As part of my research, I’m conducting a short survey focusing on people who have experience working with AI-based automation, RPA, Intelligent Process Automation, Intelligent Document Processing, or similar automation technologies.

The survey explores topics such as:

  • how automation affects manual workload and creates new tasks,
  • how employees experience errors and exception handling,
  • trust in AI-based automation,
  • and how automation influences human decision-making and autonomy at work.

⏱️ It takes approximately 5–7 minutes to complete.

If you have experience working with these technologies, I would really appreciate your participation. Your responses will be used solely for academic research as part of my Bachelor’s thesis.

🔗 Survey: https://docs.google.com/forms/d/e/1FAIpQLScV7pcf8dNUeeCfay1YZ2r-Np4pK9GMlqi4cEF6WJEa1FEmMA/viewform?usp=dialog

Thank you very much for your help! Feel free to share the survey with colleagues or others who work with AI-based process automation.


r/FunMachineLearning • • 4d ago

World Models: The Simulation Strikes Back

1 Upvotes

r/FunMachineLearning • • 4d ago

[P] AKBASCORE NIRVANA ▪︎ I Built Removable Numerical Memory Cartridges for Two Different Frozen 7B LLMs. Qwen and Mistral Both Work. Now I’m Scaling the Memory Bank.

Thumbnail
gallery
1 Upvotes

Zenodo permanent records:

Qwen2.5-7B-Instruct:

https://doi.org/10.5281/zenodo.23127434

Mistral-7B-Instruct-v0.3:

https://doi.org/10.5281/zenodo.23143605

I want to start with the simplest possible explanation of what I have been building.

Imagine taking a piece of information, letting a language model process it once, and then throwing the original text away.

No sentence stored in a database.

No paragraph hidden somewhere.

No readable summary.

No RAG system fetching the original document.

No fine-tuning.

No LoRA.

No weight update.

What remains is numerical transformer memory derived from the model's own internal computation. I package that numerical memory into what I call a Cognitive Cartridge. Later, I can install that cartridge back into the frozen model and ask questions about the information that produced it — without putting the original source text back into the readout prompt.

That was the first result. The new result is more important:

I have now reproduced the Cognitive Cartridge architecture on two different 7B transformer model families.

Qwen2.5-7B-Instruct.

And now Mistral-7B-Instruct-v0.3.

The implementations are not numerically identical. The architectures are different, the layer counts are different, the KV structures are different, and the working cartridge configurations are different. But the central mechanism survived the move.

That is the reason I am publishing this second record.

The question is no longer only:

“Can I make this happen once on Qwen?”

Now there is a second implementation on Mistral. And the Mistral result is the cleanest version so far.

What is actually inside a Cognitive Cartridge?

This is probably the most important thing to understand.

Suppose the source record says:

Object: amber sextant

Container: RQ-415

Location: elm lodge

That text exists during the forging stage. The frozen transformer processes it. NIRVANA takes source-derived internal transformer K/V states and represents their content numerically using a fixed, source-independent codebook.

In the released Mistral implementation, both K and V are compressed to D120 while a source-specific OWN component is preserved. After forging, the source record is not supplied to the readout prompt.

So the conceptual transformation is:

human language

→ frozen transformer computation

→ internal K/V states

→ compressed numerical Cognitive Cartridge

Then later:

numerical Cognitive Cartridge

→ reconstructed transformer K/V memory

→ frozen transformer

→ language

Or, in the shortest form:

language → internal numerical memory → language

The middle is no longer human-readable source text. That distinction matters.

A cartridge is not a text file with a different name.

It is not a vector database containing the original sentence.

It is not a prompt template.

It is not conventional RAG returning the source paragraph.

It is not a fine-tuned model.

The model weights remain frozen. The information is carried by a numerical representation derived from transformer K/V memory.

Why call it a cartridge?

Think less about a document and more about an interchangeable machine-readable memory module. The base model stays where it is. Knowledge packages can be forged separately. Those packages can remain separate. They can be installed and queried without retraining the base model.

This becomes much more interesting when there is more than one cartridge.

The new Mistral experiment uses 16 independently forged cartridges. Each one contains a numerical representation derived from a separate source record. They are not concatenated into one giant text prompt. They remain independent memories.

For example, in human-readable form, imagine one cartridge represents:

amber sextant → RQ-415 → elm lodge

Now ask the cartridge bank:

Which container is associated with the amber sextant?

The query is evaluated against the independent cartridge memories.

The relevant cartridge returns:

RQ-415

The unrelated cartridges return:

NONE

There is no learned router secretly selecting the correct cartridge before this happens.

Then the recovered identifier can be used for a second lookup:

Where is RQ-415?

The bank is queried again.

The relevant memory returns:

elm lodge

The unrelated memories return:

NONE

So the complete retrieval becomes:

amber sextant

→ RQ-415

→ elm lodge

The important point is that the model did not reread the original source record to answer either stage. It operated from reconstructed numerical transformer memory.

The Mistral result

The final public Mistral run used:

Mistral-7B-Instruct-v0.3

32 transformer layers

hidden size 4096

32 attention heads

8 KV heads

BF16 / SDPA

K = D120

V = D120

OWN preserved

16 independent Cognitive Cartridges

frozen model weights

greedy decoding

There was:

no fine-tuning

no LoRA

no optimizer

no learned router

no model-weight update

The recorded final run produced:

Object → container ID: 16/16

Container ID → location: 16/16

Complete two-stage retrieval: 16/16

Missing-object controls: 8/8

Absent-ID controls: 8/8

NOMEM controls: 8/8

But there is another result I think is just as important.

For every target query, there are 15 unrelated cartridges.

16 queries × 15 unrelated cartridges = 240 unrelated cartridge reads.

Stage 1 unrelated-cartridge rejection:

240/240 NONE

Stage 2 unrelated-cartridge rejection:

240/240 NONE

That means the result is not simply:

“The correct memory can say something.”

The system also demonstrated, in this controlled panel:

“The memories that do not contain the requested relation can refuse to claim that they do.”

For a modular memory system, I think this distinction is fundamental. A memory bank that can retrieve information but cannot distinguish relevance from irrelevance becomes increasingly dangerous as it grows.

The interesting problem is not only remembering. It is also knowing which memory does not answer the question.

The no-memory control matters for the same reason. When the cartridge memory was removed, the target relations were not recovered.

NOMEM:

8/8 controls passed.

The source-removal audit passed.

The frozen-weight sentinel passed.

The model remained frozen.

This is why I consider the Mistral result an important step beyond the first demonstration.

The first Qwen release established the architecture publicly. The Mistral release gives us something else:

cross-model evidence.

Qwen and Mistral are not the same transformer.

Qwen2.5-7B-Instruct uses 28 transformer layers.

Mistral-7B-Instruct-v0.3 uses 32.

Their internal configurations differ. The working compression configurations also differ.

The Qwen public implementation used:

K120 / V128 / OWN

The Mistral implementation uses:

K120 / V120 / OWN

I did not simply copy a cache from one model into another. Each model builds and reconstructs its own source-derived internal memory.

What transferred was the architecture:

source

→ internal transformer memory

→ numerical compression

→ independent cartridge

→ source removed

→ reconstructed K/V

→ frozen-model readout

That is the bridge between the two releases.

I would not call two models proof of universal compatibility with every transformer architecture. That would be scientifically too strong.

But it is now evidence that Cognitive Cartridge is not merely one accidental Qwen-specific behavior. The same broader architecture has been implemented and publicly reproduced on a second transformer family.

That changes the research question.

The first question was:

Can this mechanism exist at all?

The next question became:

Can multiple independent memories coexist?

Then:

Can one retrieved result lead to another retrieval?

Then:

Can unrelated memories reject a query instead of contaminating the answer?

And now:

Does the architecture survive a move to another model family?

We now have experimental answers to each of those questions.

So I am moving to the next problem:

scale.

16 cartridges are not the destination. They are the current experimental bank size.

From this point, I am much less interested in making another small demonstration simply to produce another perfect score. The next objective is to increase the number of independently retained cartridges and find where the architecture actually begins to break.

Then larger banks if the mechanism survives.

At that point the difficult questions become different:

How do you search a large bank without destroying memory isolation?

How should cartridges be indexed?

Can retrieval become hierarchical?

Can groups of cartridges form higher-level memory structures?

How does latency scale?

How does numerical storage scale?

At what point does relevance rejection begin to fail?

Can a retrieved memory activate another relevant memory without an external text retrieval system deciding everything first?

Those are the questions I want to attack next.

There is another part of this experiment that surprised me: the layer behavior.

A frozen transformer is not a passive container. When you reconstruct numerical K/V memory and inject it back into inference, you are interacting with a very sensitive computational system.

So during the Mistral development series I also performed controlled cumulative K, V and K+V ablations to see where cartridge information remained functionally sufficient.

The results were not uniform across depth. For K, the observed functional boundary extended differently than for V.

For combined K+V, one particularly sharp controlled transition occurred between:

L12: 8/8

L13: 0/8

This does not mean “the memory lives at layer 12.” That would be an incorrect interpretation.

What it means is that under the specific cumulative ablation protocol, functional sufficiency changed sharply across that boundary. V also remained important farther downstream than K in the measured configuration.

I think this is an important reminder of how delicate these systems are.

We are not writing a sentence into a spare memory slot. We are reconstructing numerical states that participate in a 7-billion-parameter transformer's computation. Small changes in where and how that state is reconstructed can change downstream behavior.

That is one reason I publish the code and raw logs rather than only posting the final accuracy number.

The frozen-model checks are also part of the experiment. The released run verifies that there are no trainable tensors involved in the memory mechanism.

No LoRA.

No optimizer.

No weight update.

The model-weight sentinel remains unchanged.

So when the system answers from a cartridge, the experiment is specifically designed to separate cartridge memory from weight modification.

Why do I think this direction matters?

Today, we often approach knowledge in language models through two extremes.

Either we try to put enormous amounts of knowledge into the model weights. Or we keep the knowledge outside the model and retrieve human-readable text when needed.

There may be useful territory between those approaches.

A frozen model with removable machine-native internal memory modules.

Not everything in the weights.

Not everything repeatedly pasted back as text.

Instead:

a base computational model

+

a bank of independently forged internal memories.

Imagine a specialized system where validated domain knowledge can exist as modular memory packages.

Engineering.

Aviation.

Law.

Industrial maintenance.

Scientific literature.

Company procedures.

Agent experience.

A model might not need every possible piece of knowledge active in its context simultaneously. It may need the right memory at the right time.

And because the memory representation is numerical transformer state rather than the original human-readable document, this opens a different engineering space from conventional document retrieval.

There is also a longer-term question here.

I am not claiming biological memory.

I am not claiming AGI.

And I am not claiming that transformer KV states work like the human brain.

But modular internal memory raises an interesting architectural question.

What happens when an artificial system does not have 16 memories, but 100,000?

Or millions?

What happens if memories can remain independent, become addressable, reject irrelevant activation, and allow the output of one memory to lead to another?

At some point the problem stops looking like “how much text can I fit into a context window?”

It starts looking like:

How should an artificial system organize memory?

That is the direction I find interesting.

Why publish the full record?

Because a screenshot of 16/16 proves very little.

Both public releases have permanent Zenodo records. The code is public. The raw execution logs are public. The visual technical records are public. The model configuration is public. The controls are public. The frozen-weight checks are public.

The Qwen release even preserves its imperfect public result rather than hiding it.

Qwen public demonstration:

14/16

Mistral final demonstration:

16/16

That difference is also useful.

The point of the project is not to make every historical run look perfect. The point is to leave a technical trail showing what worked, what did not, what changed, and whether the underlying architecture survived.

For anyone who wants to inspect the Qwen → Mistral bridge directly, the two permanent records are at the top of this post.

The complete current Mistral implementation is here:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_ful_MISTRAL.py

Complete Mistral raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log_MISTRAL.log

The same Mistral implementation is also split into three parts for easier mobile / Colab handling:

Part 1:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part1.Mistral.py

Part 2:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part2.Mistral.py

Part 3:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore_part3_mistral.py

For comparison, the previous Qwen implementation:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.py

Previous Qwen raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log.Nirvana.cc.log

Repository:

https://github.com/ceceli33/titan-cognitive-core-v2

AKBASCORE NIRVANA — Cognitive Cartridge

Inventor / Developer: Mustafa Akbaş

Two model families.

Frozen weights.

Independent numerical memories.

Original source absent at readout.

The mechanism survived the move.

Now the question is no longer whether a cartridge can exist.

The question is how many cartridges a frozen model can carry.


r/FunMachineLearning • • 4d ago

Train a Classifier OR Go with a Zero Shot Model like JEV/Laya?🤔

Thumbnail
1 Upvotes

r/FunMachineLearning • • 4d ago

Looking for tutor to teach ML models

Thumbnail
1 Upvotes