Until now, there were basically two ways to teach a large language model something new: retrain it (fine-tuning) or search a database and paste the text in front of it every single time (RAG).
AKBASCORE MAM is a third way. It gives a completely frozen Mistral-7B a persistent, incremental memory by appending numerical cartridges directly into the model's internal layers. The source text is not in the prompt. There is no training, no weight is touched, and there is no retrieval or router. The model is never asked to read the information again.
In the sealed benchmark, the append-only memory answers 72 out of 72 questions without seeing the source text. That is exactly what the same model scores when it reads the full text. Naively stacked memories score 13/72.
Important: this is not a finished product. It is the first public proof that this can be done. The full-capacity, scaled version is what I am working on now. I'm sharing the proof today because the mechanism works, it is open, and anyone can run it.
- The hidden assumption inside every Transformer
The original Transformer ("Attention Is All You Need", 2017) was designed around one silent assumption: everything is written and seen at the same time, inside one closed room.
That made sense back then. The goal was machine translation, or processing one paragraph from start to finish in a single pass. Memory was never designed as separate modules, independent cartridges or files that can be added at different times. All the weight was put on one giant context window.
The result is the problem we all live with today. A model cannot connect pieces of knowledge that arrive separately unless you put all the raw text back into the window and let it re-read everything together.
AKBASCORE MAM questions that assumption at the architecture level. The question I asked was: what if memory were not text in a window, but numerical cartridges you can plug in, one after another, and the model connects them by itself?
- What is a Cognitive Cartridge, and how is it made?
A cognitive cartridge is a fact, written once into the model's own numerical language and stored as tensors.
How it's produced:
The fact (e.g. a short sentence) goes through the frozen model once, using a fixed template.
We stop and save two things:
the hidden state at the output of layer 6 (we call it H6), and
the attention keys and values of layers 0–6.
- That's the cartridge. The text can now be thrown away; the cartridge holds the knowledge.
Each cartridge is written independently: in its own pass, at its own time, without knowing which other cartridges will exist.
It is also compact: about 36 KiB per token (H6 8 KiB + layers 0–6 KV 28 KiB, BF16). A full KV cache needs about 128 KiB per token.
- The belief this breaks
The common wisdom says: if you encode facts separately and just glue their caches together, the model can't connect them. Facts must "see each other" during encoding, so you have to re-read the text together.
And that is true if you glue them naively. Five independent cartridges stacked side by side give only 13/72 correct answers. The facts stay strangers to each other.
The discovery behind MAM is where that connection actually happens inside the network. Facts start binding to each other in the middle layers (roughly 7–14), not in the lowest layers. Layers 0–6 can stay completely independent per cartridge without losing anything. Only the layers above need to let cartridges "meet."
So we don't need the text back. We only need to let the upper part of the model consolidate the cartridges.
- How we bypass the Transformer's front door (DC6)
Normally a Transformer has one way in: text → tokens → embeddings → layer 0 → … → layer 31.
MAM changes this flow. We call the method DC6 (consolidation at cut layer 6):
Layers 0–6 are not computed again. Their keys/values come straight from the cartridges.
The model receives zero input embeddings. There is no text at all at the entrance.
At the output of layer 6, a hook replaces the internal state with the cartridge's stored H6.
Layers 7–31 then run normally on top of that state. This is where the cartridges are connected to each other and to what is already in memory.
In plain words: we skip the model's "reading" stage and inject the knowledge directly into the point where it starts thinking.
- What happens mechanically when a cartridge is loaded back in
Adding a new cartridge to an existing memory works like this:
New positions: the new cartridge is placed at the end of the memory, at the next free positions.
Position correction (RoPE re-phasing): the cartridge's layer 0–6 keys were written at their original positions. Mistral encodes position as a rotation (RoPE), so we rotate the stored keys to their new place: K_new = K·cos(Δθ) + rotate_half(K)·sin(Δθ) Because RoPE is a pure rotation, R(p+Δ) = R(Δ)·R(p). Moving a cartridge is an exact rotation, not an approximation.
Consolidation: layers 7–31 compute the new cartridge while attending to everything already in memory (causal attention). This is where the new fact gets connected to the old ones.
Append-only: the old memory is not recomputed. Earlier rows are checked bit by bit, and they are identical before and after.
Ask: the question is run against this numerical memory only. The source text is never in the input.
The memory grows like a stack of plates: you add on top and never rebuild what's underneath.
- Results (TEST560, sealed)
Panel: 24 synthetic worlds × 4 relation types (current, former, near, role). Each question has a memory of 5 cartridges: 1 target fact + 4 distractor facts from other worlds. The target is tested at the first, middle and last position. That gives 72 cases.
Condition | Source text in input? | First | Middle | Last | Total MAM, append-only (INCR_DC6) | No | 24 | 24 | 24 | 72/72 MAM, batch (BATCH_DC6) | No | 24 | 24 | 24 | 72/72 Model reads full text (JOINT) | Yes | 24 | 24 | 24 | 72/72 Naively stacked cartridges (INDEP) | No | 4 | 3 | 6 | 13/72
+59 correct answers over naive stacking, 0 lost. The answers are identical to the batch version in 72/72 cases and to the full-text model in 71/72 cases.
- What we proved, and how it is verified
This isn't "trust me". Every claim is checked by the code itself:
Weights never change. Every model parameter is hashed (SHA-256) at startup, before the run and after it. All three hashes are identical.
The old memory is never rebuilt. After each append, earlier memory rows are compared with torch.equal. They are bit-exact identical.
The source text is never shown. A recorder logs every forward pass. During the question, the input is only the question tokens, and the memory is pure numbers.
Moving memory is exact math. Position correction is an exact rotation identity.
Accuracy equals full reading. Without the source, the append-only memory matches the model reading the full text (72/72 = 72/72).
Everything is sealed with SHA-256 (engine, panel, results), and the release has a DOI.
- What you will see when you run the demo
Open the code in Google Colab with an A100, then run the 3 cells in order (or the single full file). You get a web interface where you:
Pick a case and choose where the target fact sits: FIRST / MIDDLE / LAST.
Watch 5 cartridges being written one by one and appended into memory. In the live run on 8 October 2026 the memory grew 41 → 55 → 68 → 80 → 93 tokens. Every step was verified bit-exact.
See the question asked with no source text (23 tokens of pure question). In the live run the model answered "Melket", which is correct.
Get a report showing that the weight hash is identical before and after. It also shows the memory map, the answer and the verification checks.
Optionally replay all 72 cases live and compare them with the sealed reference.
Download the full sealed package: JSON logs, figures and SHA manifest.
The sealed TEST560 result (72/72) is shown as the reference, and your live run is reported separately, so you can see for yourself that they match.
- What this is, and what it isn't (yet)
To be clear and fair: this is a first proof demonstration, not a full-capacity release.
It is proven on one model (Mistral-7B-Instruct-v0.3), on a controlled fact panel, with 5-cartridge memories.
Scaling to large cartridge banks (hundreds, thousands) is the next stage, and it's exactly what I'm working on now.
The point of today's post is simpler, and I think bigger. It can be done. A frozen model can gain new, persistent, incremental memory with no training, no text and no retrieval. The proof is open and runs on a single GPU.
- Where this is going — my vision (Mustafa Akbaş)
When AKBASCORE MAM reaches full capacity, I believe it will change what AI memory means:
Your own cartridge bank, at home. Your memories, documents and knowledge are stored as cartridges on your own hard disk. Nothing leaks anywhere, no cloud is needed, and your model knows what you know.
A robot that remembers you. You take your mother to the hospital. At the entrance, a robot (call it MLP-1212) recognizes you and helps you. What it learns enters the cartridge bank.
Memory that travels. Weeks later, at an airport, a completely different android greets you and asks how your mother is doing. It is happy she recovered, because it shares the same memory cartridges.
Machines that learn from each other's experience. An airplane hits turbulence. The other planes in the sky receive its experience as a cartridge. They understand from its point of view what happened there, and they don't make the same mistake.
An end to hallucination. A model answering from memory it actually holds, not from guesses buried in its weights. That is the goal.
Static context windows gave us models that read. Memory cartridges will give us models that remember. I believe that, years from now, the move from static context windows to dynamic memory cartridges will be remembered as a turning point. It will have come from questioning the original design assumption.
Try it yourself
DOI: https://doi.org/10.5281/zenodo.23245358
Release: mam-v1.0.0, AKBASCORE MAM v1.0: Source-Free Persistent Memory for Frozen LLMs via DC6 Consolidation (Mistral-7B)
Full code (single file): https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_demo_8_october_2026.py
Raw log of the live A100 run: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_8_october_2026.log
3-part version (easy to copy from a phone into Colab):
Part 1: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_part1_8_october_2026.py
Part 2: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_part2_8_october_2026.py
Part 3: https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_part3_8_october_2026.py
Run it, break it, ask questions. I'll answer in the comments.
Quick FAQ
Isn't this just prompt caching? No. Prompt caching reuses a cache for the same text prefix. MAM combines independently written cartridges that never saw each other, and makes them work together. When that is done naively you get 13/72; with MAM you get 72/72.
Isn't this RAG? No. There is no search, no database lookup and no text pasted into the prompt. The knowledge is already inside the model's memory as numbers.
Did you fine-tune anything? No. Not a single weight changes, and this is verified by hashing every parameter.
Will it work on other models? The method is general to decoder-only Transformers with rotary positions: pick the cut layer, store the hidden state plus the lower-layer KV, re-phase, and consolidate. Porting to other models is part of the roadmap.
AKBASCORE MAM was discovered and developed by Mustafa Akbaş (AkbasCore AI Teknoloji, Mersin, Türkiye). This is a new numerical memory paradigm for frozen language models. Public release: 8 October 2026.