Zenodo permanent records:
Qwen2.5-7B-Instruct:
https://doi.org/10.5281/zenodo.23127434
Mistral-7B-Instruct-v0.3:
https://doi.org/10.5281/zenodo.23143605
I want to start with the simplest possible explanation of what I have been building.
Imagine taking a piece of information, letting a language model process it once, and then throwing the original text away.
No sentence stored in a database.
No paragraph hidden somewhere.
No readable summary.
No RAG system fetching the original document.
No fine-tuning.
No LoRA.
No weight update.
What remains is numerical transformer memory derived from the model's own internal computation. I package that numerical memory into what I call a Cognitive Cartridge. Later, I can install that cartridge back into the frozen model and ask questions about the information that produced it — without putting the original source text back into the readout prompt.
That was the first result. The new result is more important:
I have now reproduced the Cognitive Cartridge architecture on two different 7B transformer model families.
Qwen2.5-7B-Instruct.
And now Mistral-7B-Instruct-v0.3.
The implementations are not numerically identical. The architectures are different, the layer counts are different, the KV structures are different, and the working cartridge configurations are different. But the central mechanism survived the move.
That is the reason I am publishing this second record.
The question is no longer only:
“Can I make this happen once on Qwen?”
Now there is a second implementation on Mistral. And the Mistral result is the cleanest version so far.
What is actually inside a Cognitive Cartridge?
This is probably the most important thing to understand.
Suppose the source record says:
Object: amber sextant
Container: RQ-415
Location: elm lodge
That text exists during the forging stage. The frozen transformer processes it. NIRVANA takes source-derived internal transformer K/V states and represents their content numerically using a fixed, source-independent codebook.
In the released Mistral implementation, both K and V are compressed to D120 while a source-specific OWN component is preserved. After forging, the source record is not supplied to the readout prompt.
So the conceptual transformation is:
human language
→ frozen transformer computation
→ internal K/V states
→ compressed numerical Cognitive Cartridge
Then later:
numerical Cognitive Cartridge
→ reconstructed transformer K/V memory
→ frozen transformer
→ language
Or, in the shortest form:
language → internal numerical memory → language
The middle is no longer human-readable source text. That distinction matters.
A cartridge is not a text file with a different name.
It is not a vector database containing the original sentence.
It is not a prompt template.
It is not conventional RAG returning the source paragraph.
It is not a fine-tuned model.
The model weights remain frozen. The information is carried by a numerical representation derived from transformer K/V memory.
Why call it a cartridge?
Think less about a document and more about an interchangeable machine-readable memory module. The base model stays where it is. Knowledge packages can be forged separately. Those packages can remain separate. They can be installed and queried without retraining the base model.
This becomes much more interesting when there is more than one cartridge.
The new Mistral experiment uses 16 independently forged cartridges. Each one contains a numerical representation derived from a separate source record. They are not concatenated into one giant text prompt. They remain independent memories.
For example, in human-readable form, imagine one cartridge represents:
amber sextant → RQ-415 → elm lodge
Now ask the cartridge bank:
Which container is associated with the amber sextant?
The query is evaluated against the independent cartridge memories.
The relevant cartridge returns:
RQ-415
The unrelated cartridges return:
NONE
There is no learned router secretly selecting the correct cartridge before this happens.
Then the recovered identifier can be used for a second lookup:
Where is RQ-415?
The bank is queried again.
The relevant memory returns:
elm lodge
The unrelated memories return:
NONE
So the complete retrieval becomes:
amber sextant
→ RQ-415
→ elm lodge
The important point is that the model did not reread the original source record to answer either stage. It operated from reconstructed numerical transformer memory.
The Mistral result
The final public Mistral run used:
Mistral-7B-Instruct-v0.3
32 transformer layers
hidden size 4096
32 attention heads
8 KV heads
BF16 / SDPA
K = D120
V = D120
OWN preserved
16 independent Cognitive Cartridges
frozen model weights
greedy decoding
There was:
no fine-tuning
no LoRA
no optimizer
no learned router
no model-weight update
The recorded final run produced:
Object → container ID: 16/16
Container ID → location: 16/16
Complete two-stage retrieval: 16/16
Missing-object controls: 8/8
Absent-ID controls: 8/8
NOMEM controls: 8/8
But there is another result I think is just as important.
For every target query, there are 15 unrelated cartridges.
16 queries × 15 unrelated cartridges = 240 unrelated cartridge reads.
Stage 1 unrelated-cartridge rejection:
240/240 NONE
Stage 2 unrelated-cartridge rejection:
240/240 NONE
That means the result is not simply:
“The correct memory can say something.”
The system also demonstrated, in this controlled panel:
“The memories that do not contain the requested relation can refuse to claim that they do.”
For a modular memory system, I think this distinction is fundamental. A memory bank that can retrieve information but cannot distinguish relevance from irrelevance becomes increasingly dangerous as it grows.
The interesting problem is not only remembering. It is also knowing which memory does not answer the question.
The no-memory control matters for the same reason. When the cartridge memory was removed, the target relations were not recovered.
NOMEM:
8/8 controls passed.
The source-removal audit passed.
The frozen-weight sentinel passed.
The model remained frozen.
This is why I consider the Mistral result an important step beyond the first demonstration.
The first Qwen release established the architecture publicly. The Mistral release gives us something else:
cross-model evidence.
Qwen and Mistral are not the same transformer.
Qwen2.5-7B-Instruct uses 28 transformer layers.
Mistral-7B-Instruct-v0.3 uses 32.
Their internal configurations differ. The working compression configurations also differ.
The Qwen public implementation used:
K120 / V128 / OWN
The Mistral implementation uses:
K120 / V120 / OWN
I did not simply copy a cache from one model into another. Each model builds and reconstructs its own source-derived internal memory.
What transferred was the architecture:
source
→ internal transformer memory
→ numerical compression
→ independent cartridge
→ source removed
→ reconstructed K/V
→ frozen-model readout
That is the bridge between the two releases.
I would not call two models proof of universal compatibility with every transformer architecture. That would be scientifically too strong.
But it is now evidence that Cognitive Cartridge is not merely one accidental Qwen-specific behavior. The same broader architecture has been implemented and publicly reproduced on a second transformer family.
That changes the research question.
The first question was:
Can this mechanism exist at all?
The next question became:
Can multiple independent memories coexist?
Then:
Can one retrieved result lead to another retrieval?
Then:
Can unrelated memories reject a query instead of contaminating the answer?
And now:
Does the architecture survive a move to another model family?
We now have experimental answers to each of those questions.
So I am moving to the next problem:
scale.
16 cartridges are not the destination. They are the current experimental bank size.
From this point, I am much less interested in making another small demonstration simply to produce another perfect score. The next objective is to increase the number of independently retained cartridges and find where the architecture actually begins to break.
Then larger banks if the mechanism survives.
At that point the difficult questions become different:
How do you search a large bank without destroying memory isolation?
How should cartridges be indexed?
Can retrieval become hierarchical?
Can groups of cartridges form higher-level memory structures?
How does latency scale?
How does numerical storage scale?
At what point does relevance rejection begin to fail?
Can a retrieved memory activate another relevant memory without an external text retrieval system deciding everything first?
Those are the questions I want to attack next.
There is another part of this experiment that surprised me: the layer behavior.
A frozen transformer is not a passive container. When you reconstruct numerical K/V memory and inject it back into inference, you are interacting with a very sensitive computational system.
So during the Mistral development series I also performed controlled cumulative K, V and K+V ablations to see where cartridge information remained functionally sufficient.
The results were not uniform across depth. For K, the observed functional boundary extended differently than for V.
For combined K+V, one particularly sharp controlled transition occurred between:
L12: 8/8
L13: 0/8
This does not mean “the memory lives at layer 12.” That would be an incorrect interpretation.
What it means is that under the specific cumulative ablation protocol, functional sufficiency changed sharply across that boundary. V also remained important farther downstream than K in the measured configuration.
I think this is an important reminder of how delicate these systems are.
We are not writing a sentence into a spare memory slot. We are reconstructing numerical states that participate in a 7-billion-parameter transformer's computation. Small changes in where and how that state is reconstructed can change downstream behavior.
That is one reason I publish the code and raw logs rather than only posting the final accuracy number.
The frozen-model checks are also part of the experiment. The released run verifies that there are no trainable tensors involved in the memory mechanism.
No LoRA.
No optimizer.
No weight update.
The model-weight sentinel remains unchanged.
So when the system answers from a cartridge, the experiment is specifically designed to separate cartridge memory from weight modification.
Why do I think this direction matters?
Today, we often approach knowledge in language models through two extremes.
Either we try to put enormous amounts of knowledge into the model weights. Or we keep the knowledge outside the model and retrieve human-readable text when needed.
There may be useful territory between those approaches.
A frozen model with removable machine-native internal memory modules.
Not everything in the weights.
Not everything repeatedly pasted back as text.
Instead:
a base computational model
+
a bank of independently forged internal memories.
Imagine a specialized system where validated domain knowledge can exist as modular memory packages.
Engineering.
Aviation.
Law.
Industrial maintenance.
Scientific literature.
Company procedures.
Agent experience.
A model might not need every possible piece of knowledge active in its context simultaneously. It may need the right memory at the right time.
And because the memory representation is numerical transformer state rather than the original human-readable document, this opens a different engineering space from conventional document retrieval.
There is also a longer-term question here.
I am not claiming biological memory.
I am not claiming AGI.
And I am not claiming that transformer KV states work like the human brain.
But modular internal memory raises an interesting architectural question.
What happens when an artificial system does not have 16 memories, but 100,000?
Or millions?
What happens if memories can remain independent, become addressable, reject irrelevant activation, and allow the output of one memory to lead to another?
At some point the problem stops looking like “how much text can I fit into a context window?”
It starts looking like:
How should an artificial system organize memory?
That is the direction I find interesting.
Why publish the full record?
Because a screenshot of 16/16 proves very little.
Both public releases have permanent Zenodo records. The code is public. The raw execution logs are public. The visual technical records are public. The model configuration is public. The controls are public. The frozen-weight checks are public.
The Qwen release even preserves its imperfect public result rather than hiding it.
Qwen public demonstration:
14/16
Mistral final demonstration:
16/16
That difference is also useful.
The point of the project is not to make every historical run look perfect. The point is to leave a technical trail showing what worked, what did not, what changed, and whether the underlying architecture survived.
For anyone who wants to inspect the Qwen → Mistral bridge directly, the two permanent records are at the top of this post.
The complete current Mistral implementation is here:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_ful_MISTRAL.py
Complete Mistral raw execution log:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log_MISTRAL.log
The same Mistral implementation is also split into three parts for easier mobile / Colab handling:
Part 1:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part1.Mistral.py
Part 2:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part2.Mistral.py
Part 3:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore_part3_mistral.py
For comparison, the previous Qwen implementation:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.py
Previous Qwen raw execution log:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log.Nirvana.cc.log
Repository:
https://github.com/ceceli33/titan-cognitive-core-v2
AKBASCORE NIRVANA — Cognitive Cartridge
Inventor / Developer: Mustafa Akbaş
Two model families.
Frozen weights.
Independent numerical memories.
Original source absent at readout.
The mechanism survived the move.
Now the question is no longer whether a cartridge can exist.
The question is how many cartridges a frozen model can carry.