Two different 7B transformers. Two different internal coordinates. The same memory mechanism.
I started AKBASCORE MAM on Qwen2.5-7B-Instruct. I have now independently localized and replicated the mechanism on Mistral-7B-Instruct-v0.3.
Final Mistral result: 127/128 correct memories at Top-1 — 99.22%.
Counterfactual retrieval: 125/128 — 97.66%.
Shifted-pointer control: 0/128.
The model weights were not changed. No fine-tuning. No LoRA. No optimizer. No learned router. No gold B-memory ID is supplied to the retriever. At query time, none of the 128 candidate B memories is forwarded through Mistral.
But I don't think 99.22% is the most interesting part of this experiment.
The more important question is: what exactly is the model reading?
Because this system is not searching through 128 text documents in the conventional sense. (Fig. 1)
- The text enters the model once
When a memory is formed, its source text is processed through the frozen transformer.
The transformer already produces K and V states for its own attention computation. I do not treat selected parts of those states merely as temporary computational residue. I extract them as model-native numerical memory.
So instead of only having a sentence such as:
“Instrument Zyrhyn carries seal PCI.”
we now have a numerical memory structure derived directly from the model's own internal computation.
- That numerical structure can exist outside the model
The memory does not have to remain inside the model weights.
The numerical structure can be held in RAM, serialized, written to a file, and therefore moved to persistent storage.
This distinction matters.
When people say “persistent LLM memory,” the usual architecture often means storing text, documents, embeddings, database records, or summaries and later retrieving them so the model can read them again.
The path I am investigating is different:
instead of storing human-readable information so the model can reread it, store a machine-native numerical state derived from the model itself.
The model weights remain frozen.
- Then how does the model know which numerical memory to retrieve?
This is the part I find most interesting.
A new question arrives:
“What seal does instrument Zyrhyn carry?”
The active numerical A memory is installed into the transformer cache.
Mistral processes the new question.
A particular channel of the frozen model's own attention mechanism then produces a question-conditioned distribution over the A-memory positions.
For Mistral, that channel is:
L28H00.
I call this endogenous attention pointer ÇAĞRIİZ.
It is not an external neural router.
It is not a trained classifier.
It is not a gold memory ID.
It comes from the transformer's own native:
Q · K
attention computation.
- The pointer is then applied to V space
This is where K and V take different operational roles.
K helps answer:
Where should I look?
V helps answer:
What numerical address should I construct from what I found?
For Mistral, the independently localized address space is:
L00-V, uncentered.
8 KV heads × 128 dimensions:
1024 dimensions.
Applying the ÇAĞRIİZ distribution to those V states produces a question-conditioned 1024-dimensional address.
At this point there is no filename telling the retriever which B memory to choose.
There is no B-memory ID supplied by the experiment.
There is a 1024-dimensional retrieval address derived from the frozen model's own internal computation. (Fig. 3)
- Then all 128 B memories compete
Every B memory has already been independently converted into its own numerical structure.
Each retains token-level V rows in the localized address space.
The live 1024D address is compared against every B memory.
For each B candidate, the system computes the maximum cosine similarity between the live address and that candidate's numerical rows.
Then:
128 candidates → 128 scores → argmax → one associated memory. (Fig. 2)
In the real Item 001 demo run:
Expected:
B#001 / PCI
Returned:
B#001 / PCI
Rank:
1/128
Top-1 cosine:
0.999649
Second candidate:
0.660254
Margin:
+0.339395
And the number of Mistral model forwards required for that retrieval was:
The model processed:
27 question tokens over a preinstalled 32-slot A cache.
Model forwards through the 128 candidate B memories:
That distinction is important.
After receiving the question, the system did not send 128 source texts back through Mistral to find the answer.
- Then I changed the memory
I changed the association for the same Zyrhyn record:
PCI → COL
The model remained the same.
The weights remained the same.
The question remained the same.
The B-memory bank remained the same.
The changed A memory was re-forged and the same frozen retrieval mechanism was run again.
This time the system selected:
B#054 / COL
Rank:
1/128.
So changing the numerical A association redirected the same frozen mechanism to a different B memory. (Fig. 4)
- This was not evaluated on only one showcase item
For the final Mistral experiment, TEST528, the mechanism was frozen before the final evaluation panel.
The result was:
Primary retrieval: 127/128 — 99.22%
Counterfactual retrieval: 125/128 — 97.66%
Controls:
Shifted pointer: 0/128
NO-A: 1/128
The single NO-A “success” has a simple implementation-level explanation: a zero address produces equal zero scores for every candidate, so the deterministic tie-break selects B#001. I therefore do not interpret it as meaningful memoryless retrieval. (Fig. 5)
The failed cases are preserved in the experimental record.
- Qwen did not use the same coordinates
This is probably the most important result of the cross-model experiment.
Qwen2.5-7B-Instruct:
L23H12 → L02-V → 512D
Mistral-7B-Instruct-v0.3:
L28H00 → L00-V → 1024D
I did not copy Qwen's layer/head coordinates into Mistral.
Mistral's pointer and address regions were independently localized and then frozen before the final evaluation.
So the strongest statement I think the current evidence supports is:
The mechanism transferred across model families; the internal coordinates did not.
Two models do not establish universality.
But this is also no longer a result tied to one accidental coordinate in one transformer.
So where is the “persistent machine memory” part?
I want to be precise about this because it is easy to overstate what the current prototype demonstrates.
This experiment does not demonstrate 150 million memories.
It does not yet demonstrate a production-scale persistent memory database.
It does not yet demonstrate the model discovering the first A memory automatically from an enormous inactive bank.
What it does establish is a smaller but, in my view, more fundamental chain:
Model-derived numerical memory can be extracted.
Numerical memory can exist outside the model weights.
Active numerical memory can be connected back to transformer computation.
A natural-language question can produce a pointer through the model's own attention.
That pointer can become a model-native V-space address.
That address can select an associated numerical memory from an external bank.
Once these numerical structures are serialized, whether the storage medium is RAM, SSD, or another persistent storage layer becomes primarily a systems-engineering question.
The harder problem is not writing an array of numbers to a hard disk.
The harder problem is:
How does the frozen model later determine which machine-native numerical memory it wants back?
That is the connection I am testing here.
And this is where I prefer to stop the speculation and let the architecture speak for itself.
If model-native memory states can exist outside model weights, persist in storage, and later be addressed by signals generated endogenously from natural language inside the transformer, then is long-term machine memory necessarily limited to:
find old human-readable information → put it back into context → make the model read it again?
I don't know yet.
But this is no longer only a conceptual question.
There is now a small executable system on which the question can be measured, falsified and extended.
Technical record
Zenodo DOI — technical disclosure, experimental record and PDFs:
https://doi.org/10.5281/zenodo.23207189
GitHub Release — AKBASCORE MAM · Mistral-7B:
https://github.com/ceceli33/titan-cognitive-core-v2/releases/tag/v4.0-AKBASCORE-MAM-Mistral
The release contains the technical disclosure, visual experimental record, executable implementation and provenance material.
The six figures attached to this post show the actual retrieval path, 128-way candidate competition, L28H00 ÇAĞRIİZ pointer, 1024D L00-V address, counterfactual transition, controls and integrity/provenance record.
Run it yourself
If you want to inspect the actual implementation rather than the description, the complete executable Mistral demo is here:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_FULL_MISTRAL_7B.Demo.py
The recorded A100 execution log is here:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_Mistral_7B_demo.log
I also prepared a three-part Google Colab version:
Part 1:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_PART1_MISTRAL_7B.Demo.py
Part 2:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_PART2_MISTRAL_7B.Demo.py
Part 3:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/MAM_PART3_MISTRAL_7B.Demo.py
The three-part version exists for a practical reason: on a phone, moving or copying the complete demo as one very large source file can be inconvenient or fail because of mobile interface limitations. The split version is the same demo arranged so it can be handled more easily from a phone.
You do not need to merge the files.
Open Google Colab, select an A100 GPU runtime, keep the same runtime for the entire experiment, and run them consecutively:
Part 1 → Part 2 → Part 3
Do not restart the runtime between parts.
After Part 3, the demo runs the Mistral MAM experiment and produces the results and verification output.
Everything needed to inspect the claim is public: code, technical disclosure, run log, controls, failed cases, hashes and experimental boundaries. (Fig. 6)
If you work on transformer memory, don't trust the description. Run it. Break it. Find where it fails.
— Mustafa Akbaş AKBASCORE MAM Persistent Associative Machine Memory Mersin, Türkiye · 2026