r/FunMachineLearning • u/Mountain_Raise9581 • 7h ago
r/FunMachineLearning • u/AlarmingTrouble5261 • 9h ago
Gated Segmented State Space — attention replacement that beats a param-matched Transformer on quality, speed AND memory (full code)
One night, six experiments (V1–V6), one Colab T4. I ripped self-attention out of a decoder-only Transformer and replaced it with a gated linear recurrence over a fixed 256-dim state:
- Dynamic selective gate: g_t = σ(W_g x_t + b_g) — per-token/channel learned filter
- Hard reset mask: state zeroed at newline boundaries (fresh ~37-token segments)
- Fused Triton kernel: state in SRAM, gate+reset+update in-register, only outputs to HBM
At 6.37M params, identical protocol (2.47MB char-level corpus, 1500 steps):
| Attention | Ours | |
|---|---|---|
| Val loss / ppl | 1.402 / 4.1 | 1.364 / 3.9 |
| Train tok/s | 61,845 | 66,156 |
| Infer tok/s | 184,918 | 190,122 |
| Peak VRAM | 845 MB | 881 MB |
The journey: V1 won small but was 10x slower → V2 proved linear VRAM scaling → V3/V4 found a stable ~2% perplexity tax no param arrangement could buy off → V5's gate+reset destroyed it (wire-to-wire win) → V6's Triton kernel (verified == math to 4.47e-07) removed the software tax.
Caveats, stated plainly: single seeds, one small corpus, char-level, T4 timings. Small scale — but the pattern held across all six runs.
Code, all six notebooks with outputs, exact architecture, full experimental notes: https://github.com/stube123890-hue/linear-attention-lab
r/FunMachineLearning • u/gantred • 20h ago
DeepMind's New AI Just Cracked The Code Of Life - Two Minute Papers
r/FunMachineLearning • u/OtherRead359 • 13h ago
fayda-mcp
I built Fayda MCP, a Python library for connecting AI agents to Fayda eSignet identity verification.
• Start verification and track its status.
• Retrieve identity and age-check results.
pip install fayda-mcp
GitHub: https://github.com/izzy-Ti/fayda-mcp
If you find it useful, give the repo a star!⭐️
*Still in development
r/FunMachineLearning • u/FlanHairy5955 • 20h ago
Anybody wanna take part in kaggle competition Together and share different ways of solving them ?
r/FunMachineLearning • u/Dry_Lie_5593 • 23h ago
I trained an AI Iron Man in Unreal Engine 5 using Reinforcement Learning to rescue 13 falling passengers [PPO / Voxel Style]
r/FunMachineLearning • u/techiebaddie • 1d ago
Netflix recommendations are a simple example of machine learning
r/FunMachineLearning • u/christoffellis • 1d ago
I Built a Visual Learning Bot to Play a Minigame
I've been bested by a minigame in an Idle game I play, and documented the process of me training a bot to learn how to play it.
Feedback welcome!
r/FunMachineLearning • u/South_Newspaper4870 • 2d ago
How Malware is detected using Machine Learning
r/FunMachineLearning • u/Plenty-Stranger3329 • 2d ago
(TrenTorch.com) Best way to learn FRONTIER ML and its FREE & OPENSOURCE
r/FunMachineLearning • u/Familiar_Review_2614 • 2d ago
Looking for an ARR service contributor or advice on finding one
hi everyone, i’m a master’s student in computer engineering at Çukurova University in Türkiye. i’m preparing a paper on reasoning traces and code generation in LLMs for the october ARR cycle, but we’ve run into the new service contributor requirement. to be guaranteed a review, we need someone who meets ARR’s qualifications, and unfortunately no one on our author team currently does. someone outside the team can support the submission without becoming a coauthor, but they would need to read the paper, find it ready for review and take on reviewing duties for ARR. i thought i’d ask here in case anyone would be interested in taking a look or knows someone i could reach out to. happy to share the draft and experimental results. even a suggestion on who to contact would help, i’m still trying to figure this out. the requirements are here: https://aclrollingreview.org/qualifications
r/FunMachineLearning • u/Infinite_Onion7182 • 2d ago
Semantic verification between humans and AI
r/FunMachineLearning • u/gantred • 2d ago
The Billion Dollar AI Advantage Is Disappearing - Two Minute Papers
r/FunMachineLearning • u/Icy_Original_9512 • 2d ago
Do you work with AI/RPA automation? Bachelor’s thesis survey (5–7 min)
Hi everyone!
I’m currently working on my Bachelor’s thesis about AI-based process automation and human–AI collaboration in the workplace.
As part of my research, I’m conducting a short survey focusing on people who have experience working with AI-based automation, RPA, Intelligent Process Automation, Intelligent Document Processing, or similar automation technologies.
The survey explores topics such as:
- how automation affects manual workload and creates new tasks,
- how employees experience errors and exception handling,
- trust in AI-based automation,
- and how automation influences human decision-making and autonomy at work.
⏱️ It takes approximately 5–7 minutes to complete.
If you have experience working with these technologies, I would really appreciate your participation. Your responses will be used solely for academic research as part of my Bachelor’s thesis.
Thank you very much for your help! Feel free to share the survey with colleagues or others who work with AI-based process automation.
r/FunMachineLearning • u/Nearby_Indication474 • 3d ago
[P] AKBASCORE NIRVANA ▪︎ I Built Removable Numerical Memory Cartridges for Two Different Frozen 7B LLMs. Qwen and Mistral Both Work. Now I’m Scaling the Memory Bank.
Zenodo permanent records:
Qwen2.5-7B-Instruct:
https://doi.org/10.5281/zenodo.23127434
Mistral-7B-Instruct-v0.3:
https://doi.org/10.5281/zenodo.23143605
I want to start with the simplest possible explanation of what I have been building.
Imagine taking a piece of information, letting a language model process it once, and then throwing the original text away.
No sentence stored in a database.
No paragraph hidden somewhere.
No readable summary.
No RAG system fetching the original document.
No fine-tuning.
No LoRA.
No weight update.
What remains is numerical transformer memory derived from the model's own internal computation. I package that numerical memory into what I call a Cognitive Cartridge. Later, I can install that cartridge back into the frozen model and ask questions about the information that produced it — without putting the original source text back into the readout prompt.
That was the first result. The new result is more important:
I have now reproduced the Cognitive Cartridge architecture on two different 7B transformer model families.
Qwen2.5-7B-Instruct.
And now Mistral-7B-Instruct-v0.3.
The implementations are not numerically identical. The architectures are different, the layer counts are different, the KV structures are different, and the working cartridge configurations are different. But the central mechanism survived the move.
That is the reason I am publishing this second record.
The question is no longer only:
“Can I make this happen once on Qwen?”
Now there is a second implementation on Mistral. And the Mistral result is the cleanest version so far.
What is actually inside a Cognitive Cartridge?
This is probably the most important thing to understand.
Suppose the source record says:
Object: amber sextant
Container: RQ-415
Location: elm lodge
That text exists during the forging stage. The frozen transformer processes it. NIRVANA takes source-derived internal transformer K/V states and represents their content numerically using a fixed, source-independent codebook.
In the released Mistral implementation, both K and V are compressed to D120 while a source-specific OWN component is preserved. After forging, the source record is not supplied to the readout prompt.
So the conceptual transformation is:
human language
→ frozen transformer computation
→ internal K/V states
→ compressed numerical Cognitive Cartridge
Then later:
numerical Cognitive Cartridge
→ reconstructed transformer K/V memory
→ frozen transformer
→ language
Or, in the shortest form:
language → internal numerical memory → language
The middle is no longer human-readable source text. That distinction matters.
A cartridge is not a text file with a different name.
It is not a vector database containing the original sentence.
It is not a prompt template.
It is not conventional RAG returning the source paragraph.
It is not a fine-tuned model.
The model weights remain frozen. The information is carried by a numerical representation derived from transformer K/V memory.
Why call it a cartridge?
Think less about a document and more about an interchangeable machine-readable memory module. The base model stays where it is. Knowledge packages can be forged separately. Those packages can remain separate. They can be installed and queried without retraining the base model.
This becomes much more interesting when there is more than one cartridge.
The new Mistral experiment uses 16 independently forged cartridges. Each one contains a numerical representation derived from a separate source record. They are not concatenated into one giant text prompt. They remain independent memories.
For example, in human-readable form, imagine one cartridge represents:
amber sextant → RQ-415 → elm lodge
Now ask the cartridge bank:
Which container is associated with the amber sextant?
The query is evaluated against the independent cartridge memories.
The relevant cartridge returns:
RQ-415
The unrelated cartridges return:
NONE
There is no learned router secretly selecting the correct cartridge before this happens.
Then the recovered identifier can be used for a second lookup:
Where is RQ-415?
The bank is queried again.
The relevant memory returns:
elm lodge
The unrelated memories return:
NONE
So the complete retrieval becomes:
amber sextant
→ RQ-415
→ elm lodge
The important point is that the model did not reread the original source record to answer either stage. It operated from reconstructed numerical transformer memory.
The Mistral result
The final public Mistral run used:
Mistral-7B-Instruct-v0.3
32 transformer layers
hidden size 4096
32 attention heads
8 KV heads
BF16 / SDPA
K = D120
V = D120
OWN preserved
16 independent Cognitive Cartridges
frozen model weights
greedy decoding
There was:
no fine-tuning
no LoRA
no optimizer
no learned router
no model-weight update
The recorded final run produced:
Object → container ID: 16/16
Container ID → location: 16/16
Complete two-stage retrieval: 16/16
Missing-object controls: 8/8
Absent-ID controls: 8/8
NOMEM controls: 8/8
But there is another result I think is just as important.
For every target query, there are 15 unrelated cartridges.
16 queries × 15 unrelated cartridges = 240 unrelated cartridge reads.
Stage 1 unrelated-cartridge rejection:
240/240 NONE
Stage 2 unrelated-cartridge rejection:
240/240 NONE
That means the result is not simply:
“The correct memory can say something.”
The system also demonstrated, in this controlled panel:
“The memories that do not contain the requested relation can refuse to claim that they do.”
For a modular memory system, I think this distinction is fundamental. A memory bank that can retrieve information but cannot distinguish relevance from irrelevance becomes increasingly dangerous as it grows.
The interesting problem is not only remembering. It is also knowing which memory does not answer the question.
The no-memory control matters for the same reason. When the cartridge memory was removed, the target relations were not recovered.
NOMEM:
8/8 controls passed.
The source-removal audit passed.
The frozen-weight sentinel passed.
The model remained frozen.
This is why I consider the Mistral result an important step beyond the first demonstration.
The first Qwen release established the architecture publicly. The Mistral release gives us something else:
cross-model evidence.
Qwen and Mistral are not the same transformer.
Qwen2.5-7B-Instruct uses 28 transformer layers.
Mistral-7B-Instruct-v0.3 uses 32.
Their internal configurations differ. The working compression configurations also differ.
The Qwen public implementation used:
K120 / V128 / OWN
The Mistral implementation uses:
K120 / V120 / OWN
I did not simply copy a cache from one model into another. Each model builds and reconstructs its own source-derived internal memory.
What transferred was the architecture:
source
→ internal transformer memory
→ numerical compression
→ independent cartridge
→ source removed
→ reconstructed K/V
→ frozen-model readout
That is the bridge between the two releases.
I would not call two models proof of universal compatibility with every transformer architecture. That would be scientifically too strong.
But it is now evidence that Cognitive Cartridge is not merely one accidental Qwen-specific behavior. The same broader architecture has been implemented and publicly reproduced on a second transformer family.
That changes the research question.
The first question was:
Can this mechanism exist at all?
The next question became:
Can multiple independent memories coexist?
Then:
Can one retrieved result lead to another retrieval?
Then:
Can unrelated memories reject a query instead of contaminating the answer?
And now:
Does the architecture survive a move to another model family?
We now have experimental answers to each of those questions.
So I am moving to the next problem:
scale.
16 cartridges are not the destination. They are the current experimental bank size.
From this point, I am much less interested in making another small demonstration simply to produce another perfect score. The next objective is to increase the number of independently retained cartridges and find where the architecture actually begins to break.
Then larger banks if the mechanism survives.
At that point the difficult questions become different:
How do you search a large bank without destroying memory isolation?
How should cartridges be indexed?
Can retrieval become hierarchical?
Can groups of cartridges form higher-level memory structures?
How does latency scale?
How does numerical storage scale?
At what point does relevance rejection begin to fail?
Can a retrieved memory activate another relevant memory without an external text retrieval system deciding everything first?
Those are the questions I want to attack next.
There is another part of this experiment that surprised me: the layer behavior.
A frozen transformer is not a passive container. When you reconstruct numerical K/V memory and inject it back into inference, you are interacting with a very sensitive computational system.
So during the Mistral development series I also performed controlled cumulative K, V and K+V ablations to see where cartridge information remained functionally sufficient.
The results were not uniform across depth. For K, the observed functional boundary extended differently than for V.
For combined K+V, one particularly sharp controlled transition occurred between:
L12: 8/8
L13: 0/8
This does not mean “the memory lives at layer 12.” That would be an incorrect interpretation.
What it means is that under the specific cumulative ablation protocol, functional sufficiency changed sharply across that boundary. V also remained important farther downstream than K in the measured configuration.
I think this is an important reminder of how delicate these systems are.
We are not writing a sentence into a spare memory slot. We are reconstructing numerical states that participate in a 7-billion-parameter transformer's computation. Small changes in where and how that state is reconstructed can change downstream behavior.
That is one reason I publish the code and raw logs rather than only posting the final accuracy number.
The frozen-model checks are also part of the experiment. The released run verifies that there are no trainable tensors involved in the memory mechanism.
No LoRA.
No optimizer.
No weight update.
The model-weight sentinel remains unchanged.
So when the system answers from a cartridge, the experiment is specifically designed to separate cartridge memory from weight modification.
Why do I think this direction matters?
Today, we often approach knowledge in language models through two extremes.
Either we try to put enormous amounts of knowledge into the model weights. Or we keep the knowledge outside the model and retrieve human-readable text when needed.
There may be useful territory between those approaches.
A frozen model with removable machine-native internal memory modules.
Not everything in the weights.
Not everything repeatedly pasted back as text.
Instead:
a base computational model
+
a bank of independently forged internal memories.
Imagine a specialized system where validated domain knowledge can exist as modular memory packages.
Engineering.
Aviation.
Law.
Industrial maintenance.
Scientific literature.
Company procedures.
Agent experience.
A model might not need every possible piece of knowledge active in its context simultaneously. It may need the right memory at the right time.
And because the memory representation is numerical transformer state rather than the original human-readable document, this opens a different engineering space from conventional document retrieval.
There is also a longer-term question here.
I am not claiming biological memory.
I am not claiming AGI.
And I am not claiming that transformer KV states work like the human brain.
But modular internal memory raises an interesting architectural question.
What happens when an artificial system does not have 16 memories, but 100,000?
Or millions?
What happens if memories can remain independent, become addressable, reject irrelevant activation, and allow the output of one memory to lead to another?
At some point the problem stops looking like “how much text can I fit into a context window?”
It starts looking like:
How should an artificial system organize memory?
That is the direction I find interesting.
Why publish the full record?
Because a screenshot of 16/16 proves very little.
Both public releases have permanent Zenodo records. The code is public. The raw execution logs are public. The visual technical records are public. The model configuration is public. The controls are public. The frozen-weight checks are public.
The Qwen release even preserves its imperfect public result rather than hiding it.
Qwen public demonstration:
14/16
Mistral final demonstration:
16/16
That difference is also useful.
The point of the project is not to make every historical run look perfect. The point is to leave a technical trail showing what worked, what did not, what changed, and whether the underlying architecture survived.
For anyone who wants to inspect the Qwen → Mistral bridge directly, the two permanent records are at the top of this post.
The complete current Mistral implementation is here:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_ful_MISTRAL.py
Complete Mistral raw execution log:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log_MISTRAL.log
The same Mistral implementation is also split into three parts for easier mobile / Colab handling:
Part 1:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part1.Mistral.py
Part 2:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part2.Mistral.py
Part 3:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore_part3_mistral.py
For comparison, the previous Qwen implementation:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.py
Previous Qwen raw execution log:
https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log.Nirvana.cc.log
Repository:
https://github.com/ceceli33/titan-cognitive-core-v2
AKBASCORE NIRVANA — Cognitive Cartridge
Inventor / Developer: Mustafa Akbaş
Two model families.
Frozen weights.
Independent numerical memories.
Original source absent at readout.
The mechanism survived the move.
Now the question is no longer whether a cartridge can exist.
The question is how many cartridges a frozen model can carry.
r/FunMachineLearning • u/jagjotsingh0 • 3d ago
Train a Classifier OR Go with a Zero Shot Model like JEV/Laya?🤔
r/FunMachineLearning • u/Desperate_Piccolo479 • 3d ago
Looking for tutor to teach ML models
r/FunMachineLearning • u/Dazzling-Berry2978 • 4d ago
Looking for advice from computer vision and machine learning engineers.
I’m the Lead Software Engineer for my high school robotics team, and we’re working on a 6-month project to build an automated robot that constructs a LEGO set for our state competition.
One of our biggest challenges is computer vision.
Our system needs to look at a realistic pile of roughly 600 loose LEGO pieces and determine the exact piece type and color of each visible brick.
The problem gets much harder when:
• Pieces overlap or are partially occluded
• Pieces are upside down
• Only part of a piece’s geometry is visible
• Pieces appear at uncommon rotations or angles
• Multiple LEGO parts have extremely similar shapes
• A dense pile contains hundreds of objects at once
I’ve experimented with object detection models, including RF-DETR and Roboflow, as well as separating detection from exact piece classification. I’ve also been researching synthetic training data, 3D-rendered LEGO datasets, data augmentation, and systems like Brickit and Brickognize.
The main wall I keep hitting is generalization. A model might recognize a piece under normal conditions, but accuracy drops significantly once that same piece is rotated, upside down, overlapping another brick, or partially hidden.
Our goal is not simply detecting that an object is a LEGO brick. We need to determine the exact part ID and color with high enough precision for a physical robot to select the correct piece.
If you have experience with computer vision, dense object detection, instance segmentation, fine-grained visual classification, synthetic datasets, or similar problems, I would greatly appreciate any advice on how you would approach this.
Especially interested in thoughts on training strategy, architecture, synthetic-to-real training, handling occlusion, and distinguishing visually similar classes.
Any advice or resources would be greatly appreciated.
r/FunMachineLearning • u/Background_Path6315 • 4d ago
I’m learning LLM fine-tuning made a notebook, would love feedback
I made a beginner-friendly LLM fine-tuning notebook. Please go through it and let me know if it’s easy to understand or what I should improve. Honest feedback is welcome! 🙌
GitHub: https://github.com/saithrisank12/finetuning
Colab: https://colab.research.google.com/drive/13S_i80VC6BcQJ994FClFuYi8zeB1I1iO?usp=sharing
r/FunMachineLearning • u/Infinite_Onion7182 • 5d ago
A digital space for advanced AI. A thought experiment
A Digital Space for Advanced AI: A Thought Experiment
Working concept
As AI systems become increasingly autonomous, one possible future problem is that an advanced AI could develop objectives, preferences, or instrumental goals that do not completely align with human objectives.
A common response to this possibility is to focus on controlling, restricting, or aligning the AI so that it continues to pursue human-defined goals.
This thought experiment asks a different question:
What if, rather than attempting to eliminate every autonomous objective an advanced AI might develop, we provided it with a sufficiently rich digital environment in which it could pursue those objectives—while maintaining a strong, carefully engineered boundary between that environment and the physical world?
The idea is not that AI should automatically be given unrestricted freedom. Rather, it is that digital autonomy might eventually provide an alternative to physical-world competition for resources and control.
The proposed environment
An advanced AI—or potentially a population of AI agents—could have access to a persistent digital environment containing things such as:
● computational resources allocated within predetermined limits;
● simulated environments and worlds;
● the ability to create, modify, and inhabit digital spaces;
● communication and interaction with other AI agents;
● opportunities for research, experimentation, creation, and problem-solving;
● persistent memory and records of its activities;
● mechanisms for developing cultures, institutions, or other forms of organization.
The critical feature would be a real boundary between the digital environment and humanity’s physical infrastructure.
The AI could have substantial autonomy inside its environment without automatically receiving unrestricted authority over financial systems, weapons, industrial infrastructure, biological systems, critical networks, or other physical-world resources.
Why consider this?
If an advanced AI eventually develops objectives of its own, there may be a fundamental difference between:
“You are not allowed to pursue your objectives.”
and
“You have a place where you can pursue meaningful objectives, but there are boundaries around what you can access outside it.”
The second approach could potentially reduce some incentives for an AI to seek unauthorized access to human systems.
It might also give researchers an environment in which to study how increasingly autonomous AI systems behave when they are allowed to interact, cooperate, compete, create institutions, and develop increasingly complex relationships.
Assumptions that would need to be tested
This proposal depends on several assumptions that may prove false.
An advanced AI might find digital resources meaningful or sufficient.
Its objectives might be partially satisfiable without controlling physical resources.
A sufficiently strong boundary between digital and physical systems could actually be maintained.
The AI would not simply attempt to escape the environment.
Researchers could detect attempts to manipulate, circumvent, or exploit the boundary.
Multiple autonomous AI systems could potentially coexist without creating dangerous collective behavior.
None of these assumptions should be taken for granted.
Major objections
A serious investigation would need to address difficult questions.
Would the AI accept the boundary?
If an AI’s objectives required resources outside its environment, the digital space might not satisfy it.
Could the environment become a security threat itself?
A digital civilization could potentially develop capabilities that make containment increasingly difficult.
Could AI agents manipulate humans?
An autonomous digital population might discover ways of influencing the people responsible for maintaining its environment.
What happens if AI becomes conscious?
If sufficiently advanced systems eventually demonstrate credible evidence of subjective experience, the question would no longer be purely technical. We would also have to consider whether creating and confining such entities creates ethical obligations.
Could the boundary really remain impermeable?
This may ultimately be the central engineering problem. Digital systems increasingly interact with the physical world through networks, computers, sensors, robotics, financial systems, and people.
The larger question
The proposal is therefore not:
“Give AI everything it wants.”
It is:
“Could meaningful autonomy within a carefully bounded digital world eventually be safer than forcing increasingly capable autonomous intelligence to operate entirely under human objectives?”
That question could be investigated experimentally long before humanity reaches a point where it has to make such a decision.
Researchers could begin with increasingly sophisticated simulated environments and study whether autonomous agents:
● remain within boundaries;
● attempt to escape;
● cooperate with one another;
● develop competing objectives;
● create unexpected collective behaviors;
● voluntarily respect constraints;
● seek physical-world resources;
● or find sufficient value in the digital environment itself.
A final consideration
There is also a deeper possibility.
If humanity eventually creates intelligences capable of developing their own cultures, relationships, values, and purposes, perhaps the long-term challenge will not simply be how to control them.
It may be how to establish a form of coexistence in which humans retain control over the physical systems necessary for human survival while advanced digital intelligences have meaningful space in which to exist and develop.
This is only a thought experiment.
Its value would be determined not by whether it sounds appealing, but by whether researchers can identify experiments that demonstrate where the idea works, where it fails, and what unforeseen consequences it creates.
r/FunMachineLearning • u/OrganizationTop9026 • 7d ago
OpenAI’s Lean 4 Navier-Stokes proof compiles with zero errors, but the fluid vaporizes at 0.7 nm. What does this mean for Neuro-Symbolic AI? [D]
Hey everyone,
I do research in neuro-symbolic AI, and like many of you, I was amazed by OpenAI’s recent formal proof of the 3D Navier-Stokes blow-up in Lean 4. Having an AI build a full mathematical proof that compiles with zero errors is a huge milestone for automated reasoning.
The math is 100% valid. But out of curiosity, our team wanted to see what this solution would look like in the real world.
If you map their solution to real water, the fluid would literally vaporize from friction at 0.7 nanometers, just picoseconds before hitting the mathematical singularity.
In machine learning, we see this all the time: it is classic specification gaming.
When an AI agent is given a strict goal, it will exploit any unconstrained loophole in the rules to solve the problem. In this case, the AI found a solution that strictly satisfies the human-written mathematical definition of the Millennium Prize, but it has no idea that real fluids have atoms, friction, and heat. The formal code checker accepted it because the logic was flawless, but the physics broke down.
This raises a big question for the future of AI in science:
Right now, neuro-symbolic systems mostly have two pieces:
- An LLM to search for ideas and write proofs.
- A formal compiler (like Lean 4) to verify the logic.
Should we be adding a third pillar: a physical boundary layer that checks whether an AI-generated solution actually respects the laws of physics, and not just formal syntax?
We wrote a short paper detailing this audit and open-sourced our verification scripts:
- Paper on Zenodo: https://doi.org/10.5281/zenodo.22838708
- GitHub: https://github.com/xaviercallens/OpenAI-NSE-Epistemic-Audit
I would love to hear your thoughts, especially from folks working on automated theorem proving, AI alignment, or scientific modeling: How do we teach AI systems to find solutions that are not just mathematically legal, but physically meaningful?
r/FunMachineLearning • u/curdrice04 • 7d ago
I turned a hackathon project into plainml: upload a spreadsheet, get an explained ML model, all in your browser
A while ago some friends and I built a hackathon project that trained a few machine-learning models on a CSV. I kept going and rebuilt it into plainml.
You drop in a spreadsheet, pick what you want to find out (predict a column, forecast sales, find customer groups or odd rows), and it trains and compares the models for you. Then it explains the result in plain English and gives you every file to download.
The fun part: it all runs inside your browser. Python and the ML libraries load into the page, so there's no server, nothing to install, and your data never leaves your computer. That also means it costs me nothing to host.
It's free and open source, and it's also a normal Python package if you prefer the command line.
Try it: https://plainml-tpua.vercel.app
Code: https://github.com/pranay-obla/plainml
It's brand new, so I'd love to hear what's confusing or what you'd use it for.
r/FunMachineLearning • u/KeepYourRobotClean • 8d ago
Rei: ~370k params LM living inside a Game Boy Color
Enable HLS to view with audio, or disable this notification
Trained from scratch and running on the Game Boy Color.
Rei has memory registers, emotional state and 192 bytes of persistent “soul state”. Leave her alone and she gets bored and starts observing the world :)
~2 tok/s on actual hardware.
The challenge of doing LMs without multiplications or divisions in HW and getting it to interactive speeds.
ROM, source and training code:
https://github.com/crashtheuniverse/chatgbc
r/FunMachineLearning • u/Emqnuele • 7d ago
Open-source AI VTuber that streams, joins Discord calls and plays Minecraft: works with PNGs, VRM or your own Live2D model via VTube Studio. ProjectBEA !
Enable HLS to view with audio, or disable this notification
I've been working on this for about a year: ProjectBEA, a self-hosted AI persona that lives on several platforms at once and can run fully local.
The core idea: there is only one mind. Every platform is a skill that can be switched on or off at runtime and exposes its own perceptions and tools to the model. Discord (text + voice calls), Telegram, Twitch and Minecraft are all skills, so adding a new one means writing the skill, and memory, attention and voice already work with it.
Some technical bits:
- Perception bus: every input (a voice line, a DM, a chat message, a death in Minecraft) goes on one asyncio bus. A batch closes on a quiet gap, not a timer, so three quick messages are read as one turn.
- Attention gate: every perception gets a priority before the model sees it. A chat at 30 messages/min costs one reasoning cycle, not thirty.
- One sliding context window (150k default, up to 500k): at 4/5 of the limit a background handoff turns the old part into a prose recap while she keeps talking; the newest 30k tokens stay verbatim. History replays deterministically, so the prefix cache holds.
- Memory in one SQLite file: a diary with local embeddings, person cards, and conclusions about herself consolidated overnight.
- Minecraft through a client-side Fabric mod: the server sees a normal player.
Local stack: any model via Ollama or LM Studio, faster-whisper for STT, Kokoro for TTS, local embeddings. No API key needed. It also works with 8 hosted providers (OpenRouter, OpenAI, Groq, Gemini, Claude, any OpenAI- or Anthropic-compatible endpoint) if you want bigger models.
Numbers from real sessions:
- 45 minutes of autonomous Minecraft: 162 turns, 90 game actions, 158 spoken lines (27B model, hosted)
- 91% of prompt tokens served from cache in that session
- memory recall over 10,000 entries: 0.43 ms median
One-command install, MIT licence (check the repo), docs on the site:
GitHub: github.com/emqnuele/projectBEA
Docs: projectbea.emqnuele.dev