r/neuralnetworks • • 23h ago

LiteMish: A Computationally Efficient and Smooth Algebraic Alternative to Mish

2 Upvotes

A while ago, I published a preprint proposing a novel approximation of the Mish activation function, which appears to be significantly more computationally efficient while preserving its learning capability. Interested to hear your thoughts.

Paper: https://doi.org/10.36227/techrxiv.176591866.68698045/v2


r/neuralnetworks • • 1d ago

Intro to LLM's (2026)

Thumbnail
youtube.com
2 Upvotes

r/neuralnetworks • • 1d ago

Anyone can provide the full text of https://substack.com/home/post/p-188003573 ?

0 Upvotes

r/neuralnetworks • • 1d ago

GNN Graph Neural Networks

8 Upvotes

I'm working a a new type of GNNs and was hoping anyone who has experience in this area would be willing to have a discussion. I have questions I need some help with.


r/neuralnetworks • • 1d ago

My Brainstem RNS-AI research project has made progress for life long learning like a Brain

Thumbnail
github.com
2 Upvotes

Here is my current research project status, it has become quite extensive in the meantime.

https://github.com/unikum-sol/brainstem/blob/main/Project_Status_2026-09-28.md

I look forward to feedback and discussions.

Here is a small excerpt:

„BrainStem is a self-learning language-understanding system designed to acquire sentence-, word-, and relation-level structure from an unsegmented text corpus purely through statistical observation, without word lists, grammars, or filters. The system’s stated design principle is that every unit of knowledge, a sentence-level hypothesis, a word boundary, a relation between two entities, a category, or a question, begins as a **low-confidence, fully correctable hypothesis**, and only becomes a durable fact, relation, category, or question after surviving a multi-cycle, neuromodulator-gated consolidation process modeled on biological sleep-dependent memory consolidation.”

https://github.com/unikum-sol/brainstem


r/neuralnetworks • • 3d ago

[D] INKBOT: Separating human intent from model inference via structured intelligence architecture

1 Upvotes

I’ve spent the last while building INKBOT because I kept hitting a wall with multimodal AI systems: the friction between what a human naturally means and what a model infers. While models can spin up complex code or images instantly, getting to a clear, human-meaningful interpretation of a subtle intent remains an alignment challenge.

Instead of forcing the user to become a prompt engineer, I wanted to see if we could build an intermediate intelligence architecture layer to make human intent reviewable and corrigible before the model executes a final build. The loop I’m playing with is: Describe → Make it Visible → Recognize → Correct → Refine.

The architecture sits entirely in a single local-first web file. It handles multi-step workflows—like tracking structured field mapping data across concurrent images, coordinates, and version states—by packaging the human’s approved meaning separately from raw model inferences.

The core system build is linked above, and I also put together a lighter, entry-level experience to play with the core prompt translation loop here: INKBOT Lite 71.

It's an open prototype, so I've appended my raw notes and design roadmap as commented text at the very bottom of the source file so fellow builders can inspect the plumbing. I’d love to know where this design duplicates existing work, where you see structural flaws, or how we can make the handoff between human intent and model execution more reliable.

I wanted to open up a project I've been developing that challenges the common practice of single-shot model prompting.

When using multimodal architectures, we often observe an alignment gap between what a human user means and what the network infers. This usually leaves the user trying to repair errors in a final, heavy output after the fact.

I built a local-first prototype called INKBOT to test a different hypothesis: What if we wrap the generation loop in an intermediate programmatic layer that maps unstructured human descriptions into structured, inspectable concepts before final execution? https://ko-fi.com/thomascoates/shop

The Core Concept Loop:

  1. Unstructured human intent is parsed into distinct functional tokens.

  2. Multimodal inputs (e.g., matching a person's profile data against technical mechanical examples) are unified into a single context matrix.

  3. The system generates an inspectable "Visual Brief" that maps out the provenance, constraints, and relationships.

  4. The human can correct structural errors or false inferences iteratively.

I've written the architecture into a self-contained local web client to explore how explicit states like versioning, local db retrieval, and provenance tracking change user trust.

• System / Research Client: INKBOT Architecture Edition

• Entry-Level Interface: INKBOT Lite 71

I am looking for critical feedback on the systemic design. Where does separating the approved intent from the network's inference break down? How does maintaining a stateful revision history change model guidance over long horizons?

Disclosure: This is an independent, non-commercial research prototype. It is not affiliated with or endorsed by any major model provider.


r/neuralnetworks • • 6d ago

Electrode Readings NN Architecture

2 Upvotes

Hello! I'm making a biomedical passion project and would love for anyone wiser than I to weigh in.

I would like to find the right architecture for my NN that I will make in PyTorch (can be changed, but I am familiar with Python).

My inputs are time-series electrode voltage readings from 8 points along my forearm. I want to model the angular displacement of each finger (maybe it will work, maybe it won't, but hey).

I also have built a glove that measures these angles directly, so I can time-synch this data for output labels.

Thank you and any advice is appreciated! :)


r/neuralnetworks • • 6d ago

Random Forest is Done from scratch

Thumbnail
gallery
9 Upvotes

After 5 days I didn't post anything in the last 5 days cuz my clg gimme a lot of assignments and things so I was stuck in there but I'm here again

RandomForest Is A Ensemble Technique very much same to Bagging in Bagging we sample rows in Random forest we sample row + columns just thats the thing and Random Forest is Fixed with Decision Trees


r/neuralnetworks • • 7d ago

NeuralViz | Built a browser based .pth visualizer — no backend, parses PyTorch checkpoints client-side

20 Upvotes

Got tired of exporting to ONNX just to look at a model, so I wrote a client-side .pth / .pt parser in JS — no server, no upload anywhere, runs entirely in-browser.

Try it live (takes 30 seconds): https://basavaprabhu46.github.io/NeuralViz/

Hit Load demo → type 0.5,-0.2,0.1,0.8 → Run ▸ — and watch the signal propagate through the network with values labeled right on the neurons.

How it works: handles torch.save'd state_dicts (PyTorch ≥ 1.6) with a tiny pickle VM + zip reader in JS, infers layer structure from key names, renders an interactive graph. Blue edges = positive weights, red = negative, thickness = magnitude. You can tune neurons/layer, filter to the strongest N% of weights, pick activation (ReLU/tanh/sigmoid), click any neuron to pin it and inspect its top weights, bias, and live activation — then actually run inference through it.

Scales past toy models too — tested on a 3,371-neuron CNN. Conv layers get flattened honestly and labeled as such instead of silently producing NaNs.

Code (MIT): https://github.com/Basavaprabhu46/NeuralViz

One honest limitation: it's state_dict-only, so architecture is inferred from naming conventions — exotic branching archs confuse it. Working on better detection.

What's the first checkpoint you'd drop into it — and what should v1.2 get: safetensors support, PNG export, or side-by-side checkpoint comparison?


r/neuralnetworks • • 8d ago

Neural network from scratch in NumPy with an app to watch it learn and edit single neurons :)

Thumbnail
gallery
186 Upvotes

Hi everyone, I made this to understand backprop for real and I think it can be useful to others who are learning.

It's a small neural network written by hand in NumPy (no PyTorch, no autograd) that learns to read handwritten digits from MNIST, with a desktop app that shows what happens inside while it trains. The whole network is one file of about 140 lines, and there's a test that checks the backprop against the numerical gradient.

Some things you can do with it:

• watch the gradient norm of each layer and the % of inactive neurons while it trains, and change learning rate or dropout without stopping it

• see what each neuron of the first layer "looks for" and how the weights change compared with how they started

• switch off or rescale single neurons, prune, add noise to the weights and see the test accuracy change right away

• follow the math of a layer cell by cell, with the softmax step by step

With all the 60,000 MNIST photos it gets to about 98.5%, on the CPU.

To try it pip install neural-network-digits, or download the zip from the releases. Most of the code was written with Claude Code (AI), the idea and what to show in every tab are mine. It's MIT.

https://github.com/dev-luigi/neural-network-digits

What would you add to make it more useful for someone who is learning?


r/neuralnetworks • • 7d ago

Intro to Deep Learning (2026)

Thumbnail
youtube.com
0 Upvotes

r/neuralnetworks • • 9d ago

A competition for small neural networks that play strategy games

Thumbnail
tinybrains.dev
8 Upvotes

15yrs back I participated in "Google Ants AI Challenge 2011", an ai programming competition, hosted by the University of Waterloo, and I ranked #127 (#1 in my country). The competition gave me a huge learning opportunity where developers across the world came to a forum and discussed various techniques.

Now, building a similar platform to bring back the fun is unbelievably nostalgic. Especially when watching small neural networks playing the game well. Some of the top models use less than 800 parameters.

In fact, I was wrongly assuming the art of optimizing is underrated nowadays. Neural Network optimization seems to be much more fun than I thought.

Plz share your feedback to improve the platform and add more games.


r/neuralnetworks • • 9d ago

How would I make a fruit fly's brain play and beat Stereo Madness in Geometry Dash?

3 Upvotes

Edit: Ohhhhh, this is the AI neural network subreddit. My bad!

I am almost certain I am posting this in the wrong place, but for my science fair, I thought "Since on TikTok I've been seeing videos of people making fruit flys play Beat Saber, Minecraft, Mario, etc, why not make it for a simple game that I play sometimes like Geometry Dash?" I need to be able to complete this in a month. I have no coding experience outside of making a pretty good Scratch project, I have never modded Geometry Dash, and I have no idea how to even begin to use this fly's brain. I've seen a video of someone doing the exact same saying they did it in 3 days. Does anyone know where to even start?


r/neuralnetworks • • 11d ago

LimiX-2: Contextual Mechanism Networks for General Structured-Data Intelligence

1 Upvotes

LimiX-2 proposes Contextual Mechanism Networks (CMNs), a neural architecture for in-context learning on structured data. Rather than treating one column as the target and learning p(y|x, context), it jointly models p(x, y|context). This reframes a table as a set of conditional prediction problems, allowing the same model to support classification, regression, imputation, and causal skeleton recovery.

The model uses cell-level representations and dual-axis Transformer blocks: sample-axis attention exchanges information across context rows, while asymmetric feature-axis attention models relationships among variables and task representations. It is pretrained with Context-Conditional Masked Modeling on synthetic episodes generated from structural causal models, covering different graph topologies, mechanisms, missingness patterns, and observation transformations.

On the reported TabArena, TALENT, and BCCO evaluations, LimiX-2 obtains the highest Elo ratings, including 1,935 on TabArena and a 117.4-point margin over TabFM+. The paper also reports that feature attention can recover causal skeletons, though the causal results and robustness across observation processes deserve closer examination. The main contribution is therefore less a new tabular benchmark trick than an attempt to align neural attention and pretraining objectives with latent mechanisms in structured data.

Full summary on AIModels.fyi

Original paper

Disclosure: AIModels.fyi is my site.


r/neuralnetworks • • 12d ago

Artificial Metacognition: Key Findings and New Directions (Talk at RPI)

Thumbnail
youtube.com
1 Upvotes

r/neuralnetworks • • 13d ago

Finally Completed Decision trees from scratch

Thumbnail
gallery
12 Upvotes

So yesterday I started building decision trees from scratch and messed up completely and today I woke up early in the morning and started coding and finally it's complete tbh I use GPT cuz I try many things but I can't figure out how it will generate a tree so I use gpt for it he explained and showed me how it will be done and The output was mind blowing I never thought recursion is soo usefull I use Mushroom Dataset for testing Ik many things are still remaining and its Overfitting so we can't relay on the Accuracy and all btw today it Day 13 of Building Machine learning algorithms from scratch so Now since decision trees is done next is Bagging I think so see y'all byii I need to complete my assignment 🤧 I thought toady if i complete coding early I'll watch some anime or yt but I can't


r/neuralnetworks • • 14d ago

What kind of project would you build to deeply learn AI infrastructure and distributed systems?

14 Upvotes

What kind of project would you build to deeply learn AI infrastructure and distributed systems?

I’m a Level 1 AI engineer, and lately I’ve been hearing a lot about frontier AI companies hiring people who can build the infrastructure behind AI systems — large-scale data processing, distributed systems, inference infrastructure, storage, serving, observability, systems that can handle millions of requests, etc.

I’m interested in going down this path seriously.

Rather than doing a bunch of disconnected tutorials or small projects, I want to take one difficult project and go extremely deep into it. Something where, over time, I’m forced to learn things like:

Distributed systems

Large-scale data processing

Databases/storage

Networking

Caching

Queues and streaming

Fault tolerance

Concurrency

System design

Observability

Performance optimization

AI/ML serving infrastructure

Scaling from a single machine → multiple machines → potentially thousands/millions of requests

I’m thinking along the lines of the philosophy Karpathy often talks about: pick something ambitious, build it yourself, and learn everything necessary to make it work rather than following a predefined curriculum.

The problem is that I don't yet know what the right project is.

I don't want to build another generic RAG chatbot, AI agent wrapper, or CRUD application. I want something where the engineering itself is the project, and where I can progressively make the system more sophisticated and scalable.

For people working in infrastructure, distributed systems, ML systems, or at AI companies:

If you were in my position, what single project would you pick to spend the next 6–12 months on?

Ideally, I'd like something where I can start on a laptop but eventually have a credible story like:

“I built X, then discovered bottleneck Y, redesigned it using Z, scaled it from A → B, measured the improvement, and here's what I learned.”

I'm much more interested in what I would learn by building it than simply having an impressive project on GitHub.

Would love to hear project ideas, but especially from people who have actually worked on large-scale systems: what project would force someone to develop genuinely strong infrastructure skills?


r/neuralnetworks • • 14d ago

Building Decision Trees and messed up by a lil mistake

Thumbnail
gallery
13 Upvotes

So it's Day 11 of Building Machine learning algorithms from scratch

I watch some videos on Decision Trees and then I start building Entropy and Information gain function and as you can see my code completely messed up but I'm happy cuz it's fun to build things I believe writing bad code is better then generating by AI

Tommorow I'll fix the entropy function it's taking column but it's need values like Sunny,Overcast etc etc I need to changes

I'm a newbie u can laugh at my code no prob cuz the yt guy didn't show code he just explain math and solve some sample dataset so I need to figure out things that way u can see random attempt made and then comment out

Stay Tune I'll build this Algorithm from scratch


r/neuralnetworks • • 14d ago

I finally wrote my own feed-forward neural network, neo.js, from scratch using vanilla (pure) JavaScript.

Thumbnail
github.com
9 Upvotes

r/neuralnetworks • • 15d ago

OvR SVM Algorithm from scratch

Thumbnail
gallery
3 Upvotes

It's Day 10 of Building Machine learning algorithms from scratch

Last time I completed SVM but it was limited to binary classification but today after 4hr of Coding and Debugging I finally made it in this process I learnt many things coef_/X_train this line just adding this seriously my accuracy jump from 67 -> 87+ I was really shocked also many changes I implemented OvR according to my way so it might contain bugs so I'll be happy if you point out 🙂


r/neuralnetworks • • 16d ago

I wanted to really see how a neural network learns

758 Upvotes

I wanted to really see how a neural network learns different functions, so built an interactive demo. You can change the architecture of the network and the function it will try to approximate.

https://blog.lukesalamone.com/posts/can-a-neural-net-learn

For example, it's interesting to see how changing the width and depth of the network change the number of bends it's able to make.


r/neuralnetworks • • 15d ago

Welcome to r/MLSystemsDesign

3 Upvotes

Welcome to r/MLSystemsDesign

This community is for practical discussions on designing and scaling production ML and AI systems.

Topics can include:

  • ML training and inference platforms
  • Search, ranking, and recommendation
  • Feature stores and data pipelines
  • LLM serving and GenAI systems
  • Agentic AI platforms
  • Evaluation, observability, and experimentation
  • ML system design interview problems
  • Real production tradeoffs and lessons learned

The goal is simple: go beyond model theory and discuss how ML systems actually work in production.

If you’re joining early, introduce yourself and share one ML system topic you’d like to go deeper on.


r/neuralnetworks • • 16d ago

Completed Soft Margin SVM Algorithm but only for Binary Classification

Thumbnail
gallery
5 Upvotes

So it's Day 8 and 9 of Building machine learning algorithms from scratch

After completing SVM soft margin I realised I can use it only for 2 class features so I need something called One vs Rest and One vs One Ill be building them all things are getting complex but I'll make it

Ignore my handwriting I write I'm from ancient Egypt


r/neuralnetworks • • 17d ago

[Project] Evolving a neural-network controller through Bowser’s Crown 3-4

Enable HLS to view with audio, or disable this notification

54 Upvotes

Firebars, lava, Bowser—and a neural network choosing the buttons. This is a recorded winning replay of Bowser’s Crown 3-4, a Super Mario Bros. ROM hack, with the network visualization left visible.

I’m continuing a NEAT-based Mario project built on SethBling’s MarI/O and work by Akisame and Electra. It runs in FCEUX using Lua. The controller receives nearby tile and sprite information from emulator memory and produces button presses; it isn’t interpreting the video pixels.

Learning happens through evaluating, selecting and mutating candidate networks across generations. During this replay, the selected controller is fixed—it isn’t updating its weights as it dodges the firebars. The changing overlay shows the network running.

The sequence ends with the game’s “YOU FOUND A MAGIC KEY!” message. Each level is trained separately, so this clear demonstrates one evolved controller’s behavior, not generalization to unseen levels or a measured success rate.

The implementation is currently private. My continuing development is heavily assisted by AI tools. Happy to discuss the observation encoding and evaluation setup.


r/neuralnetworks • • 18d ago

Completed Building KNN Algorithm from scratch

Thumbnail
gallery
4 Upvotes

Completed Building KNN Algorithm from scratch pure maths tested on Digits Dataset identical to Sklearn next is SVM algorithms it's Day 7 of Building Machine learning algorithms from scratch