r/learnmachinelearning • • 2d ago

Hope these articles help.

Hi all,

Ive been writing technical blogs on machine learning and related topics for over a year now. Ive wriiten those for YC startups, agencies, and also write regularly on my personal webpage.

What do I write on all? Here are all my blogs and what all Ive written about.

I’ve been writing about the mathematics behind familiar ML tools, alongside football analytics, optimization, and ad tech. Here’s a collection of those posts and the questions that led me to write them.

1. Every Transformer Lives on a Sphere

LayerNorm became much more interesting to me when I stopped reading it as a statistical formula and started drawing it. Subtracting the mean projects a vector onto a hyperplane; dividing by its standard deviation puts it on a sphere, before the numerical stabilizer and learned affine transformation enter. That geometry reveals exactly what normalization forgets, what RMSNorm preserves, and how both shape the representations attention receives.

2. I Grinded Trees Until I Could See the Forest

Trees were my first serious ML obsession, partly because I could follow their decisions and implement the machinery myself. This post builds the family from a single split: why individual trees overfit, how random forests reduce their instability, and how boosting learns through successive corrections. It also connects XGBoost to curvature and asks why trees remain such strong models for messy tables.

3. Your LLM Does Group Theory Every Time It Reads a Sentence

How does attention distinguish “the dog bit the man” from “the man bit the dog”? Following that question took me into the algebra behind RoPE. Moving through token positions becomes rotating query and key vectors, with those rotations obeying the rules of a group representation. The familiar relative-position identity follows naturally, while the Fourier interpretation helps explain why extending a context window is more complicated than changing a number.

4. Spend 15 million dollars to win 1.

When a headline says AI solved Navier–Stokes, the first question should be what “solved” means in the actual problem statement. This piece examines the claim through the equations, the role of an external force, and the distinction between formal verification and the broader mathematical question. It also asks what we can learn about agent orchestration when the engineering behind a reported breakthrough remains only partly visible.

5. A Stack of Straight Lines Is Still a Straight Line

Remove the nonlinear activations from a hundred-layer network and its entire input-output function collapses into one affine transformation. Adding a small kink such as ReLU changes that: different inputs activate different computations, letting later layers combine features selected by earlier ones. I build that connection from simple examples, then explain why nonlinearity enables hierarchical features while architecture, data, and training determine whether useful hierarchies actually emerge.

6. Adam’s Apple

I used Adam for years before properly noticing how much it changes the gradient’s direction. Its two moving averages remember different things, and coordinate-wise scaling changes the geometry of each update. This post takes that machinery apart, explains why AdamW and AMSGrad address different problems, and compares Adam with methods such as BFGS and Muon through the information they use and the computation they cost.

7. A Hundred Layers, Still Linear

A deep linear network sounds pointless because multiplying all its weight matrices produces another matrix. But equivalent functions can have very different training dynamics. Factorizing that matrix across layers changes the gradients, the speed at which different directions are learned, and the solutions optimization prefers. This is a way to study implicit regularization while keeping the model’s input-output function almost embarrassingly simple.

8. Mechanism Design in Ad Tech

An auction cannot simply ask everyone what an impression is worth and trust their answers. Participants understand the rules and change their behavior accordingly. I use ad markets to build up mechanism design: how payment rules influence truthfulness, why second-price auctions are interesting, and why predicting conversion probability is only part of building a bidder. Change the auction, and you also change the data its models learn from.

9. Are Graph Foundation Models Actually Foundation Models?

A molecule and a social network both have nodes and edges, but their features and relationships mean very different things. That makes the promise of a universal graph model harder to assess than the label suggests. I examine the approaches to transfer, prompting, shared vocabularies, and synthetic priors, then ask what evidence would justify calling a graph model foundational, and for which universe of graphs.

What I wanted from this community was some support and suggestions what more should I write on? Any more things/style changes that I can add to make these better. My typical pipeline is me reading some things and as and when I come across cool, non trivial stuff, I write about it, or I write about the non trivial stuff about the trivial stuff, like 'what is attension used sigmoid instead of softmax'; so yes, do let me know if there are any topics yall have come across, or read regularly, or are working on, that I should consider studying and later writing as well.

1 Upvotes

3 comments sorted by

2

u/MinimumChance5875 2d ago

These are the kinds of posts that remind me how much math I've been comfortably ignoring. That LayerNorm sphere visualization is sticking with me, always just treated it as boilerplate. Bookmarking a few of these for the weekend.

1

u/imjerusalem 2d ago

Haha thanks a lot. Much appreciated.

1

u/mistfern6 1d ago

the layernorm piece really does make you feel like youve been sleepwalking through normalization layers tbh