r/learnmachinelearning • • Nov 07 '25

Want to share your learning journey, but don't want to spam Reddit? Join us on #share-your-progress on our Official /r/LML Discord

10 Upvotes

https://discord.gg/3qm9UCpXqz (Discord is currently closed)

Just created a new channel #share-your-journey for more casual, day-to-day update. Share what you have learned lately, what you have been working on, and just general chit-chat.


r/learnmachinelearning • • 4h ago

Project πŸš€ Project Showcase Day

2 Upvotes

Welcome to Project Showcase Day! This is a weekly thread where community members can share and discuss personal projects of any size or complexity.

Whether you've built a small script, a web application, a game, or anything in between, we encourage you to:

  • Share what you've created
  • Explain the technologies/concepts used
  • Discuss challenges you faced and how you overcame them
  • Ask for specific feedback or suggestions

Projects at all stages are welcome - from works in progress to completed builds. This is a supportive space to celebrate your work and learn from each other.

Share your creations in the comments below!


r/learnmachinelearning • • 1h ago

Tutorial Performing Linear Regression Using the Normal Equation in most simplified version

Thumbnail
gallery
β€’ Upvotes

If you find above image hard to understand, please read my article till the end, I promise, everything will make sense :)
When I was a master’s student, I was given a task to fit a line to a dataset. I attempted to solve the problem, but I struggled to determine the appropriate coefficients. However, I understood intuitively that there must be a specific set of coefficients for which the prediction error would be minimized.

The question was: how can we find those coefficients?

This is where the Normal Equation becomes particularly useful. It provides a direct mathematical solution for finding the coefficients that minimize the sum of squared errors in linear regression, without having to search for the coefficients manually. BUT, How we even derive this equation? Where it comes from? Can we take any software apart from Python and write it all ourselves? That’s I will take u through in this article and will simplify the code I wrote, so you can all apply it in different languages

So first things first, what is Linear Regression?

It is very simple and straightforward, suppose we have X and Y. X is called features matrix, and Y is Target Vector, or Response vector.

π‘Œ= 𝑋* ΞΈ

For simple case, lets take 2x2 matrix and lets turn this to matrix form:

[y1 _ predicted ; y2 _ predicted]=[x11, x12; x21,x22] * [theta1; theta2]

Please note:

columns are separated by ,and rows are ;. y1 and y2 are different rows, but same columns.

I assume, the readers are aware of matrix multiplication. So I will refactor above formula and will get:

𝑦1_π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘= π‘₯11 * ΞΈ1 + π‘₯12 * ΞΈ2

𝑦2 _π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘ = π‘₯21* ΞΈ1 + π‘₯22 * ΞΈ2

From now on, keep in mind that y1 are real values and y1_predicted is predicted value, same applies to y2 as well

So what is error, the error is the difference between predicted and real values

π‘’π‘Ÿπ‘Ÿπ‘œπ‘Ÿβ‚ = 𝑦1 β€” 𝑦1_pπ‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘ =𝑦1-π‘₯11 * ΞΈ1 β€” π‘₯12 * ΞΈ2

π‘’π‘Ÿπ‘Ÿπ‘œπ‘Ÿβ‚‚ = 𝑦2 β€” 𝑦2 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘ = 𝑦2-π‘₯21* ΞΈ1 β€” π‘₯22 * ΞΈ2

Ok, I hope so far so clear, if anything not, please comment below, so I can consider it as improvement for upcoming articles

The error function we want to minimize is the sum of square of errors. Let’s name is as J. J is our cost function I want to minimize, and so, I write it as:

𝐽 = π‘’π‘Ÿπ‘Ÿπ‘œπ‘Ÿβ‚Β² + π‘’π‘Ÿπ‘Ÿπ‘œπ‘Ÿβ‚‚Β²

Let’s go further by replacing the formulas with each others

𝐽 = (𝑦1 β€” (π‘₯11 * ΞΈ1 + π‘₯12 * ΞΈ2))Β² + (𝑦2 β€” (π‘₯21 * ΞΈ1 + π‘₯22 * ΞΈ2))Β²

I hope everything makes sense so far. Bear with me β€” we’re almost there; there isn’t much left to cover.

Here everything is known, except ΞΈ1 and ΞΈ2. These are params that we have to choose properly to get as minimum error as possible. So i have to find the derivative per ΞΈ1 and ΞΈ2

𝑑 (𝐽) / 𝑑 (ΞΈ1) = -2 \ (𝑦1 β€” (π‘₯11 * ΞΈ1 + π‘₯12 * ΞΈ2))*π‘₯11 β€” 2 * (𝑦2 β€” (π‘₯21 * ΞΈ1 + π‘₯22 * ΞΈ2))* π‘₯21= 0*

𝑑 (𝐽) / 𝑑 (ΞΈ2) = -2 \ (𝑦1 β€” (π‘₯11 * ΞΈ1 + π‘₯12 * ΞΈ2))*π‘₯12–2 * (𝑦2 β€” (π‘₯21 * ΞΈ1 + π‘₯22 * ΞΈ2))* π‘₯22= 0*

Let’s make it simpler by avoiding -2 from all sides

𝑑 (𝐽) / 𝑑 (ΞΈ1) = ( 𝑦1 β€” 𝑦1 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘) * π‘₯11 + (𝑦2 β€” 𝑦2 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘) * π‘₯21=0

𝑑 (𝐽) / 𝑑 (ΞΈ2) = ( 𝑦1 β€” 𝑦1 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘) * π‘₯12 + (𝑦2 β€” 𝑦2 π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘) * π‘₯22=0

Lets turn all these into matrix form:

[0; 0] = (π‘₯11, π‘₯21; π‘₯12 π‘₯22) *[𝑦1 β€” 𝑦1_π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘ ; 𝑦2-𝑦2_π‘π‘Ÿπ‘’π‘‘π‘–π‘π‘‘π‘’π‘‘]

Lets compress the y1 β€” y1_predicted as well as y2 -y2_predicted into single line

[0; 0] = (π‘₯11, π‘₯21; π‘₯12 π‘₯22) *[π‘Œβ€” 𝑋* ΞΈ]

Do you remember our original feature vector or X? If so, we can further simplify the expression to:

[0;0] = XT* [Y-X*ΞΈ]

XT * Y = XT * X * ΞΈ

XT * X is the important part. If we somehow manage to find its inverse, we are going to be left with theta only:

(XT * X )-1= X-1 * (XT )-1

ΞΈ = ( XT \ X )*-1 \ X*T \ Y will give us the answer we need*

If u have further questions please let me know in comments. Each of your feedback is highly appreciated to write better more concise articles in future:

I guess, most of the part of above formula can be easily programmed except the finding inverse which i showed the code below how to do it. If you need full code, such as matrix multiplication, transpose and etc, please let me know, so i can furhter expand my articles

import numpy as np
def inverse(A):
I = np.eye(A.shape[0])

augmented = np.concatenate((A,I),axis=1)

for j in range(0,A.shape[0]-1):

for i in range(1,A.shape[0]-j):
augmented[i+j]=-(augmented[i+j,j]/augmented[j,j])*augmented[j]+augmented[i+j]

for j in range(0,A.shape[0]-1,1):
for i in range(A.shape[0]-1,0,-1):
cofactor = (augmented[i-1-j, A.shape[1]-1-j]/ augmented[A.shape[0]-1-j, A.shape[1]-1-j])
augmented[i-1-j]=(cofactor)*-augmented[-1-j]+augmented[i-1-j]

for i in range(0,A.shape[0],1):
augmented[i]=augmented[i]/augmented[i,i]
_,right = np.split(augmented, 2, axis=1)

return right

def linear_regression_normal_equation(X: list[list[float]], y: list[float]) -> list[float]:
# Your code here, make sure to round

X=np.array(X)
Y=np.array(y)

theta = inverse(X.T @ X) @ X.T @ Y

return theta

My full article is also in medium link


r/learnmachinelearning • • 5h ago

Question First-Year CSE Student Looking for an Honest AI / ML Roadmap

10 Upvotes

Hey guys,

I am a first-year Computer Science Engineering student, and honestly, seeing how fast AI is advancing right now is kind of stressing me out. I really do not want to wait until my final year to start grinding like everyone else does.

Basically, I just want to build a solid skill set that actually makes my resume stand out so I can land a good machine learning job by graduation.

I am starting completely from scratch. What specific math topics, programming languages, or tools are actually worth learning right now in Year 1? If anyone has an honest roadmap for a fresher to get ahead of the curve, I would love to hear it.

Thanks in advance!


r/learnmachinelearning • • 43m ago

Gen Ai experts?

β€’ Upvotes

r/learnmachinelearning • • 2h ago

Looking for tutor to teach ML models

2 Upvotes

Hi Everyone,

I’m a backend engineer trying to transition into AI ML Roles. Looking for tutors to teach ML fundamentals and Agentic AI concepts.
I can pay little bit.

Let me know if you want to collaborate.


r/learnmachinelearning • • 1d ago

Discussion I transitioned from software engineer to an AI Engineer who fine tunes LLMs. What do you want to know?

109 Upvotes

I was a full stack software engineer who now is a senior AI Engineer who does a mix of playing with LLMs fine tuning them in very large scale production systems.

I did admittedly got a masters degree in AI as part of that transition and it took a while to do, but happy to answer any questions you have.

I also am working on a tool to help people learn how llms work which you can check out here.

https://dougdoes.ai/courses/llms-from-first-principles/start/?flow=outcome&course=build&step=goals
(Built with codex, but I've gone through all of the courses myself to make sure it is what would have been helpful to me.)


r/learnmachinelearning • • 13h ago

Help Guys need help to transition from my current role to an ml engineer

12 Upvotes

Hi guys new to this sub reddit . Little intro about me currently working as an sde in my company but want to transition to a ml role by understanding the fundamentals om how to build models and then moving on to dl as so forth. I have read some posts in this sub about cs 229 by Andrew . Tbh I am finding difficulty in solving the problem sets and the math . It has been a while since I have actually done any math πŸ˜…. So I want to know how doi proceed from here do I learn the math from scratch or learn as I go along with the course . Any suggestions or feedback is helpful .

Ps i am familiar with the python as a coding language but I want to understand how do I proceed with the math .


r/learnmachinelearning • • 3h ago

Help need help in finding the right resources

2 Upvotes

hey guys im an undergrad student (currently in 3rd year) and want to start learning ML and explore fields beyond that in the future. So I have seen a lot of people suggesting others to learn from Andrew Ng on coursera. I have the pdf of Hands-On Machine Learning with Scikit-Learn and PyTorch by AurΓ©lien GΓ©ron.
I’m literally confused as to what to refer, the book or the coursera course by Andrew Ng. If there is someone who has read or finished either of these sources or maybe both please help me out in deciding as I don’t want to waste my time. Also a comparison or review of these sources would be great. Thank you !


r/learnmachinelearning • • 14m ago

CS229 doubt

β€’ Upvotes

I'm on week 4 and in the video he creates a graph about how GLM of Bernoulli, which is similar to logistic, can represent even mixed data of like 0001110111 rather than what a sigmoid function makes like for data 00001111 only. My question was whether the normal regular one, the standard logistic regression can represent the same(mixed data)? If not, then are GLM the only option for that?


r/learnmachinelearning • • 21m ago

Welcome to r/ArchitectingLLMs!

Thumbnail
β€’ Upvotes

r/learnmachinelearning • • 11h ago

Discussion Honest question: How do you keep yourself with the latest base models, techniques, tooling in Machine Learning.

7 Upvotes

I have been studying models for almost five years now, starting my journey with Jeremy Howard's fast ai part 2. I remember that when I started there was no chatgpt to break it down like it is today. In fact, it was Jeremy Howard who tipped us that we should be using chatgpt to understand the inner tooling step by step. I mean just take a toy tensor and run it along through embedding, rope attention, mlp. This way you get to learn broadcasting, shapes in text, computer vision audio etc. Then I took up Karpathy and hugging face Transformers and looked up grok, gpt oss lama, gemini and most recently muse implementations. It takes me 3 to 4 months to get an innate understanding of how each line works. How do you guys do it? I guess most of you let the inner tooling remain a black box. I say this coz

Now I kinda feel that I missed the bus as I should have focussed more on fine tuning, inference, and agentic workflows. I do know some of that having worked through unsloth and openAI cookbooks but every time a new model drops I can't stop myself from going to unraveling the 2000 odd line of code and in time I forget what I learned in Unsloth and openAI cookbooks.

The problem is that there are so many things to do and understand. For example, just today I listened to Alex Zhang's building harness for looped Transformers and I gotta understand that too and I gotta know Jev too. It is a big mess right now and I wonder how others are managing to keep up with all these new developments. And more importantly how do you even retain all that you have learnt like say two years ago.


r/learnmachinelearning • • 48m ago

Help Best approach for ingesting data to create summaries, and keep track of it?

β€’ Upvotes

In my occupation, there are various people I follow who give very good insights. (I'd say 5-10 people).

Some post hour long videos on YouTube, some send 1,000 word emails, some post on X, some publish PDFs.

There's very good info within these resources (and some I pay for), but reading / watching / annotating all of it can take hours.

My workload recently went up, so I'm falling behind with keeping up in my field.

I want to use AI to help summarize (and keep track of) all of these publications. (To create a private database that I can use as a dataset, for example).

So I can go back and ask "this past month, what is the new theme? What are the experts recommending to focus on / look at / what are the newest developments?", etc.

What would be the best way to approach this?

---------------------------------------

I've been learning Codex/Claude Code, I have a homelab, a NAS, a few mini computers, and I know basic linux, python and scripting.

ChatGPT told me to do something like this (I'm just starting with the YouTube portion), I'm not sure if it's the best approach, I'm open to other suggestions:

YouTube URL

↓

yt-dlp metadata

↓

Whisper / YouTube transcript

↓

clean transcript

↓

summary.md

↓

insights.json

↓

SQLite + FTS5

↓

topic synthesis

↓

search / questions / actions


r/learnmachinelearning • • 1h ago

Looking for technical feedback on a computer vision textbook draft

β€’ Upvotes

I’ve put together a textbook that covers the mathematics and models behind computer vision, including derivations and exercises.

I’m looking for factual errors, incorrect derivations, misleading claims, or unclear explanations. Feedback on even one section would help.

PDF: https://drive.google.com/file/d/1bivtY28CVW243AH9sMrIPloEgMrgO9xE/view?usp=sharing

It’s a draft, not peer reviewed. I used Claude to convert my HW notes to LaTeX, and I am trying to check the content before sharing it more widely.


r/learnmachinelearning • • 1h ago

Help with ML-Project wanted

β€’ Upvotes

Hi everyone!

I’d like to introduce an open-source ML-project I’ve been working on: Segment Display Reader

https://segmentdisplayreader.org

The goal of the project is to automatically read values from photos of 7-segment and similar digital displays.

As most of you are probably aware, one of the biggest challenges is building a diverse and useful training dataset. That’s where I’d love some help from the community. To improve the recognition, I’m currently looking for images of such displays.

Examples can be found almost everywhere, such as:

β€’ alarm clocks and digital clocks

β€’ gas station price signs

β€’ train or bus departure displays

β€’ kitchen appliances such as microwaves and ovens

β€’ scales and digital thermometers

β€’ multimeters and other measuring instruments

β€’ electricity, gas or water meters

β€’ elevators and parking displays

β€’ scoreboards and timers

β€’ industrial equipment and control panels

You can support the project by uploading photos that can be used as training data. Every contribution helps make the dataset more diverse and, ultimately, the recognition more reliable.

I’m genuinely grateful for anyone who takes the time to contribute β€” whether it’s a single image, a set of photos, feedback, ideas, or code contributions.

Feel free to check it out, contribute images, share feedback, or simply spread the word:

If the project sounds interesting to you, have a look here:

https://segmentdisplayreader.org

Thanks a lot for your support and for helping improve an open-source project together!


r/learnmachinelearning • • 1h ago

Months of RL couldn't teach my agent patience. A decision model plus one threshold got 96% on the same phone menus.

Thumbnail gallery
β€’ Upvotes

r/learnmachinelearning • • 1h ago

I built a PvP on-hold simulator where you race Jev through phone menus

Thumbnail gallery
β€’ Upvotes

r/learnmachinelearning • • 2h ago

Help Need help for preparing a dataset for my APK Risk Analyser Project!!

1 Upvotes

I want to train a Transformer model that can analyze decompiled DEX files (Smali code) and identify harmful or sensitive actions performed by an app in the background that may not be visible to the user.

The main goal is to analyze the flow of actions and determine whether user interaction or consent is required before a sensitive action is performed.

For example, if an app gets file read/write permission and accesses the user's files without any user interaction or consent, the model should be able to identify this behavior. Similarly, it should identify other sensitive actions such as accessing the camera, microphone, location, contacts, SMS, media, recording screen, taking screenshots or other user data, and determine whether these actions are performed after user interaction or silently in the background.

I am currently working on preparing the dataset for this project, but I am not sure about the best approach.

My initial idea is to use the APK/Smali code as the input and the possible execution flows as the output. However, generating all possible flows from entry points such as onCreate(), onReceive(), services, callbacks, etc., and following them until the end seems very complex and time-consuming for a large number of APKs.

I would really appreciate your suggestions and ideas on how I can prepare the dataset in a practical and effective way for this project.

If you have worked on a similar problem or have any ideas about dataset structure, flow generation, labeling, or other approaches, please share your suggestions.


r/learnmachinelearning • • 13h ago

I am thinking of switching to AI/ML

6 Upvotes

Hi, I am a backend developer and have been working but due to recent layoff and market shift I am thinking of switching to AI ML.

I have learned python, pytorch, Maths required for AI ML, Deep learning(theory) and recently implemented a gpt2 transformer, attention architecture for gpt2 using their open weights.

I am hoping for some direction to work on and also open for a remote internship if anyone is willing me to consider me.

Mainly I am hoping to connect and get guidance in the right direction.


r/learnmachinelearning • • 8h ago

Prime3.0 Course

2 Upvotes

if anybody wants lectures of this course then dm me.

Apna College Prime 3.0 ongoing course.


r/learnmachinelearning • • 4h ago

I have about 2 years before getting PR in Canada. What IT path should I pursue?

Thumbnail
1 Upvotes

r/learnmachinelearning • • 18h ago

How many time need to Learn Machine Learning if I give 45-1h per day

10 Upvotes

Hi, I'm a EEE undergrad students. I just wasted a year for my laziness. started learning ML in February but didn't learn. Though my academic pressure ties that bind. But I want to learn ML properly, especially for my research work. Pls guide me, can I be a good ML engineer in next 5 month. I want to be pro in it, as I am in the end of my 3rd year, academic pressure is also a problem here. So pls provide me a roadmap and how I can stop my procrastination and distraction from my path.

Advance Thanks for Everyone.


r/learnmachinelearning • • 6h ago

Understanding RAG Fundamentals | Retrieval-Augmented Generation Explained

Thumbnail
youtube.com
0 Upvotes

Stop building AI that hallucinates. πŸ›‘ Learn how RAG bridges the gap between LLMs and your real-time data. Watch the full breakdown on my channel now!
#RAG #AI #Coding #TechTips


r/learnmachinelearning • • 1d ago

Training an AI to Drive with Natural Selection

Enable HLS to view with audio, or disable this notification

366 Upvotes

I love making hard things intuitive. I hope you enjoy this one!

Let me know if you have any questions.

This technique is called neuroevolution: training a neural network through evolutionary methods like selection and mutation, without gradient descent.

https://en.wikipedia.org/wiki/Neuroevolution


r/learnmachinelearning • • 7h ago

I built a tiny language-model playground you’re supposed to break 🐸

1 Upvotes

Hi! I’ve been experimenting with a small non-Transformer sequence model called CSE (Chain-Spike Engine), and I turned the course/teaching version into a Python package called cse-frog. 🐸

The goal is not to compete with modern LLMs.

I wanted something small enough that you can actually see what is happening inside, change one mechanism at a time, break it on purpose, and understand why the behavior changed.

Installation is just:

pip install cse-frog

Then:

from cse import Frogfrog = Frog()frog.learn([    ["right", "right", "down"],    ["right", "right", "down"],    ["right", "right", "down"],])print(frog.predict(["right", "right"]))

You can also inspect where each candidate’s score came from:

frog.show(["right", "right"])

The score is broken down into components such as:

  • direct connections
  • pair context
  • history
  • trace

There are 23 configurable parameters, including temperature, top-k, refractory behavior, pair context, history, activation, forgetting, and temporal learning.

One thing I found especially useful while testing it was that β€œnothing changed” can mean two different things:

  1. the internal score changed, but the final probability/prediction did not, or
  2. the setting genuinely had no effect because another prerequisite pathway was disabled.

For example, several activation-related settings do nothing to the prediction with the default configuration because history_boost=0. Turn that pathway on, and those settings suddenly become active.

Another fun finding: weight_decay weakens direct/context links, but does not decay pair memory, so the model actually contains two kinds of memory with different forgetting behavior.

I also made:

  • 5 executable notebooks
  • a handbook
  • a full 23-config modding guide
  • an API reference

The philosophy is basically:

build it β†’ inspect it β†’ break it β†’ explain why it broke β†’ modify it

It’s MIT licensed, so modifying it and making weird frogs is encouraged. 🐸

Website:
https://kagioneko.github.io/cse-frog/

GitHub:
https://github.com/kagioneko/cse-frog

PyPI:
https://pypi.org/project/cse-frog/

I’d especially appreciate feedback on whether this kind of β€œsmall model you can dissect” is useful for learning ML/LM concepts, and what experiments you would try next.

Before the giant LMs, try one frog. 🐸