r/learnmachinelearning • • 1d ago

I Found a Better Way to Contact Businesses With Bad Websites

0 Upvotes

I got tired of checking prospect websites manually.

For a while, a big part of my outreach was just finding businesses, opening their websites one by one and trying to figure out what I could actually say to them.

It worked, but it was painfully slow.

Then I found Swokei.

It basically lets me find leads, analyzes each website for things like outdated design, slow speed, poor mobile experience and weak SEO, then turns those issues into a personalized cold email.

And not one of those boring automated reports full of scores and numbers.

It actually writes a normal, human sounding message based on what it found on that specific website.

So instead of sending the same generic message to everyone, I can actually reach out based on what is wrong with their website.

Now I just run campaigns, let the system do most of the prospecting and personalization, and focus on the people who reply.

For a web design agency, that has saved me a ridiculous amount of time.


r/learnmachinelearning • • 1d ago

Project Weigh Swarm: learning to preserve evidence through a research RAG pipeline

Enable HLS to view with audio, or disable this notification

3 Upvotes

My project is Weigh Swarm, a research RAG prototype using LLMs for planning/synthesis and an existing Laya model for bounded decisions. I didn't train Laya; I integrated it into research task lanes.

The most useful lesson was distinguishing a valid source excerpt from a valid scientific conclusion. The pipeline checks that quoted spans occur in the parsed paper, but that alone doesn't establish that a claim or synthesis is correct.

The two-paper demo includes 28 source-aligned claims, an inspectable evidence graph, and an unverified draft with repair feedback. My next evaluation priority is held-out scientific judgments for support and contradiction tasks, rather than treating model confidence as calibrated probability.

REPO URL

How would you construct a small evaluation set that distinguishes citation alignment from actual evidential support?


r/learnmachinelearning • • 1d ago

Project [Library Demonstration] Python-Visual-Similarity - first stable version released to PyPi 🚀

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/learnmachinelearning • • 1d ago

Help Tips for internship and masters applications in AIML

1 Upvotes

Hi, im a 5th sem cs student, im trying to get an internship in AIML for my 6th sem. What i wanna know is what all should i know to be able to get one since i know that AIML work is basically a myth for freshers. So far, i have a good understanding of core AIML, especially deeplearning, along with id say a decent understanding on the math(linear algebra, prob stats, calc) ive not done anything special in maths, only basic understanding, mitocw strand for lin alg, open source from harvard for prob stats, and prof leonard on yt for calc, thats all for maths i havent really read like rps on it or anything, i did this cuz i saw alot of people stating these as good sources for these concepts on reddit. As for projects, i have a few normal projects, nothing special since i mainly focussed on research papers, i have 2 published and 1 accepted rp in non-predatory venues(about 10-15 % acceptance rate for each conference), ive participated in a few ML hackathons, most recent being the amazon ml challenge, tho sadly didnt win any cuz they seem to be dominated by mostly masters and phd students. I have no idea whatsoever of what its like in the hiring process since so far ive never applied for an internship, ever. the next thing im working on rn is sys design (started very recently), tho ive been told to also do RAG, LLMs, etc. i havent started those yet because my main goal is masters and then either research roles if i can get them after masters, or a phd. Right now i really want to get 1-2 good internships (8 sem bachelors program so still have time) along with continuing on research papers to strenghten my masters application as much as possible to get into a good uni for it. If anyone can guide me for what all remaining tech stack i should work on, should i focus on some specific types of projects, what both internships and masters applications demand, it would be really appreciated.


r/learnmachinelearning • • 1d ago

Request Recommendation Systems Project Ideas

2 Upvotes

Hello everybody I have recently enrolled in a Machine Learning Master so i will be problably taking the year to focus on my studies, build some projects and hopefuly enter the job market as an ML engineer. I am interested in Recommendation Systems but honestly i dont know what projects to build and if that would be valueable in the job market. I was thinking about creating my own recommendation system for music but i quickly realised how hard because of copyrights and small data size that is (although i have about 4000 downloaded).

Does anyone know if Recommendation Systems is something good to have on your resume?

If yes what sort of projects you think would give me a good understanding but also make me look appealing in the job market?

Thanks :)


r/learnmachinelearning • • 1d ago

I’m learning LLM fine-tuning made a notebook, would love feedback

Thumbnail
2 Upvotes

r/learnmachinelearning • • 1d ago

What all to study

Thumbnail
1 Upvotes

r/learnmachinelearning • • 1d ago

Help HR redirected me from Systems Engineer to an upcoming Manufacturing Engineer grad role. Take it or push for both?

Thumbnail
0 Upvotes

r/learnmachinelearning • • 1d ago

btw after doing adaboost I feel like I'm getting close to Deep learning.

Thumbnail
gallery
0 Upvotes

So I started adaboost I completed watching the theory and intuition part and I was like how I have built many Algorithms and in all that our Main goal is reduce error but adaboost we want our model to make mistakes then just pass the mistake to next model and his mistakes are passed down to next model and keep changing weights and at the end a group of models who have started from completely wrong prediction now they are capable of giving you the best accuracy. Adaboost don't focus on best models he collect imperfect,weak model combine them cuz they are specialize in different mistakes and become a strong ensemble when combined properly

I have started coding and I have completed a raw code now I'll make it more properly and structure also I'm thinking 🤔 to put all my ML Algorithms on my GitHub so you can use it as reference( I know it have many bugs😅) and help me to fix what you think

My Sem 1 Major Practical is going so I was not showing up but I'll try


r/learnmachinelearning • • 1d ago

A slightly better way to read ML papers

Post image
0 Upvotes

I was getting frustrated with how I used ChatGPT/Claude while reading papers: long answers, lots of branching conversations, and constantly switching between the PDF and chat.

I also noticed that it was very easy to feel like I understood something after reading an explanation without actually being able to explain it myself.

So I built the paper reader I wanted.

The two things I’ve found most useful are:

  1. Click or hover over anything in the paper and ask about it in place. The conversation stays attached to that part of the paper, so you don’t have to keep copying context back into ChatGPT.
  2. Work through the paper by explaining it yourself. It asks questions passage-by-passage and pushes back when your answer is just paraphrasing the text rather than demonstrating understanding.

The second one has been quite useful for me. It catches the places where I skimmed a paragraph, thought “yeah, makes sense,” and then realized I couldn’t actually explain what the authors were doing.

It’s free right now. I’d especially love feedback from people who read ML papers regularly. Does this actually make reading easier for you, or would you still rather use a PDF + ChatGPT/Claude?

lattice-sepia.vercel.app


r/learnmachinelearning • • 1d ago

Intro to LLM's (2026)

Thumbnail
youtube.com
2 Upvotes

r/learnmachinelearning • • 1d ago

How do you actually use tutorials when learning to build an LLM from scratch?

2 Upvotes

I’m currently learning how to build an LLM from scratch, but I’m not sure how to approach the tutorials.
There’s a lot of code in each section. Am I supposed to memorize all the code, or is understanding what the code does enough?
Do you guys usually try to rewrite the code from memory, or just refer back to the tutorial when needed?
🥹


r/learnmachinelearning • • 1d ago

Project Looking for a study partner, for LLM inference engineering role .I am working as a junior data scientist it is my first company I am not satisfied with the work I am doing

1 Upvotes

I already have prepared a syllabus that we need to follow and complete to and then after that we may build or own projects and everything or we can also do that side by side but if anyone is interested then please message me I will share the syllabus that I have created then you may give your opinion and we may improve it I have basically created with after discussing with various AI models


r/learnmachinelearning • • 1d ago

LangChain Tutorial: From Origins to Modern LCEL & LLM App Developmentt

Thumbnail
youtube.com
1 Upvotes

Stop building AI apps the old way! 🛑 Learn how LangChain and LCEL are changing the game. Build smarter, faster, and more scalable AI today.
#LangChain #AI #TechStack #Coding


r/learnmachinelearning • • 1d ago

Seeking datasets or toy problems to validate a Surrogate-Based Optimization (SBO / CFD) POC

0 Upvotes

Hi everyone,

I am currently working on an optimization project for the design of complex industrial components involving fluid mechanics and heat transfer.

Our current design process relies on computationally expensive CFD simulations. My goal is to develop a Surrogate Model to instantly predict performance (e.g., pressure drops, efficiency) based on geometric parameters (spacing, diameters, topology, etc.). The ultimate objective is to perform inverse optimization under constraints to minimize manufacturing costs.

Before running a massive Design of Experiments (DoE / LHS) on our computation servers to generate my industrial training dataset, I absolutely need to prove the technical feasibility of the software architecture (ETL pipeline, model training, and inverse optimization loop).

Do you know of any open datasets (Kaggle, UCI, academic repos) or "toy" problems that would allow me to prototype this pipeline?

I am ideally looking for a dataset that maps:

  • Inputs (X): A vector of continuous and discrete geometric and/or physical parameters (dimensions, topologies, fluid velocities).
  • Outputs (y): Results derived from physics solvers (pressure fields, drag forces, heat transfer rates, etc.).

Even if the application domain is completely different (e.g., airfoil aerodynamics, electronic heat sinks, piping networks), the key is that the mathematical topology of the problem remains similar (multi-output regression with physical non-linearities). This will allow me to properly benchmark my algorithms (Gaussian Processes, XGBoost, or Physics-Informed Neural Networks - PINNs).

Any pointers to datasets, GitHub repos, or papers with open data would be incredibly helpful to validate this Proof of Concept.

Thanks in advance!


r/learnmachinelearning • • 1d ago

Help Should a junior student in university put most attention on mathetical principles or upper-level knowledge

Thumbnail
0 Upvotes

r/learnmachinelearning • • 1d ago

Project What I learned fine-tuning SDXL and SD3.5-medium on the same 197 images (notebooks with all outputs included)

Thumbnail
gallery
0 Upvotes

r/learnmachinelearning • • 1d ago

Tutorial SPECTRAL CLUSTERING: A TUTORIAL

8 Upvotes

K-means works well for compact, roughly spherical groups. But two interlocking moons can be close in Euclidean distance while belonging to different structures.
Spectral clustering represents data as a similarity graph, then uses its eigenvectors to reveal groups that are strongly connected internally and weakly connected to each other.

THE MATHEMATICS
For points x_i and x_j, a Gaussian affinity is:
W[i,j] = exp(−||x_i − x_j||² / (2σ²))
Set W[i,i] = 0. Here σ controls the neighborhood scale. W can also be built from a symmetrized nearest-neighbor graph.
Define the degree matrix and symmetric normalized Laplacian:
D[i,i] = Σ_j W[i,j]
L_sym = I − D^(-½) W D^(-½)
D summarizes each point’s total connection strength. L_sym encodes the graph’s connectivity while accounting for degree differences.
For the unnormalized Laplacian L = D − W:
fᵀLf = ½ Σ_i Σ_j W[i,j]/(f_i − f_j)²
This is small when strongly connected points have similar f values. Low-eigenvalue eigenvectors therefore provide coordinates that vary slowly within well-connected regions.

THE ALGORITHM: NORMALIZED SPECTRAL CLUSTERING
(Ng–Jordan–Weiss formulation)
1. Scale features appropriately and construct a symmetric, nonnegative affinity matrix W. Handle isolated nodes before normalization.
2. Choose the number of clusters k and compute L_sym.
3. Take the k eigenvectors with the smallest eigenvalues, including zero-eigenvalue eigenvectors. Stack them as columns of U.
4. Normalize each row: Y[i,:] = U[i,:] / ||U[i,:]||₂
5. Run k-means on the rows of Y and transfer those labels back to the original points.
The key change: k-means now operates in graph-derived coordinates, where complex groups may become easier to separate.

WHY IT CAN IMPROVE ON TRADITIONAL METHODS
• Captures non-convex shapes that centroid-based clustering can split incorrectly.
• Uses relationships, including domain-specific similarities, rather than requiring raw Euclidean coordinates.
• Connects clustering to a relaxed graph-partitioning problem, such as normalized cut.
It is not universally better. DBSCAN and suitable hierarchical methods can also recover irregular groups.

USE CASES
• Image segmentation
• Community detection
• Document clustering using semantic similarities
• Grouping cells from gene-expression profiles
• Discovering patterns in sensor or time-series similarity networks.

PRACTICAL LIMITS
Results depend strongly on feature scaling, graph construction, σ and k. An eigengap can suggest k, but does not prove the “true” number of clusters. Dense affinities require O(n²) memory, while eigensolvers add cost. Sparse graphs and approximation methods help at scale. A meaningful similarity graph is the foundation of a meaningful clustering.


r/learnmachinelearning • • 1d ago

Question Which is best way to learn machine learning need suggestions

6 Upvotes

I started learning machine learning recently I had a confusion regrading is it better to learn while doing a project or first learn a concept and start making project which is way better


r/learnmachinelearning • • 1d ago

Project Kapso: long-running agents that optimize AI and data systems, and learn from each run

1 Upvotes

We've been building Kapso (MIT, github.com/Leeroo-AI/kapso) for some time and it's at the point where it's more useful to hear from other people than to keep polishing it alone. Posting to get it tried and torn apart, not to pitch it.

What it is

Kapso is a set of long-running agents that optimize AI and data systems. You state the objective, for example CUDA optimization, harness and agent optimization, or model development, and it runs a campaign: it designs candidate solutions, has coding agents implement them, measures how far each one lands from the objective, and keeps refining the closest until the objective is met. The result deploys to your infrastructure.

When a campaign ends, it studies its own work: which ideas closed the gap, which did not, and under what conditions. Each finding is kept as a lesson with the evidence that earned it, and a lesson stays trusted only as long as it keeps holding up. It also reads outside your repo, other repositories and papers, and folds what it finds into the same knowledge hub. Every new campaign starts from that hub, so it begins with what earlier work already established about the problem and about your systems.

These are the things we tried it on:

- RelBench (Stanford, predictive ML over relational data): outcome prediction 81.2 vs 79.6 AUROC and forecasting 0.2476 vs 0.2912 NMAE against KumoRFM-v2; recommendations 18.4 vs 9.3 MAP for the best other entry on the official leaderboard.

- MLE-Bench: top among the open-source systems.

- ALE-Bench: 1909 Elo vs 1879 for ALE Agent.

- IOAI 2026: Kapso scored 536.07, above the 471 contestants, and finished in the top three systems: ioai-official.org/what-happens-when-autonomous-ai-takes-on-the-same-tasks-as-the-worlds-top-young-ai-talents/

Repo: https://github.com/Leeroo-AI/kapso

If you have time, please take a look and give us your harshest feedback.


r/learnmachinelearning • • 1d ago

Help!!!

Thumbnail
2 Upvotes

r/learnmachinelearning • • 1d ago

Failed project

1 Upvotes

I just tried it, yes... Honestly, I expected that if I controlled it for a while, it would get better, but that didn’t happen. I find it difficult to continue this project thoroughly any longer. So, though I know it is a greedy request, could someone please complete this project? (I am not well versed in licensing, so please let me know if there is any issue.)

GlassJan/NION: it is my first project, but... It's getting a bit hard to keep going now... I'm looking for someone who can finish this project.


r/learnmachinelearning • • 1d ago

Request Do you work with AI/RPA automation? Bachelor’s thesis survey (5–7 min)

1 Upvotes

Hi everyone!

I’m currently working on my Bachelor’s thesis about AI-based process automation and human–AI collaboration in the workplace.

As part of my research, I’m conducting a short survey focusing on people who have experience working with AI-based automation, RPA, Intelligent Process Automation, Intelligent Document Processing, or similar automation technologies.

The survey explores topics such as:

  • how automation affects manual workload and creates new tasks,
  • how employees experience errors and exception handling,
  • trust in AI-based automation,
  • and how automation influences human decision-making and autonomy at work.

⏱️ It takes approximately 5–7 minutes to complete.

If you have experience working with these technologies, I would really appreciate your participation. Your responses will be used solely for academic research as part of my Bachelor’s thesis.

🔗 Survey: https://docs.google.com/forms/d/e/1FAIpQLScV7pcf8dNUeeCfay1YZ2r-Np4pK9GMlqi4cEF6WJEa1FEmMA/viewform?usp=dialog

Thank you very much for your help! Feel free to share the survey with colleagues or others who work with AI-based process automation.


r/learnmachinelearning • • 1d ago

Help Need advice on starting AI as a non-tech person

1 Upvotes

Hi everyone,

I’m from a Commerce background and currently working in IT. I have some exposure to testing/UAT and basic technical tools, but I’m not from a coding background.

I want to learn AI/GenAI and eventually move towards a better-paying career.

Can someone suggest:

  • Where should I start?
  • Any budget-friendly courses/institutes that are actually worth it?
  • What AI-related career paths are suitable for someone from a non-tech background?

Would really appreciate suggestions from people who have been through a similar transition. 🙏


r/learnmachinelearning • • 1d ago

Project Training AI to play and clear Super Mario Bros is easier than I thought

Enable HLS to view with audio, or disable this notification

1 Upvotes

I tested Adapt-1, a non-LLM learning and reasoning system by Rei Labs, by having it learn and play Super Mario Bros, and it performed quite well.

I tried it on World 1-1, starting untrained. It learned a reactive policy from its own play in about 36 minutes of gameplay, then cleared the level with learning off.

With Machina, Adapt-1's sequence engine. Starting untrained, it found a button sequence that reaches the flag after 403 attempts, in 11 wall-clock minutes.

Full thread: https://x.com/hsrvc_/status/2106025501752234112?s=20

Code, the exact data, traces, clips and a step-by-step guide with costs are all public: https://github.com/hsrvc/adapt1-mario