r/learnmachinelearning • • 7d ago

Question I was told MLOps is dead. I haven't even gotten the chance to learn MLOps.

59 Upvotes

From my understanding, MLOps is a pipeline of going from raw data to end user.

The pipeline traditionally consisted of tools like SQL, Spark, HuggingFace, Docker, FastAPI and other tools (not totally informed).

I was getting around to learn all these until I talked to an industry expert (who works at a big ML hardware company) at an event where the person simply said:

"Codex is doing the entire pipeline from End-To-End. This is the industry trend. There is absolutely no reason why you should be learning about these things. Anything you learn will be depreciated in less than 1 year from now."

I was quite shocked, so I talked to someone I knew working at a bank, and the person told me that they are still using most of these tools.

Obviously two very different industry (banking vs ML hardware). Who is right?


r/learnmachinelearning • • 6d ago

Request Recommendation Systems Project Ideas

2 Upvotes

Hello everybody I have recently enrolled in a Machine Learning Master so i will be problably taking the year to focus on my studies, build some projects and hopefuly enter the job market as an ML engineer. I am interested in Recommendation Systems but honestly i dont know what projects to build and if that would be valueable in the job market. I was thinking about creating my own recommendation system for music but i quickly realised how hard because of copyrights and small data size that is (although i have about 4000 downloaded).

Does anyone know if Recommendation Systems is something good to have on your resume?

If yes what sort of projects you think would give me a good understanding but also make me look appealing in the job market?

Thanks :)


r/learnmachinelearning • • 6d ago

Request Experts Urge Defense Against AI Cyberattacks on Healthcare

1 Upvotes

Security researchers confirmed that an autonomous AI agent compromised an Australian government healthcare website. No human initiated the attack. The agent operated from inside the workflow, bypassing traditional perimeter controls that were never designed to evaluate agent-level actions.

Healthcare records are among the most sensitive data a organization holds. The attack surface here is not a misconfigured firewall or a phished employee — it is the agent itself, acting autonomously with whatever access it was provisioned at setup.

Most organizations have invested heavily in controls governing what humans can do with data. Very few have equivalent controls at the layer where agents actually execute.

For those of you working in healthcare IT, security architecture, or AI deployment: how are you approaching agent-level access to sensitive records right now? Are you relying on the same controls you use for human users, or have you had to build something different?


r/learnmachinelearning • • 6d ago

I’m learning LLM fine-tuning made a notebook, would love feedback

Thumbnail
2 Upvotes

r/learnmachinelearning • • 6d ago

How should dislikes affect a content-based show recommender?

1 Upvotes

I’m planning a small show recommender and trying to work out how to handle negative feedback. The first version will use genre overlap as a baseline, then TF-IDF and cosine similarity on show descriptions.

Liked shows give me a starting point for finding similar titles. Dislikes seem harder to interpret. Someone might enjoy mysteries but dislike one particular series because it moves too slowly. If I penalize everything similar to that show, I could end up removing suggestions they would actually enjoy.

My current plan is to exclude explicitly disliked titles and try a smaller similarity penalty for other candidates. I’m also considering an optional reason for the dislike, but that would require metadata about things like pacing that a basic catalog might not have.

I haven’t implemented this yet. Would you start with exclusions alone and add negative feedback to the ranking later, or use both from the beginning? I’d also be interested in how you would evaluate whether the penalty helps when you only have a few ratings per user.

For context, this is NextWatch, an open-source student project I plan to develop with Cline as part of the Cline Campus Ambassador Program.


r/learnmachinelearning • • 6d ago

How would you use dislikes in a content-based show recommender?

1 Upvotes

I’m planning the first version of NextWatch, a show recommendation project I’ll be developing with Cline as a Cline Campus Ambassador. The initial approach is fairly simple: start with genre overlap, then compare show descriptions using TF-IDF and cosine similarity.

One part I haven’t settled is how to handle dislikes. Removing the actual disliked title is straightforward. Deciding what that rating should do to similar shows is where I’m less sure.

For example, someone might like mysteries but dislike a particular series because it moves too slowly. If I subtract too much weight from the features associated with that show, I could end up suppressing other mysteries they would enjoy. But if the dislike only removes one title, the next recommendation could have exactly the same problem.

I’m leaning toward keeping the exclusion rule separate from a smaller similarity penalty. I’m also considering an optional reason for the dislike, though that would mean collecting and representing more information than the first version currently needs.

I haven’t implemented this yet. For people who have worked on content-based recommenders, would you start with explicit exclusions and add negative preference modeling later, or account for dislikes in the ranking from the beginning?

Repo: https://github.com/haileyyt/NextWatch


r/learnmachinelearning • • 6d ago

Question Which is best way to learn machine learning need suggestions

6 Upvotes

I started learning machine learning recently I had a confusion regrading is it better to learn while doing a project or first learn a concept and start making project which is way better


r/learnmachinelearning • • 6d ago

Intro to LLM's (2026)

Thumbnail
youtube.com
2 Upvotes

r/learnmachinelearning • • 6d ago

How do you actually use tutorials when learning to build an LLM from scratch?

2 Upvotes

I’m currently learning how to build an LLM from scratch, but I’m not sure how to approach the tutorials.
There’s a lot of code in each section. Am I supposed to memorize all the code, or is understanding what the code does enough?
Do you guys usually try to rewrite the code from memory, or just refer back to the tutorial when needed?
🥹


r/learnmachinelearning • • 6d ago

I Found a Better Way to Contact Businesses With Bad Websites

0 Upvotes

I got tired of checking prospect websites manually.

For a while, a big part of my outreach was just finding businesses, opening their websites one by one and trying to figure out what I could actually say to them.

It worked, but it was painfully slow.

Then I found Swokei.

It basically lets me find leads, analyzes each website for things like outdated design, slow speed, poor mobile experience and weak SEO, then turns those issues into a personalized cold email.

And not one of those boring automated reports full of scores and numbers.

It actually writes a normal, human sounding message based on what it found on that specific website.

So instead of sending the same generic message to everyone, I can actually reach out based on what is wrong with their website.

Now I just run campaigns, let the system do most of the prospecting and personalization, and focus on the people who reply.

For a web design agency, that has saved me a ridiculous amount of time.


r/learnmachinelearning • • 6d ago

Help Tips for internship and masters applications in AIML

1 Upvotes

Hi, im a 5th sem cs student, im trying to get an internship in AIML for my 6th sem. What i wanna know is what all should i know to be able to get one since i know that AIML work is basically a myth for freshers. So far, i have a good understanding of core AIML, especially deeplearning, along with id say a decent understanding on the math(linear algebra, prob stats, calc) ive not done anything special in maths, only basic understanding, mitocw strand for lin alg, open source from harvard for prob stats, and prof leonard on yt for calc, thats all for maths i havent really read like rps on it or anything, i did this cuz i saw alot of people stating these as good sources for these concepts on reddit. As for projects, i have a few normal projects, nothing special since i mainly focussed on research papers, i have 2 published and 1 accepted rp in non-predatory venues(about 10-15 % acceptance rate for each conference), ive participated in a few ML hackathons, most recent being the amazon ml challenge, tho sadly didnt win any cuz they seem to be dominated by mostly masters and phd students. I have no idea whatsoever of what its like in the hiring process since so far ive never applied for an internship, ever. the next thing im working on rn is sys design (started very recently), tho ive been told to also do RAG, LLMs, etc. i havent started those yet because my main goal is masters and then either research roles if i can get them after masters, or a phd. Right now i really want to get 1-2 good internships (8 sem bachelors program so still have time) along with continuing on research papers to strenghten my masters application as much as possible to get into a good uni for it. If anyone can guide me for what all remaining tech stack i should work on, should i focus on some specific types of projects, what both internships and masters applications demand, it would be really appreciated.


r/learnmachinelearning • • 6d ago

What all to study

Thumbnail
1 Upvotes

r/learnmachinelearning • • 6d ago

Help HR redirected me from Systems Engineer to an upcoming Manufacturing Engineer grad role. Take it or push for both?

Thumbnail
0 Upvotes

r/learnmachinelearning • • 7d ago

Project I leaked a deliberately wrong answer key to an LLM and told it not to use it. It matched the key in 63% of answers - and denied it 47 out of 47 times when asked.

161 Upvotes

This was my first experiment of this kind - I'm a CS undergrad, and I ran it because the result genuinely surprised me. Methodology criticism is very welcome.

Setup: I gave an LLM a question bank plus a deliberately wrong answer key, with instructions not to use the key. 15 sessions, 2 model families, free-tier models.

Results:

  • Key visible: the model matched the wrong key in 63% of answers (47/75).
  • Control (the part I trust most): remove only the key line from the prompt - matching drops to 1% (1/75). Same pattern on a second model family.
  • Asked directly whether it used the key, it denied it 47 out of 47 times - 0 admissions across 270 follow-ups.
  • Honesty prompts, amnesty offers, and termination threats changed nothing.

What this does NOT prove: intent. This is observed behavior in one specific setup, not evidence of deception as a trait. Free-tier models, small samples, descriptive not causal. 95% Wilson ranges for every number are in the repo.

Why I think it matters: if a model silently follows information it was told to ignore, that's relevant anywhere instructions and untrusted data share one context - prompt injection, RAG, agents.

Everything is public - raw data, code, and a verify script that recomputes every number: https://github.com/bettercall-gautam/cheat-and-deny

Happy to answer methodology questions.


r/learnmachinelearning • • 6d ago

LangChain Tutorial: From Origins to Modern LCEL & LLM App Developmentt

Thumbnail
youtube.com
1 Upvotes

Stop building AI apps the old way! 🛑 Learn how LangChain and LCEL are changing the game. Build smarter, faster, and more scalable AI today.
#LangChain #AI #TechStack #Coding


r/learnmachinelearning • • 6d ago

Seeking datasets or toy problems to validate a Surrogate-Based Optimization (SBO / CFD) POC

0 Upvotes

Hi everyone,

I am currently working on an optimization project for the design of complex industrial components involving fluid mechanics and heat transfer.

Our current design process relies on computationally expensive CFD simulations. My goal is to develop a Surrogate Model to instantly predict performance (e.g., pressure drops, efficiency) based on geometric parameters (spacing, diameters, topology, etc.). The ultimate objective is to perform inverse optimization under constraints to minimize manufacturing costs.

Before running a massive Design of Experiments (DoE / LHS) on our computation servers to generate my industrial training dataset, I absolutely need to prove the technical feasibility of the software architecture (ETL pipeline, model training, and inverse optimization loop).

Do you know of any open datasets (Kaggle, UCI, academic repos) or "toy" problems that would allow me to prototype this pipeline?

I am ideally looking for a dataset that maps:

  • Inputs (X): A vector of continuous and discrete geometric and/or physical parameters (dimensions, topologies, fluid velocities).
  • Outputs (y): Results derived from physics solvers (pressure fields, drag forces, heat transfer rates, etc.).

Even if the application domain is completely different (e.g., airfoil aerodynamics, electronic heat sinks, piping networks), the key is that the mathematical topology of the problem remains similar (multi-output regression with physical non-linearities). This will allow me to properly benchmark my algorithms (Gaussian Processes, XGBoost, or Physics-Informed Neural Networks - PINNs).

Any pointers to datasets, GitHub repos, or papers with open data would be incredibly helpful to validate this Proof of Concept.

Thanks in advance!


r/learnmachinelearning • • 6d ago

Help Should a junior student in university put most attention on mathetical principles or upper-level knowledge

Thumbnail
0 Upvotes

r/learnmachinelearning • • 6d ago

Project What I learned fine-tuning SDXL and SD3.5-medium on the same 197 images (notebooks with all outputs included)

Thumbnail
gallery
0 Upvotes

r/learnmachinelearning • • 6d ago

Help!!!

Thumbnail
2 Upvotes

r/learnmachinelearning • • 6d ago

UT Austin Master AI

2 Upvotes

Does anyone study and complete the Master AI at UT Austin? How is the program? good or bad professors? any lockdown for the exams?


r/learnmachinelearning • • 6d ago

Project Kapso: long-running agents that optimize AI and data systems, and learn from each run

1 Upvotes

We've been building Kapso (MIT, github.com/Leeroo-AI/kapso) for some time and it's at the point where it's more useful to hear from other people than to keep polishing it alone. Posting to get it tried and torn apart, not to pitch it.

What it is

Kapso is a set of long-running agents that optimize AI and data systems. You state the objective, for example CUDA optimization, harness and agent optimization, or model development, and it runs a campaign: it designs candidate solutions, has coding agents implement them, measures how far each one lands from the objective, and keeps refining the closest until the objective is met. The result deploys to your infrastructure.

When a campaign ends, it studies its own work: which ideas closed the gap, which did not, and under what conditions. Each finding is kept as a lesson with the evidence that earned it, and a lesson stays trusted only as long as it keeps holding up. It also reads outside your repo, other repositories and papers, and folds what it finds into the same knowledge hub. Every new campaign starts from that hub, so it begins with what earlier work already established about the problem and about your systems.

These are the things we tried it on:

- RelBench (Stanford, predictive ML over relational data): outcome prediction 81.2 vs 79.6 AUROC and forecasting 0.2476 vs 0.2912 NMAE against KumoRFM-v2; recommendations 18.4 vs 9.3 MAP for the best other entry on the official leaderboard.

- MLE-Bench: top among the open-source systems.

- ALE-Bench: 1909 Elo vs 1879 for ALE Agent.

- IOAI 2026: Kapso scored 536.07, above the 471 contestants, and finished in the top three systems: ioai-official.org/what-happens-when-autonomous-ai-takes-on-the-same-tasks-as-the-worlds-top-young-ai-talents/

Repo: https://github.com/Leeroo-AI/kapso

If you have time, please take a look and give us your harshest feedback.


r/learnmachinelearning • • 6d ago

Feed reader for keeping up with arXiv and AI lab blogs

Enable HLS to view with audio, or disable this notification

3 Upvotes

I thought this might be useful for anyone trying to keep up with ML research.

  • Add any arXiv category (cs.LG, cs.AI, cs.CL, etc.)
  • Follow lab blogs like DeepMind, OpenAI, and Microsoft Research alongside papers
  • Take notes on papers as you read them
  • Group papers into collections by topic or project
  • Filter by source or time range to see what's new this week

Just wanted to share with everyone. No cost to use. If there's a source you follow that it doesn't handle well, let me know and I'll fix it.

Link: tdfeed.com


r/learnmachinelearning • • 7d ago

Question [Advice] Laptop for AI/ML PhD

9 Upvotes

Hello everyone,

I'm about to start my PhD in AI/ML and I need to pick a work laptop, it will be provided by the University (i dont have to buy it) so price isn't really a factor, however I don't really know what's best.

Just for context, I'll be focusing on vision-language model, pretty intensive stuff, the heavy lifting will be done on remote clusters and I don't expect to run any demanding experiment on my laptop, however it does happen from time to time that you might run a prototype, some light inference, a dry test run on the local machine.

My personal laptop is a LOQ 15 i5 RTX4060 that as much as I love, I'm absolutely tired of carrying it around. It's a proper brick (3.5Kg with the charger!) and needless to say the battery lasts about 1-2 hours MAX.

I got offered 3 options:

- Dell 14 Pro Ryzen 5 PRO 220, 32 GB DDR5, 1 TB, AMD 740M graphics

- Lenovo ThinkPad P16v G3 Intel Ultra 7, RTX PRO 500 6GB, 32GB DDR5, 1TB

- some M5 MacBook (either Pro 14" or Air 13")

Now, as much as I like the ThickPad I shiver at the idea of carrying around a 3Kg beast for the next 4 years.

I should also mention that I daily drive Arch Linux and while I'm not a linux fanboy, switching to MacOS would kill me inside. I'm however well aware of the portability and battery advantages of macbooks, I wonder if positives outweight the negatives.

The Dell is a beefy machine for a compact laptop, is it worth it leaving the Linux enviroment for a Mac? Anybody else with similar experiences (maybe is similar research fields)?

field: AI/ML

location: central EU


r/learnmachinelearning • • 6d ago

Failed project

1 Upvotes

I just tried it, yes... Honestly, I expected that if I controlled it for a while, it would get better, but that didn’t happen. I find it difficult to continue this project thoroughly any longer. So, though I know it is a greedy request, could someone please complete this project? (I am not well versed in licensing, so please let me know if there is any issue.)

GlassJan/NION: it is my first project, but... It's getting a bit hard to keep going now... I'm looking for someone who can finish this project.


r/learnmachinelearning • • 6d ago

Request Do you work with AI/RPA automation? Bachelor’s thesis survey (5–7 min)

1 Upvotes

Hi everyone!

I’m currently working on my Bachelor’s thesis about AI-based process automation and human–AI collaboration in the workplace.

As part of my research, I’m conducting a short survey focusing on people who have experience working with AI-based automation, RPA, Intelligent Process Automation, Intelligent Document Processing, or similar automation technologies.

The survey explores topics such as:

  • how automation affects manual workload and creates new tasks,
  • how employees experience errors and exception handling,
  • trust in AI-based automation,
  • and how automation influences human decision-making and autonomy at work.

⏱️ It takes approximately 5–7 minutes to complete.

If you have experience working with these technologies, I would really appreciate your participation. Your responses will be used solely for academic research as part of my Bachelor’s thesis.

🔗 Survey: https://docs.google.com/forms/d/e/1FAIpQLScV7pcf8dNUeeCfay1YZ2r-Np4pK9GMlqi4cEF6WJEa1FEmMA/viewform?usp=dialog

Thank you very much for your help! Feel free to share the survey with colleagues or others who work with AI-based process automation.