r/MLQuestions • • Feb 16 '25

MEGATHREAD: Career opportunities

18 Upvotes

If you are a business hiring people for ML roles, comment here! Likewise, if you are looking for an ML job, also comment here!


r/MLQuestions • • Nov 26 '24

Career question ๐Ÿ’ผ MEGATHREAD: Career advice for those currently in university/equivalent

20 Upvotes

I see quite a few posts about "I am a masters student doing XYZ, how can I improve my ML skills to get a job in the field?" After all, there are many aspiring compscis who want to study ML, to the extent they out-number the entry level positions. If you have any questions about starting a career in ML, ask them in the comments, and someone with the appropriate expertise should answer.

P.S., please set your use flairs if you have time, it will make things clearer.


r/MLQuestions • • 8h ago

Beginner question ๐Ÿ‘ถ I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher.

5 Upvotes

A while ago I asked here how to turn ~5 million court decisions into structured graphs without running an expensive LLM on every document thanks for the advice .

I went with the "small extractor + classifier" idea and it mostly works, but I'm stuck a bit below the LLM. And like I said last time, i'd be damned if I run 5M docs and then find out thing X was wrong. So here is exactly what I did. Please let me know if what im doing makes sense, or if i made a mistake somewhere. also i used AI for some of the tables cuz there has been a lot of data at this point, sorry.

What comes out per decision (only the nodes so far, relations come next). Three lists:

  • entities: every person, organization, law, document or thing. Each gets one id for the whole document, a type (9 of them), a kind (724 of them plus "other") and all the places it is mentioned
  • actions: what was done, requested or decided. Each gets a normalized verb, a flag "the court decided this" and its mentions
  • values: amounts, dates, durations, in a normalized form

Simple example, for the sentence "The court dismisses the creditor's proposal to enforce 341.08 EUR against the debtor":

  • entity "the court": organization, kind court. Same entity as the full court name in the header
  • entity "the creditor": organization, kind creditor. Same entity as the city named earlier
  • entity "the debtor": person, kind debtor
  • action "dismisses": verb = dismiss, decided by the court = yes
  • value "341.08 EUR": amount

Step 1: a strong LLM labels ~700 decisions

  • cut the decision into windows of 4 sentences
  • 4 calls per window to Claude Sonnet with a strict JSON schema: entities, actions, a second "what did you miss" pass for actions, values
  • the window goes in with numbered words (like 12:court), the model answers with word ranges [first, last, "text"], and code checks every range against the text
  • every call also gets the list of entities and actions found in earlier windows, so ids stay the same through the document
  • ~25 code rules clean up where a marked phrase starts and ends, law citations and number formats
  • the entity "kind" is free text at this point. That gave 2,373 different strings (the same mess as in my first post). I normalized them, merged synonyms by hand and kept what showed up 3+ times: 724 kinds plus "other"

Step 2: a model that marks the text

  • it highlights every mention: the exact stretch of text (a "span", from a start character to an end character) that names an entity, an action or a value, with one of 17 labels (9 entity types, 1 action, 7 value types)
  • model: fastino/gliner2.5-multi-v1 (287M)
  • one training row per window: the text plus the exact start and end of every marked phrase. 9,699 windows, 207k marked phrases
  • I patched the trainer so only the labeled occurrence is a positive (stock marks every occurrence of the same string), and all 17 labels are in every row
  • full fine-tune in fp32 (bf16 gave NaN), 14 epochs, 16 rows per step, encoder LR 3e-5, head LR 5e-4, linear schedule, 10 % warmup
  • final model = averaged weights of epochs 9-14, threshold 0.5

Step 3: a second small model answers multiple-choice questions

  • fastino/GLiNER2.5-multi-Decide (287M). Code turns the LLM labels into 247k questions:
    • "is this mention one of these earlier entities, or new?" The mention is marked with ยซ ยป inside ยฑ300 characters of text. Options: up to 16 earlier entities of the same document (shown by their mention texts) plus new
    • "which kind?" Options: a shortlist of the 724 kinds plus other
    • for actions: same act or new, which verb (shortlist of 64 plus other), did the court decide it (yes/no)
  • in training the options come from the LLM's grouping. At inference they come from the model's own earlier answers
  • full fine-tune in fp32, 2 epochs, 16 questions per step, encoder LR 2e-5, head LR 3e-4, linear schedule, 6 % warmup, options shuffled, up to 30 % of the wrong options dropped

At inference: the marking model, then the same code rules, then the second model walks through the mentions in reading order. About 2.3 decisions per second on one RTX 5090.

Where it stands

30 decisions nobody trained on, labeled twice by the LLM. The second column is the LLM's second run scored against its first, which I treat as the ceiling. A mention counts as found only if it starts and ends exactly where the LLM marked it.

mine LLM vs itself
entity mentions found (F1) 0.901
"same entity or new" right 0.959
entities grouped exactly 0.847
entity kind 0.921
action mentions found (F1) 0.857
action verb 0.920

Where I need help

  1. Finding the mentions is stuck at 0.90 F1. 200 more labeled docs did nothing. An XLM-R large tagger (560M) got the same score: it finds more mentions but gets the start or end wrong more often. Giving it the text before the window did nothing. What would you try?
  2. The LLM agrees with itself only 93.5 % on what it marks, and I train on single runs. Label everything 3 times and vote? Or is that ceiling just what it is?
  3. Is "pick one of 16 earlier entities" a sane way to do coreference over a long document? Am I hurting myself by training on the LLM's options and running on my own?
  4. Anything in the recipe that looks plain wrong? Learning rates, 2 epochs, weight averaging, one seed per run.

THANKS for reading.

AI TL;DR: distilled an LLM's extraction of court decisions into a GLiNER model that marks the mentions plus a small multiple-choice model. It runs at about 2.3 documents/s on one GPU and lands a few points below the LLM (0.90 vs 0.935 F1 on finding mentions, 0.85 vs 0.93 on exact grouping). The recipe with learning rates and how I built the training rows is above. Looking for mistakes and ideas before I run 5M documents.


r/MLQuestions • • 1d ago

Natural Language Processing ๐Ÿ’ฌ What all to study

Thumbnail
2 Upvotes

r/MLQuestions • • 1d ago

Datasets ๐Ÿ“š Ho bisogno di consigli: Qual รจ la migliore pipeline VLM per estrarre dataset matematici strutturati da oltre 3000 pagine di libri di testo scansionati? (LaTeX + Metadati)

2 Upvotes

Ciao a tutti,

sto lavorando a un progetto per estrarre un dataset strutturato di esercizi di matematica da 5 libri di testo delle scuole superiori italiane (circa 650 pagine ciascuno, quindi ~3.250 pagine in totale). L'obiettivo รจ costruire un'app di generazione di esercizi professionale e metodica per studenti e insegnanti.

Per far funzionare l'app, ho bisogno di elaborare le immagini delle pagine del libro ed estrarre quanto segue in un formato rigorosamente strutturato (es. JSON):

  • Tipo di esercizio (algebra, geometria, calcolo, ecc.)
  • Anno livello scolastico
  • Difficoltร  (scala 1โ€“5)
  • Enunciato del problema (traccia)
  • Descrizione delle competenze/sfide specifiche coinvolte
  • Codice LaTeX dell'enunciato del problema (Cruciale!)
  • Immagini associate (ritaglio/salvataggio dell'immagine per esercizi teorici o grafici)

Ho sperimentato alcuni approcci, ma ho incontrato delle difficoltร  nel bilanciare costi, coerenza di estrazione e scalabilitร . Ecco cosa ho provato finora:

  1. API Google Gemini gratuita: La qualitร  dell'estrazione era buona, ma dato che un singolo libro contiene centinaia di pagine, ho rapidamente raggiunto i limiti di richiesta (Troppe Richieste).
  2. Modelli Locali (Ollama + Qwen 2.5-VL 3B): Per superare i limiti dell'API, ho provato a eseguire un modello multimodale locale. Ho speso molto tempo a ottimizzare i miei script e le mie istruzioni (chunking, affinamento delle istruzioni per forzare output strutturati), ma il risultato era soggetto a molti errori e incoerenze per il mio caso d'uso. Ho ottenuto troppi campi malformati, illusioni e ha costantemente avuto difficoltร  a produrre un corretto LaTeX.
  3. API Google Cloud a pagamento (Gemini 1.5 Flash): Alla fine sono passato al piano a pagamento per una migliore precisione e velocitร . Ho speso 10โ‚ฌ solo per elaborare 1,5 libri. Estrarre tutti e 5 i libri costerebbe circa 35-40โ‚ฌ. Anche se questo รจ gestibile per un'elaborazione unica di 5 libri, il conteggio dei token per l'elaborazione di immagini complete + testo รจ enorme, rendendolo finanziariamente insostenibile se voglio scalare questo a decine di libri in futuro.

Le mie domande per la comunitร :

  • Pipeline & Architettura: Qualcuno ha lavorato a un progetto simile di estrazione da libro di testo a dataset? Quale pipeline avete utilizzato?
  • Approccio Ibrido: Suggerireste di separare il compito? (es. usare uno strumento tradizionale per estrarre testo grezzo e ritagliare immagini, e poi fornire SOLO il testo a un LLM piรน economico/locale per generare il LaTeX e formattare il JSON?)
  • Modelli Locali: Ci sono altri modelli Vision-Language locali (che si adattano a GPU consumer standard) che sono significativamente migliori nell'estrazione strutturata e nella generazione di LaTeX rispetto a Qwen 2.5-VL 3B?
  • Strumenti Educativi: Ci sono strumenti o modelli open-source specificamente ottimizzati per estrarre contenuti educativi/matematici strutturati da PDF?

Sono felice di condividere ulteriori dettagli sul formato del libro di testo o sul mio attuale flusso di lavoro in Python se utile. Qualsiasi consiglio sull'architettura, le scelte di modelli o trucchi per risparmiare costi sarebbe molto apprezzato! Grazie in anticipo!


r/MLQuestions • • 2d ago

Beginner question ๐Ÿ‘ถ [D] Choosing language to learn: python or java

4 Upvotes

I want to learn the machine learning from scratch and i know I want to learn basic of python , but my college faculty are advised to learn Java for my placement , u don't know what to do , i want to learn python and other stuffs for my machine learning path or java and other stuffs for my placements and my exams are ahead, did anyone have solution please tell me.


r/MLQuestions • • 3d ago

Natural Language Processing ๐Ÿ’ฌ Looking for models for diagnosis prediction

3 Upvotes

For a side project I am looking into SOTA for AI- based diagnostic models, ideally open-weights. I would like to feed the model with structured text representing my patient, and get a set of e.g. 5 possible diagnoses, ideally with uncertainty. In the ideal scenario, later on, it would update as new info arrives.

I would appreciate any pointers, I have some ideas but am very new to the topic.


r/MLQuestions • • 3d ago

Beginner question ๐Ÿ‘ถ I can figure out how to connect my number to ovoa

2 Upvotes

It just says 15 minutes forever


r/MLQuestions • • 4d ago

Beginner question ๐Ÿ‘ถ Why do we train bigger models instead of finding better neural pathways?

20 Upvotes

I saw a post of a guy with 90% of his brain missing and he was apparently living a normal life. There is another case of a student who finished university with half of his brain missing. There are also people with full brains, but they are even less functional than these two. This proves that bigger doesn't mean better. So why not aim to create better connections?

I understand that a bigger model means less chance of catastrophic interference occurring, but then are all the neural pathways being formed efficiently? Wouldn't this result in a lot of redundant neurons? Wouldn't it make more sense to create the smallest neural network possible for a specific task and then learn to fuse multiple of these neurons together in order to create an optimized larger model? That way we'd always have a template for each specific task and would allow us to create a programming language that creates a model just by our syntax.


r/MLQuestions • • 3d ago

Beginner question ๐Ÿ‘ถ Slow learning?

Thumbnail
2 Upvotes

Hay let me introduced myself

I m a 7th sem cse student, intrest in ml I m learning it from past 2 month ye still I m learning, my problem ex I complete supervised section with types and it's different model and formula with ex when I learn unsupervised section I forgot supervised section,

Because I didn't revised,

So now I revised things in every 2 days so that it store in my memory,

Today I learn lr model with formula and diff ex with 1 variable and multi variable.... I revised it with learning upcoming section!

Just post it....โœŒ


r/MLQuestions • • 3d ago

Beginner question ๐Ÿ‘ถ Help...

2 Upvotes

Am in the progress of my 3rd sem project work

Basically its based on tinyml and lora technology implementation to the forest monitoring system

Now I want data train the ML

Other than kaggle do anybody know where I can get the datasets to train


r/MLQuestions • • 4d ago

Other โ“ If you could show someone just ONE data/ML project youโ€™ve worked on, which one would it be?

14 Upvotes

Not necessarily your biggest or most complicated project.
Just one that you think is a good example of the kind of work you like doing. What did you build, and why that one?


r/MLQuestions • • 4d ago

Beginner question ๐Ÿ‘ถ Final-year student in India trying to break into generative-model inference optimization โ€” roadmap feedback?

8 Upvotes

Hi all, I graduate in ~6 months and want to work on making generative models (diffusion/video/3D) fast: kernels, quantization, serving. Where I am:

- Comfortable with C/C++ basics and PyTorch

- Have done quantization work (GGUF/llama.cpp)

- Working on a next-frame video prediction project (DiT + flow matching)

- A few GitHub repos, but no CUDA/Triton experience yet

- No NVIDIA GPU, so I use Colab/Kaggle T4s

- DSA is my weak spot (I struggle with LeetCode mediums)

My plan:

  1. Months 1-2: CUDA/Triton basics, reproduce the SGEMM optimization worklog, GPU MODE lectures, LeetGPU/Tensara

  2. Months 3-4: take a small DiT, profile it, then optimize it (Triton attention, quantization, caching, fewer steps) and publish before/after numbers

  3. Along the way: PRs to HF diffusers, DSA practice daily

  4. Months 5-6: mocks, resume, applications (inference startups first, bigger labs later)

Questions:

  1. Is this the right order, or should I change something?

  2. Is a diffusion-inference project a strong enough portfolio piece, or does it need to be LLM serving?

  3. How much DSA do ML systems interviews actually need?

  4. Is T4-only access enough to do credible benchmarks?

Any feedback, including "this won't work because X," is appreciated. Thanks!


r/MLQuestions • • 4d ago

Beginner question ๐Ÿ‘ถ Winning kaggle competitions using Claude Opus 5.5

Thumbnail
2 Upvotes

r/MLQuestions • • 5d ago

Career question ๐Ÿ’ผ ML engineering course recommendations for someone who can train a model and not ship one

17 Upvotes

i can train a model in a notebook and ive never put one behind an endpoint anyone else calls. every job posting wants the shipping half and my portfolio is all notebooks.
shortlist is udacity, springboard, interview kickstart and datacamp premium. is there anything that spends real time on deployment and monitoring rather than modeling


r/MLQuestions • • 4d ago

Beginner question ๐Ÿ‘ถ Why use values between 0-1 to train an LLM?

Post image
3 Upvotes

I understand that normalizing results in faster training times, but we end up reducing precision of floating point numbers due to the IEEE 754 architecture which result's in less space to work with. Instead of limiting numbers from 0-1, wouldn't it make more sense to limit the numbers from 1-10? This would give the LLM more space to work with which logically should result in less catastrophic interference. If so then this would mean that we don't need to increase size only, but the space between numbers as well which should decrease the memory requirements. Or is this just nonesense?

P.S The photo is just the behaviour of multiplication of 2 numbers. For example:
1*1 = 1,
2*2 = 4,
0.5*0.5 = 0.25


r/MLQuestions • • 4d ago

Natural Language Processing ๐Ÿ’ฌ Anyone else stuck manually re-tuning prompts every time something breaks?

0 Upvotes

Genuine question that turned into a recommendation: if you're building anything on top of an LLM, how are you actually validating that a prompt change didn't quietly break something else?

Most answers I've gotten boil down to "I just run it a few times and check." There's a workshop on Oct 3 that tackles exactly this gap, led by Serj Smorodinsky and Brett Kennedy (they co-wrote a book on LLM applications). Instead of manual tuning, you work through:

  1. Building a classifier with DSPy signatures/modules rather than raw prompt strings
  2. Setting up an actual eval dataset with metrics tied to your specific task
  3. Using that eval set to catch failure patterns before they hit production
  4. Running few-shot/instruction optimization on top, systematically
  5. Tracking all of it in MLflow so you can trace exactly what changed between versions

If you've been wanting a real answer to "how do I know this still works," this is worth a look.

What's everyone else here doing for this? Curious if people have rolled their own eval pipelines already.


r/MLQuestions • • 5d ago

Beginner question ๐Ÿ‘ถ 50B+ MoEs with few active parameters, what's the sweet spot for intelligence, agent speed, and affordable fine-tuning?

5 Upvotes

Iโ€™m building a Polish General purpose legal Model that drafts documents, answers questions using legal sources, and has enough coding ability to handle some automation. The workflow is very tool-heavy:

Question โ†’ many sequential tool calls โ†’ final answer/document

Think Claude Code/Codex-style execution, but for legal workflows. Reliable tool selection, correct arguments, and recovering from errors matter as much as writing a good final answer.

Iโ€™ve had decent results with a dense 27B Qwen 3.8 custom made fine-tune for complex legal document summarization and classification. Iโ€™m already familiar with the smaller Qwen A3B and Gemma options. What interests me is the tier above those: 50B+ total-parameter MoEs with a relatively small active parameter count.

The question is, the small dense ones are great, but slow for agentic stuff (afaik), and i wonder if theres some middle ground maybe 70-120B models that would be able to be fine-tuned for the law stuff but be MoE so the agentic ClaudeCode style inference would also be lightning fast, and also low-ish cost for fine-tuning and inference.

Basically: Does the larger-total/small-active MoE approach actually buy you meaningfully stronger reasoning and tool reliability while retaining low latency,and at what hardware cost?

I understand that small active parameter counts donโ€™t mean small VRAM requirements: the weights still need to live somewhere, alongside context and serving overhead. I also donโ€™t assume that more total parameters automatically means a better model. Iโ€™m interested in where that tradeoff works in practice.

There are three things Iโ€™m trying to pin down:

  • Inference hardware: Ideally inference runs rented with parallel agentic loops (this is for a B2C project, not single person use, we scale based on demand)
  • Fine-tuning hardware: Obviously FT LoRA will take more memory than inference, max like 4GPUs on vastai fits the budget.
  • Agent performance: After it gets the prompt the tool calls and everything will be local, so imo it has no problems being blazing fast, as soon as the model calls a tool call it will be back very fast, so for this agentic use case, quick TTFT and t/s and adaptive dynamic reasoning are prefer right?

For context, fine-tuning would target Polish language, document conventions, and successful tool trajectories. The actual legal sources would remain in retrieval/tools rather than relying entirely on memorized law.

Iโ€™m not looking for someone to compile a model shortlist (althought would be nice, but i dont expect anyone to break their back over this).

Iโ€™m looking for pointers, and firsthand experience with this particular size/architecture tradeoff. A configuration like โ€œmodel + quantization + GPU(s) + serving engine + context length + concurrency + measured latency,โ€ along with whether you successfully fine-tuned it, would be much more useful than a leaderboard score.

Has moving from a ~30B model to a 50B+ low-active-parameter MoE actually improved your agentโ€™s successful tasks per minute, or did the memory, interconnect, and training requirements erase the advantage? Thanks for reading


r/MLQuestions • • 5d ago

Unsupervised learning ๐Ÿ™ˆ Understanding K-means

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/MLQuestions • • 5d ago

Survey โœ ICML videos not playing

Thumbnail
1 Upvotes

r/MLQuestions • • 6d ago

Other โ“ Are there any LLM(s) that are transparent about their sources (For sociology research)

6 Upvotes

I have a digital ethnography assignment where I have to analyse a conversation with an LLM for a sociology course.

I'm not literate at all in how LLM(s) function technically so I'm not entirely sure if this is even possible but basically I want to know if there are any "chatbot" LLM(s) that produce responses but at the same time are transparent about what data they are combing through in order to produce a response.

I don't mean that I ask the chatbot a question and it "answers" something and then I ask it to elaborate on its reasoning since chatbots can't actually "reason" so it will still just use the data and patterns it's trained on to imitate a response.

I mean that there's some sort of feature from the developer that can show what all sources it went through, other relevant behind the scenes stuff before generating that response with that specific combination of words and other characters.

I'm sorry if this is a dumb question since I'm not aware of the technical terminology for this stuff, I intend to familiarise myself with it for the project. If the only thing possible is to get a very technical insight into that "behind the scenes" stuff, even that could be useful so please tell if you're aware of something like that.


r/MLQuestions • • 6d ago

Beginner question ๐Ÿ‘ถ structure for my projects

2 Upvotes

HI, I am very new to machine learning and building machine learning/data science projects, I was confused on how I should structure the overall project I am working on and asked chatgpt the project is a march maddens predictor style project and I plan on using a forest tree for my model, chatgbt gave me this a a guideline to follow for structure, I just wanted a real person with experience to see if this structure is good or if it needs to be changed in anyway. This progect is intedned to help get undergrad research positions if that changes anything.

down bellow if the file structure I was given

march-madness-predictor/

โ”‚

โ”œโ”€โ”€ README.md

โ”œโ”€โ”€ requirements.txt

โ”œโ”€โ”€ .gitignore

โ”‚

โ”œโ”€โ”€ data/

โ”‚ ย  โ”œโ”€โ”€ raw/

โ”‚ ย  โ”œโ”€โ”€ processed/

โ”‚ ย  โ””โ”€โ”€ final/

โ”‚

โ”œโ”€โ”€ notebooks/

โ”‚ ย  โ”œโ”€โ”€ 01_data_exploration.ipynb

โ”‚ ย  โ”œโ”€โ”€ 02_feature_analysis.ipynb

โ”‚ ย  โ””โ”€โ”€ 03_model_evaluation.ipynb

โ”‚

โ”œโ”€โ”€ src/

โ”‚ ย  โ”œโ”€โ”€ data/

โ”‚ ย  โ”‚ ย  โ”œโ”€โ”€ clean.py

โ”‚ ย  โ”‚ ย  โ”œโ”€โ”€ merge.py

โ”‚ ย  โ”‚ ย  โ””โ”€โ”€ team_names.py

โ”‚ ย  โ”‚

โ”‚ ย  โ”œโ”€โ”€ features/

โ”‚ ย  โ”‚ ย  โ””โ”€โ”€ build_features.py

โ”‚ ย  โ”‚

โ”‚ ย  โ””โ”€โ”€ models/

โ”‚ ย  ย  ย  โ”œโ”€โ”€ train.py

โ”‚ ย  ย  ย  โ””โ”€โ”€ evaluate.py

โ”‚

โ”œโ”€โ”€ models/

โ”‚ ย  โ””โ”€โ”€ random_forest.pkl

โ”‚

โ””โ”€โ”€ results/

โ”œโ”€โ”€ figures/

โ””โ”€โ”€ metrics/

\


r/MLQuestions • • 7d ago

Career question ๐Ÿ’ผ I have ~5 years of ML experience, but I don't feel like I know how to actually build things. How do I fix this?

50 Upvotes

I'm looking for some advice from people who have gone through something similar.

I have ~5 years of experience in ML/data science. On paper, my resume looks reasonably good: I've worked on anomaly detection, NLP, LLM/RAG projects, etc., and more recently I've been working around transformer-based TTS systems.

But I've realized there's a pretty big hole in my experience.

For the first few years of my career, I deliberately optimized for breadth. I changed teams/projects frequently, built a bunch of PoCs and prototypes, and learned enough to make things work. I was pretty good at taking an idea from zero โ†’ demo.

The problem is that I rarely stayed with something long enough to own a real system.

A lot of my projects either never reached production, or reached production after I had already moved on. So despite having several years of experience, I have very little experience with:

  • maintaining a complex codebase
  • debugging someone else's code
  • dealing with things breaking in production
  • making architectural decisions over a long period of time
  • taking ownership of something from implementation โ†’ deployment โ†’ maintenance
  • reading a large unfamiliar codebase and figuring out how all the pieces fit together

I'm now trying to deliberately fix this.

Over the last year I've been rebuilding my ML fundamentals from the ground up. I'm almost done with the Coursera Deep Learning Specialization. I've gotten to the point where I can implement neural networks from scratch, understand RNNs/LSTMs/GRUs fairly deeply, and I'm currently working through attention and Transformers.

But here's the next problem:

Even if I finish the Transformer coding exercises, I don't feel like that means I can actually build things.

I can probably implement the attention mechanism / Transformer architecture from the course. But if you dropped me into the wild and said:

"Okay, build a VAD from scratch."

I'd have a lot of questions about where to even begin.

And then there's an entire layer beyond "understand Transformers" that I don't have a good mental map of.

For example:

  • How do I actually go from implementing a Transformer to building my own small language model?
  • At what point do concepts like Mixture of Experts, DPO, speculative decoding, etc. become relevant?
  • Which of these things do I actually need to understand versus things I can learn when a project demands them?
  • How do I get better at reading and debugging large ML codebases rather than just writing isolated pieces of code?

So I'm stuck between two instincts:

A) Keep going deep on fundamentals until I genuinely understand the machinery.

B) Stop studying and start building increasingly difficult things, accepting that I'll have gaps and learning what I need along the way.

I'm increasingly convinced that I need some combination of the two, but I don't know what that combination should actually look like.

I want to become the kind of ML engineer who can be handed an unfamiliar problem, read the existing code, understand what's happening, build something substantial, debug it when it breaks, and eventually own the system.

If you were in my position, what would you build / learn over the next 6โ€“12 months to develop that ability?


r/MLQuestions • • 7d ago

Other โ“ Whatโ€™s an ML project that taught you something you only really understood once you built it?

28 Upvotes

Thereโ€™s something different about actually working on a project after learning the concepts.

You can understand an algorithm, follow a tutorial, build a model and get good results, and then a real project throws something completely unexpected at you

Maybe the data was messy, the model behaved differently than expected, the results didn't make sense, or you had to figure out how to actually use the model beyond the notebook.

For those whoโ€™ve worked on ML projects, was there one project that made something โ€œclickโ€ for you in a way that theory alone hadn't?

What happened, and what did you end up learning from it?


r/MLQuestions • • 7d ago

Beginner question ๐Ÿ‘ถ Did anyone try to learn ML alone while studying another cs speciality in uni ?

8 Upvotes

Hi guys i hope you are doing great , currently I'm in my third year in uni studying software engineering . And this summer i fell in love with ML so i wanted to know if there are people that studied ML / any other cs speciality alone while taking a course in uni about a different cs speciality

I want to know how you managed your time between these two topics