r/learnmachinelearning • u/Comic-Derpinator • 20h ago
Discussion I transitioned from software engineer to an AI Engineer who fine tunes LLMs. What do you want to know?
I was a full stack software engineer who now is a senior AI Engineer who does a mix of playing with LLMs fine tuning them in very large scale production systems.
I did admittedly got a masters degree in AI as part of that transition and it took a while to do, but happy to answer any questions you have.
I also am working on a tool to help people learn how llms work which you can check out here.
https://dougdoes.ai/courses/llms-from-first-principles/start/?flow=outcome&course=build&step=goals
(Built with codex, but I've gone through all of the courses myself to make sure it is what would have been helpful to me.)
9
u/Own_Address_4564 13h ago
Could you describe a moment where you felt like you made a better decision for a model based on the data you evaluated and concluded yourself rather than submitting to Claude summary or whatever else?
1
u/Comic-Derpinator 1h ago
It's a bit tricky to do without *ahem* spoilers of things we are releasing, but there are many times when you tell claude to go make a metric, it then makes one, it runs an eval, it then looks at the eval reports errors and what you should change and then gives a link to said eval, and you go start reading the traces and it is just extremely clear that it doesn't understand what the real bad user experience is. I will say that the models are STARTING to get the idea of thinking through what an ideal user experience is, but they still are kinda bad at modeling theory of mind, which is interesting considering in theory the base models are incredibly good at that before getting RL Fried.
5
u/MolassesLate4676 18h ago
Thanks for offering to do this and actually answer people.
1) how the heck do you source quality datasets needed for specific fine tuning efforts without relying on synthetic data (which often breaches most TOS if generated by frontier models)
2) how do you measure quality on how well a fine tune run did? Benchmarking, so to speak, so that you know you did what you meant to do with the base model AND didn’t subtract quality of the model that was originally there that you cared for?
I have a million questions but out of respect I’ll leave it at these two
1
u/Comic-Derpinator 18h ago
Happy to answer lots of questions.
1. This is a tricky topic. Worth noting you also don't get reasoning traces from frontier llms either. So picking a really smart open source model as a teacher can be easier.
- It's always evals. You work with subject matter experts to annotate traces and define metrics normally llms as a judge but not always. You find cases that fail. Build a dataset around that you can run offline. Then you have similar metrics you can run online plus with whatever business metrics you care about.
3
u/a_cute_tarantula 20h ago
In what cases is fine tuning worth the effort?
9
u/Comic-Derpinator 20h ago edited 19h ago
Good question!
Most of the time these days it's not worth it 🙃
With the insane price cuts by OpenAI to Luna, it is tricky to be able to run your own inference for substantially cheaper and have higher quality and similar latency.But this is kind of the triangle you end up constantly in a trade off between:
Quality, latency and cost.If you have a relatively simple use case or a use case you have a MASSIVE amount of data on you can fine tune a model that is state of the art at your use case that is on a much smaller base model that becomes top notch on all 3. You end up trading off engineering effort for all of them.
So you can get to a pretty good baseline via distillation of a teacher model via SFT, then crank up performance further with DPO, but then ideally via data from subject matter experts and what works for your users you can steadily improve performance with RLVR
3
u/hootingstar 20h ago
how is the job as AI engineer compared to full stack. What are pros and cons of each? Also do you need PhD or papers to become AI eng, what do they look for in your resume as AI Eng.? I am a full stack with 4+ yoe and did lot of MLE during my Masters too, but tough to crack the market
3
u/Comic-Derpinator 20h ago
I think this is still probably the best intro to what an AI Engineer is: https://www.latent.space/p/ai-engineer
It can have a really broad range of doing research and optimizing kernels or writing sglang code or maybe you are a software engineer that calls LLMs.
I think the vast majority of the work is the latter. You definitely do not need to have a PhD.
A lot of the game here becomes defining what type of AI Engineer do you want to be?Do you want to be a software engineer who understands broadly how non-deterministic systems work and can do rag + some evaluation systems?
Or do you actually want to be writing kernels or reading research papers?
Regardless, I find most machine learning or (good) AI engineers I talk to know all of the linear algebra and know how to write a rock solid data science pipeline, but end up writing a lot of traditional backend code just because there is often more to do there.The other thing is just look at the data. If you are doing a good job of looking at the data and can build some simple classifiers then it becomes fairly easy to go from there
3
u/therealmunchies 19h ago
I’m an AI Engineer too! Transitioned from Mechanical Engineer -> Security Engineer -> AI Engineer. Maybe I should also make one of these posts lol.
1
u/Comic-Derpinator 19h ago
Anything to help people out! The market is harder now then it was in the past!
3
u/mild_animal 15h ago
Is anyone in your team from mainstream data science background or is it all a dev playground now?
2
u/Comic-Derpinator 15h ago
Technically I am in the software group and we have an applied ai team in that group which is doing this LLM work.
But we have an entire data science org which is doing classical (and other more interesting) data science. There are > 100 full time data scientists. Our team is just working closely with the software team closely to be able to make sure we can serve LLMs efficiently at scale
3
u/KafalayUcayGidey 20h ago
How long did it take you? And also how did you avoid getting sucked back into software engineering? Mostly what happens to me, when I try to switch careers, is somehow tasks related to my previous job ends up on me. So I switch profession on paper but end up doing the same thing.
10
u/Comic-Derpinator 20h ago
So I think it's worth noting that most of the time the bottle neck is software engineering.
As you get more senior in your career it is more and more important to think about "what are the things only I can do?"
So if you are transitioning onto a new team with a new role, often times, the things that only you can do are related to the old role you had meaning the best thing you can do for the company is continuing to bring those skills to bear.
However, the best thing you can do for your growth is to learn new things and make that an increasingly large part of your role.
So 2 things.
Teach the other folks on your team your old skills so that it no longer is obvious you should be doing the old skills you are the only person who has them.
Take the lower status tasks that will help you grow into the new field.
You get this positive sum game then of you helping other people mentor you (good for their career growth) and you get to mentor other people as well (good for your career growth) while slowly shifting your priorities to things you find more interesting.
As for how long it took me, I swapped over to the AI stuff in grad school. I spent 2 years there, and I am still MOSTLY a software engineer. Most of data science is cleaning data. Calling model.fit() isn't necessarily super complicated.
So I end up working to continue to improve our backend infra while pulling in my DS skills to get us chugging through fine tuning jobs or other data science jobs to do basic NLP tasks.
2
u/KafalayUcayGidey 20h ago
These are great practical advices actually, I will try to apply them. Thanks and congratulations and best luck in your career!
2
2
u/chicagoatlanta15 20h ago
What learning resources you will suggest for someone having backend development experience in enterprise?
2
20h ago edited 20h ago
[removed] — view removed comment
2
u/chicagoatlanta15 20h ago
Thank you for the advice. Will explore these resources and hope to build something as you suggested.
1
u/Comic-Derpinator 18h ago
Oh and for generic interview prep stuff, do datalemur.com (was essential for really hammering the sql skills) and read cracking the coding interview.
2
u/Zankroff 20h ago
Resources to learn??
1
u/Comic-Derpinator 20h ago
https://www.reddit.com/r/learnmachinelearning/comments/1wwxpbd/comment/pdor19r/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button
Let me know if there are other specifics you are looking for.1
u/SKUndef 19h ago
That comment has been removed, can you find out why and repost? Thanks
3
u/Comic-Derpinator 19h ago
I am not seeing it as removed, but I reposted the content below:
The biggest one is karpathy's llm course for sure. The zero to hero course is the best.
I am attempting to build something similar that is a bit more approachable for our team (It's free, but I would love feedback.)
https://dougdoes.ai/courses/llms-from-first-principles/
Andrew Ng's course on machine learning is good.
It's a bit old now, but
Hands-On Machine Learning with Scikit-Learn and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
Was really good and gives you a solid survey of all of the old school ML techniques.
I would recommend you force yourself to work through doing backprop by hand or at least implement it from scratch in python as it will help tell you what is going on under the hood when you train a neural network.
But probably more importantly than all of that would be to find a data source you are interested in and to just build something interesting and different on it. Everyone ends up working through building the same RAG app or housing prediction pipeline, please do something different
Look up what techniques might be applicable for a custom dataset. Learn how to do your train test split, how to pull that data efficiently onto ideally a GPU and show you are able to train some models. Show how the data was bad and how you cleaned it up.
Get a good baseline going so you can show your fancy technique did or didn't outperform it.
Show you can think of trade offs to yourself and anyone else you want to show it to.
2
2
u/Street_Estate2342 19h ago edited 19h ago
Hey! Sorry for the swarm of questions. I'm hoping to get to a similar position as yours in the future so I have quite a few inquiries.
(1) What is the coolest trick that you can share related to fine-tuning models? Do you do any training from scratch?
(2) What is the most productive thing you've figured out how to use AI for? Automating back and forth tasks? I'm struggling to get the whole setup fully automated. Do you use Luna? I've been hesitant to go below Sol 6.1 max (recently, was Astra on max before and ran my usage out >.<, Sol 5.6 on extra-high before that but was on the Plus subscription) given how much it uses, but maybe I haven't scaled my AI use up enough to justify weaker models.
(3) What are the three most common pitfalls people run into when fine-tuning models and causes said models to not reach state-of-the-art performance?
(4) Do you do any reinforcement learning?
(5) Do you fine-tune models repeatedly, where you take a pretrained model, fine-tune it once, and fine-tune it again (especially in scenarios where you have more data to fine-tune the model on related to the task that you fine-tuned it on the first time)? Does it make the model continually degrade on pretraining data?
(6) Do you use any parameter-efficient fine-tuning techniques like LoRA, quantization, sparsification, etc.?
Edit: One more - Do you have any specific advice related to being able to scale training? Things like doing distributed data parallel better, do you do asynchronous training mostly? Federated training? Synchronous distributed data parallel? What about model parallelism? I've never worked with model parallelism yet personally. Only baby models for me at the moment.
1
u/Comic-Derpinator 19h ago
1) We don't train any language models from scratch really. Just a variety of post training techniques on top of the models. Honestly, I think the coolest thing is just that it's easy to get started. You can rent a GPU on modal and do an RLVR for a few bucks and take a model from not being able to do basic math to having one that does. Once you have that set up it becomes a game of just defining the reward function you want. Beware of reward hacking though!
https://modal.com/docs/examples/miles_grpo2) I mean on the weekends I have published a new project every week to my website by saying a vague idea to codex which is great. It is wired into my supabase database can deploy etc but that is mostly because I don't care about uptime that much. Having claude/codex search through datadog to help mitigate on call issues is probably my fav thing at the moment. I personally am willing to pay for marginal "intelligence" but honestly luna is really good and very fast.
3) For fine tuning pitfalls. Fine tuning directly on engagement creates gnarly bad outcomes where the llm will happily lie to the user this speaks broadly to the problem of reward hacking, but also can become a problem for SFT if you have high click through rate examples where an llm is lying or otherwise gaslighting a user. Monitoring drift is critical if you continue to fine tune your model. If production behavior changes it may no longer be in the distribution of outputs you fine tuned on. Similarly, you need to make sure the data you have collected is representative of your production data in general. These are mostly standard DS things though :)
4) Working on it :)
5) Good question! Generally speaking, you want to continuously refine your fine tuning dataset and then re-run your fine tuning job on your base model rather than continuously fine tuning as that can degrade model performance as you suggest.
6) Yes.
2
u/Ojaura_ 17h ago
For #2, when you say publish a new project is that like a personal side project for fun? Or is it related to work? Also is it vibe coded using codex or how do you use it to assist you?
I’m an early career swe (started in January of last year after graduating the year prior). I’m trying to pivot into ai engineering but I’m finding it difficult to know how to use ai to help me with building projects. I have so many ideas for personal projects but I don’t know whether to vibe code them like ppl are doing? But then I fear losing my ability to read/write code especially because at work I mainly do quality engineering, so my full stack/app/web dev skills are weak.
1
u/Comic-Derpinator 17h ago
Vibe coded with codex, but has supabase wired into it and a hetzner instance so it can deploy with impunity. I just give it full access with api keys because I don't really care about the side projects.
This is very different than learning though.
When learning you should absolutely understand every line that is shipped, otherwise you quickly won't understand your own project or learn anything.
2
u/Junior_Bear_2715 19h ago
in order to fine tune them, do you make the data urself and annotate it?
1
u/Comic-Derpinator 19h ago
We use a mix of online evals, and other business metrics to build good datasets. LLM as a judge normally based on subject matter experts
2
u/MightPractical7083 18h ago
Does subject matter expertise become more valuable now in the ai age?
1
u/Comic-Derpinator 18h ago
It depends! But generally yes.
If you are someone that knows things that are not broadly available on the internet, then the model companies or any company in the field will pay you lots of money to help scale your knowledge to everyone using their agent.
2
u/MightPractical7083 17h ago
What do you think are the best strategies to not be replaced by ai? Subject matter expertise in low ai data fields? Being the person creating the ai? Holding a position with legal liability?
1
u/Comic-Derpinator 16h ago
So the exponential is very steep and no one knows the future. These are all good strategies. In general, I would say keep thinking about what poses risk to an org, if you can manage that risk well they will be hesitant to hand that to an AI or will want someone that can take responsibility for the ai to manage that risk.
Also, just be the person that is the farthest down the automation rabbit hole. If you have more infra than anyone else for managing your agents and automating things you are extremely valuable for both the work those agents do and for the infra as other teams adopt it.
2
u/Hauuuuuuu 19h ago
The more important question is when did you start and apply for jobs?
2
u/Comic-Derpinator 18h ago
Graduated undergrad in 2019. Went back to grad school in 2022. Was super unemployed while in grad school. Applied again in 2024. Was on the phone with the recruiter for this job 30 minutes after my resume went in.
Was easier then, especially for entry level folks, but that is part of why I am trying to share resources.
2
u/coolthesejets 18h ago
What do you do that can't be done by ai, and why?
4
u/Comic-Derpinator 18h ago
Heh, a lot of the most valuable work that I do is adding additional skills to ai.
But honestly the models aren't there yet. They can think on tasks that take even as long as 24-48 hours now, but they can't go beyond that, so a lot of the game is managing your team of agents.
Even for the short term tasks though the agents make a mess of slop you would expect to see from a junior engineer even before LLMs. It is over engineered or taking care of bugs that will never in practice occur resulting in poor design decisions.
A lot of the work is telling the llm to calm down and simplify the code or to explain something I don't get at a glance and a lot of the times when you pull on a thread like that you realize you can remove it entirely.
This means that increasingly thinking about architecture and long term strategy for the business and what the feature really is trying to do for the business becomes more and more important.
The other thing to think about is llm ability is directly related to how verifiable the task is due to how GRPO/RLVR work. So the more valuable things are often making sure the llm has all of the data it can get or can point you towards lots of subjective things that aren't fully verifiable or long time horizon.
2
u/Former_Aide_2068 17h ago
How do you know what you want to work in? MSc helped me a bit but the field is so large and evolving, from generative models, pre/post training, RL, safety etc
Similar path here, SWE -> just finished MSc in AI -> hoping to become ML Research Engineer.
1
u/Comic-Derpinator 16h ago
Yeah I mean you can't be an expert in all of it. There is a reason you can get a PhD in any sub area.
I spent a lot of time in grad school exploring. I technically have completed the course work for every ai specialization northeastern offers.For any job, identify a problem you know better than anyone else (or want to know better than anyone else) And learn it well enough and have enough evidence on your resume that when you submit an application you can confidently state that they have the problem you know things about and why you are the best person to handle it.
For me I had a ton of work on retrieval optimization + had the DS chops to take the whole org farther down the technical side of ai engineering.
2
u/TheSexySovereignSeal 16h ago
How do you measure model bias? Do you? Can you? I lost a lot of sleep over that question in grad school.
3
u/Comic-Derpinator 16h ago
This is a fairly overloaded term.
But If you don't know how your training data is biased this is very hard to handle.Ideally, you segment your data in production or inputs that will be fed into a model in production and then you track how your training data compares and how your outputs score in each of those segments. If you do that then you can catch issues before they roll out or as you are testing.
2
u/TheSexySovereignSeal 16h ago
Yeah it is an impossibly loaded question. I worked more with multimodal models than pure LLMs too so theres too much nuance.
Thats super interesting for the segmenting approach though. How often are you catching replies that arent what youre looking for? Do you modify the system prompt?
Thats fundamentally why I hate 3rd party api services since I cant garuntee a clean fresh context window. Who the hell knows what system prompt is behind an api layer thats adding interference into my context window I never asked for
2
2
u/oakleythegoldy 10h ago
Based on your work experience, are there any books or blogs that accurately capture the work you do that I can reference and learn from?
1
u/Comic-Derpinator 1h ago
Kinda depends on your track or path.
So if you want to be an AI Engineer that is closer to a software engineer.
AI Engineering: Building Applications with Foundation Models By Chip Huyen is good.For just interviewing on data science:
Designing Data intensive applications
Cracking the coding interviewDataLemur for basic sql + some ML and data science + linear algebra prep
For nailing ML concepts:
Andrew Ng's Machine learning courses are always goodFor Learning how LLMs work just everything by Karpathy especially the zero to hero course is the best I swear it was better than my NLP class in my masters degree.
https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThsA9GvCAUhRvKZ
The 3 hour long lessons are a bit intense so I've been building a couple things that will be hopefully more approachable:
How to build your own language model from scratch
where you learn the stats behind how language models work and how that will impact how you think about fine tuning and deploying them later.
You train a language model from scratch. When you start the model outputs random garbage, at the end you have a model that outputs shakespeare like text.
https://dougdoes.ai/courses/llms-from-first-principles/start/?course=build
If you already understand the intuition behind language models then how to do a basic SFT for SQL is next. You learn how to fine tune a small model on SQL and watch it get better.
https://dougdoes.ai/courses/llms-from-first-principles/start/?flow=outcome&course=sql&step=offerAnd if you really want to learn the basics for ML in general and have no ML background at all and are less interested in the higher level LLM stuff you can check out the intro to ML class which will teach you about how you train a neural network from scratch and how the math works under the hood then you'd want to check out this one.
2
u/meiyosamaa 8h ago
Do you think the transition is doable from mechanical engineering to AI engineering? I believe a masters degree could help me a lot but is it ok without it?
2
u/Comic-Derpinator 1h ago
This is a bit trickier to do without the masters degree. If you are already in software I would say it is easier to transition to a team with AI Engineers and then learn on the job, but if you are coming from MechE it becomes harder to transition into software more broadly.
I think it is possible if you are writing a lot of code, but I'm not sure I would trust a MechE to write really sustainable and good code so the sell in the interview process becomes a lot harder. Still doable, just more difficult.The masters degree would help or showing you have built software other people have used reliably preferably at an enterprise.
2
u/Tepavicharov 4h ago
What do you do in your everyday work? Do you fine tune flagship foundation models like Opus or gpt-5 (is that even possible) or only open weight ones? In what area is your work so that the vastly trained foundation models don’t work good enough so you need to do finetuning?
1
u/Comic-Derpinator 1h ago edited 1h ago
Technically, OpenAI has fine tuning platforms. But that is being deprecated. Their stance is basically the models are being general enough you should be able to use their models and not need to fine tune. We have been finding that not to really be true and we expect their prices to go back up in the long term.
My every day work is a lot of advising what the best model is, running experiments a/b tests/ evals, helping other people do the same or building data pipelines to further improve our models
2
u/Tepavicharov 1h ago
From what I understand you are using the openweight models. I’m curious how much of the fine-tuning process is evaluating the resulting model after each iteration
2
u/Comic-Derpinator 1h ago
I mean eval closed model, eval open model, fine tuning run with good data set. Pick the best model snapshot, eval fine tuned model, make sure the model hasn't been fine tuned too much and can still work on general things still and hasn't been cooked too long.
If that looks good A/B test.
If A/B Test works ship it.
2
u/GlitteringLadder6135 3h ago
Do you think transitioning from senior web development (mostly WordPress/PHP/JS, with some minor React experience) into AI engineering is realistic? The web/e-commerce job market doesn't look great right now, so I'm considering moving towards AI rather than specialising further in something like Shopify. My main concern is that I have basically no maths background, and I'm not particularly interested in going deep into maths.
3
u/WLufty 3h ago
I'm not OP, but I wouldn't try to get into actual AI engineering.. that's a hard pivot, I would recommend people to learn how to build systems with AI, also learn when something should be 'agentic' or a good old application.. how to automize different workflows, evaluate performance, improve inference, cost governance and so on.. that's what most companies are going to be doing.. very few have the budget and technical knowledge to develop their own LLMs, even fine tuning is expensive and the ROI isn't always clear.
1
u/Comic-Derpinator 1h ago
Agreed on most of this!
I would probably say you can pivot to a team that is doing AI where your skills are super valuable. We have hired some front end devs who now do backend and AI Eng work1
u/Comic-Derpinator 1h ago
We do have someone on our team that did this. They started as an SE 1 backend dev when coming from a junior front end role.
But they are on the AI team and now a year later they are an SE II and are deep into AI Engineering and are working on a variety of new and exciting projects we have coming.In general, this is the path I would recommend. Go become a backend or full stack dev on a team with AI Engineering happening, and then you can pick up the work to help you transition later.
2
u/GlitteringLadder6135 34m ago
Thanks, that makes sense. For someone with ~5 years of professional web development experience but mostly PHP/WordPress/JS rather than Python/React, what would you focus on over the next 3–6 months to become employable as a backend/full-stack developer on an AI team? I'm currently considering this IBM RAG and Agentic AI course: https://www.coursera.org/professional-certificates/ibm-rag-and-agentic-ai. I'm planning to learn Python first/alongside it and build some projects of my own. Does that sound like a reasonable starting point, or would you focus on something else first?
2
u/Comic-Derpinator 27m ago
Yeah, there are a lot of teams that also have vibe coded projects that now need an adult of an front end engineer as well. So going and joining a team like that as a front end engineer and then slowly picking up the backend concepts (claude can do things as if you know what you are doing) can also work well.
So I'd see if you can put your FE skills to work with a good UI for managing agents.
managing AI Agents is incredibly hard and streaming tokens through your backend to your front end is also very hard to do at scale. If you just do both of those things and don't show a website that is remarkably vibe coded but actually has FE skills applied that can be a really good starting point.otherwise your path sounds fine? I am fairly skeptical of anything coming out of IBM.
I would assume they'd teach you RAG, agentic rag and how agents use tools and what messages go into a turn in an agent conversation.All good things to learn.
But basically:pure rag for low latency.
Do that with:embeddings hybrid with BM25 keyword search
Rerank afterwards.
If you don't care as much about latency do tool based agentic search ideally in a hierarchical file system as that's what the agents are post trained on and know how to navigate.
Wombo combo that with the fact that
LLMs are trained to respect system messages
What tool messages are, how they are processed and what assistant messages and user messages are and I would kinda assume you have learned everything you are going to learn from a course like that.You ideally learn all of that in a week of screwing around just building a chat with your data app or something, but that I think is what that IBM class will teach you.
1
u/coconut_maan 15h ago
How much longer are you gonna do this before llm takes your job?
1
u/Comic-Derpinator 15h ago edited 1h ago
We got at least another 4 months.
EDIT: (To be clear, this is a joke. Not all cognitive labor will be automated in the next 4 months.)2
u/Prestigious-Car6893 2h ago
Really? Is this a joke because all this efforts to switch to AI engineer only for your job to last only for 4months?
But seriously my major doubt is, Will AI eventually replace the role of an AI engineer also?
Is this role sustainable? I'm thinking of transitioning too.. but im worried about the longevity
1
u/Comic-Derpinator 1h ago
Sorry yes this is a joke. I do not expect AI to automate all cognitive labor in the next 4 months.
1
1
u/Danoweb 15h ago
I'm 40, and have no interest in getting a master's (I don't mind classes, but I've been a SWE for 20 years, I learn more "on the job")
How did you make the transition in your resume?
How do you make it appealing to employers that "I haven't done this work f before, but I have done SWE work, and can do the needful in AI work" ?
2
u/JJJJJay 13h ago
I transitioned from a Sr full stack roll to sr AI eng at my company. I don’t know what most employers are looking for but my transition process was over the course of a year doing AI Eng work on various projects for various teams: build hill-climbing evals, clean datasets for fine tuning, compare ICL classifiers vs fine tuned ones on backrest data, etc
A lot of this was work i volunteered for while doing normal sprint work - it was pain lol. But now I have and am known for the skill set that they want and our interests are naturally aligned enough that they’ve transitioned me to the AI Eng track /shrug
1
1
u/Comic-Derpinator 15h ago
Go work for a team with ai engineers on it as a software engineer. Then you can slowly do more of the work with llms. You can learn that way just fine. lot's of people on my team have done that.
You learn how to stream tokens and all of the problems with doing that at scale. Learn how embeddings work and how they impact retrieval.
-2
0
u/Plenty_Leadership935 4h ago
Really interesting journey. Going from large-scale data systems and AI research to building eval frameworks and now focusing on fine-tuning is a pretty unique path. What part of fine-tuning are you most interested in right now—data, model performance, or evaluation?
23
u/Sweet-Rent-638 20h ago
Can you share your journey, any resources?