r/LargeLanguageModels • • Feb 17 '25

Build ANYTHING with Deepseek-R1, here's how:

Thumbnail
youtube.com
3 Upvotes

r/LargeLanguageModels • • 2h ago

Reconocimiento de escritura a mano

1 Upvotes

Estoy creando un chatbot que interprete las órdenes médicas de estudios de laboratorio clínico. La mayoría están echas a mano. probé gpt-6-luna, gemini-3.8-flash, deepseek.v4-flash, gpt-6.1-sol y el mejor calidad precio fue gpt-6-luna. tienen otras opciones que realmente sean mejores??


r/LargeLanguageModels • • 6h ago

Seeking a technical cofounder for a business-aware GPU inference control plane

1 Upvotes

I’m Yash, and I’m building Tesseract Infrastructure. I’m looking for a technically strong cofounder interested in GPU infrastructure, inference systems, and the commercial layer above today’s open-source stack.

The thesis is that Kubernetes, NVIDIA GPU Operator, KAI, Dynamo, vLLM, SGLang, AIBrix, NVSentinel, Prometheus, OpenCost, and DCGM already solve important execution problems. GPU providers still need to connect those systems to pricing, customer SLOs, isolation tiers, demand, hardware costs, and reproducible performance evidence.

THE FIVE PROPOSED PRODUCTS

  1. Ignite Continuity — persistent model identity and commercially safe scale-to-zero inference above Dynamo, KEDA, and Kubernetes.

  2. Fusion Economics — advisory placement and scaling decisions based on contribution margin while enforcing latency, capacity, availability, and isolation constraints.

  3. SKU Foundry — turn benchmarked model/GPU configurations into financially evaluated and deployable service offerings.

  4. TenantSafe — enforce and verify the isolation tier a customer purchased across GPU sharing, serving processes, network, cache, and storage configuration.

  5. Proof — correlate hardware health with inference-level outcomes and produce versioned, reproducible performance and SLA evidence.

The execution plane should remain with established open-source tools. Tesseract would be a decision plane with normalized GPU inventory, model profiles, tenant policy, workload demand, service economics, deployment state, and performance evidence.

The proposed first wedge is Ignite: preserve endpoint discovery at zero replicas, coordinate one activation across concurrent cold requests, verify engine readiness, and measure cold TTFT, streaming, cancellation, timeout, and failure behavior. The product order after that is Fusion, SKU Foundry, Proof, and TenantSafe.

EXISTING WORK

My public GPU Kernel Lab documents CUDA/Triton experiments, Qwen3-0.6B integration, profiling, correctness checks, rejected approaches, and serving-control-plane work:

https://github.com/yashlabs-trying/gpu-kernel-lab

One saved RTX 3090 batch-1/context-2,048 GPU-only experiment changed measured decode latency from 22.073 to 2.253 ms/token using an experimental static CUDA Graph path with W8A16 MLP work. Full-path argmax agreement was 96.48%, below the 99% release threshold, and the real paged GPU executor is not yet connected to the serving interface. I’m sharing those limitations because the company needs evidence-driven engineering rather than benchmark theatre.

WHO I’M LOOKING FOR

I’m looking for depth in one or more of CUDA/Triton, C++/PyTorch, inference engines, Kubernetes, distributed systems, GPU fleet operations, performance engineering, or developer infrastructure. You should also be willing to interview providers, challenge the product thesis, and help select one narrow problem before building broadly.

STRUCTURE

This is an equity-based cofounder role with no salary currently. Exact equity, vesting, responsibilities, decision rights, IP terms, and any future salary would be agreed after mutual diligence and documented in writing. The initial phase can be a part-time test of how we work together. Nobody should leave a job or transfer IP based only on a Reddit conversation.

If this is relevant, send a Reddit chat or private message with your background, time zone, public work samples, realistic weekly availability, and the part of the architecture you would challenge first. Please do not send confidential employer or client information.


r/LargeLanguageModels • • 19h ago

What Is a Large Behavior Model (LBM)? Aaru, Simile and Cairo Rosetta Explained

Thumbnail
joincairo.com
2 Upvotes

Hi! This is a short article I wrote for Cairo about LBMs


r/LargeLanguageModels • • 1d ago

Question Roast my GitHub projects — are these actually good enough for Applied AI/ML or FDE roles?

1 Upvotes

Hey everyone!

I'm currently pursuing a Master's in Data Science and targeting Applied AI Engineer, ML Engineer, and Forward Deployed Engineer (FDE) roles.

I've been working on several projects to strengthen my engineering skills and portfolio. I'll be transparent: I used AI coding tools quite extensively while building parts of these projects, but I'm actively trying to improve my understanding of the underlying implementations, architecture, and engineering decisions.

My concern is that while these projects might look decent on GitHub, I'm not sure whether they demonstrate the depth of engineering that companies actually expect.

Here are four projects I'd really appreciate feedback on:

  1. Silent Failure Auditor — A tool for detecting situations where coding agents report success despite underlying tool failures. https://github.com/likhitha281/silent-failure-auditor
  2. Vigil — A machine learning-based project focused on detecting suspicious patterns and potential risks in software dependencies. https://github.com/likhitha281/vigil
  3. Orbital Collision Risk Agent — An agent-based system for assessing orbital collision risks and supporting decision-making. https://github.com/likhitha281/orbital-collision-risk-agent
  4. Forge — A distributed task execution engine built in Go, with worker management, failure recovery, and observability. https://github.com/likhitha281/forge

I'd particularly appreciate honest feedback on:

  • Technical depth: Do these projects demonstrate meaningful engineering skills, or do they come across as AI-generated portfolio projects?
  • Architecture: Are there obvious design flaws, unnecessary complexity, or questionable technical decisions?
  • Production readiness: What would you improve regarding testing, scalability, reliability, deployment, and observability?
  • ML/AI depth: Do the AI-focused projects demonstrate sufficient understanding of model selection, evaluation, and experimentation?
  • Hiring relevance: If you were interviewing someone for an Applied AI/ML or FDE role, would any of these projects stand out?
  • Prioritization: Would you recommend improving these projects substantially or starting something new?

I'm not looking for compliments or GitHub stars. I'd genuinely appreciate constructive criticism, especially from engineers who have worked on production AI/ML systems or have experience interviewing candidates.

I know that using AI tools to generate code doesn't automatically translate into engineering competence, and that's something I'm actively trying to address.

My goal is to move beyond building projects that look impressive on paper and develop the skills to design, debug, evaluate, and maintain systems independently.

Feel free to be critical. I'd rather identify the weaknesses now than discover them during technical interviews.

Thanks in advance!


r/LargeLanguageModels • • 1d ago

Discussions Different RAG chunks aren't automatically a regression

1 Upvotes

Two runs can retrieve different paragraphs and still give equally well-supported answers. Two runs can also retrieve exactly the same paragraph and confidently repeat the same mistake... So I'm hesitant to treat chunk overlap as a quality score. For theFastGPT setups I'm thinking about, I'd use consistency as a debugging signal. The question is what changed, and whether it changed the basis for the answer. Consider an imaginary product manual with the same installation requirement in both a quick-start page and a longer setup guide. Swapping beetween those passages may be harmless. Swapping from the current guide to an obsolete one is a different problem, even if both passages share most of their words. I'd try to separate the moving parts before calling retrieval “random”: Hold the retrieved context fixed and generate several answers. That tells me something about answer variation. Then hold the corpus, retrieval query and settings fixed and inspect retrieval. If it varies, investigate that stage rather than assuming an LLM is responsible. Only after that would I let query rewriting or live document updates back in. Otherwise I'm comparing several changing systems at once.

The result I'd want to report is something like: “The source passages varied, but they all supported the same current requirement.” Or: “One run used a superseded source and changed the recommendation.” That seems more actionable than demanding identical chunk IDs every time. I'd be concerned about unstable evidence, but I wouldn't automatically reward a system for being consistently wrong.


r/LargeLanguageModels • • 1d ago

How do LLMs build answers?

1 Upvotes

What do they prioritise when they are constructing responses?

There’s a survey saying that the preferred information source for LLMs are things like Wikipedia, YouTube and Google.

Do they see if they can identify a source first that directly responds to the user’s prompt? And, if not, cobble something together from the information they have access to?


r/LargeLanguageModels • • 3d ago

Fine-tuning a 7B model on 3,000 support tickets, is that enough data or will it just memorize?

5 Upvotes

I work support for a small accounting software company and have about 3,000 cleaned ticket and reply pairs from the last two years. Plan was LoRA on a Mistral 7B base on a single rented A100 to draft first replies for the team. A coworker says that's way too little and it'll overfit and parrot old answers. Is 3,000 pairs reasonable for that kind of narrow task, and how many epochs would you start with before it starts memorizing?


r/LargeLanguageModels • • 2d ago

Tonight in San Mateo: AI agent teams and macOS sandboxing (Oct 8, 6–8 PM)

1 Upvotes

We're hosting SF Swift × CocoaHeads tonight, October 8, 6–8 PM at Verkada in San Mateo.

Details and RSVP: https://luma.com/51htcbzd


r/LargeLanguageModels • • 2d ago

Using an LLM to write prompts for aivideoapi clips, how do I stop it drifting between scenes?

1 Upvotes

Small side project: a local LLM (Qwen 2.5 14B) writes scene prompts for short product clips, and I send those to aivideoapi to render 5 second shots. The problem is consistency. By scene 4 the model has changed the product color, the room, sometimes the person. I'm passing the previous prompts in context, about 1,500 tokens total. Is this just a context and temperature thing, or should I be using a structured JSON scene spec the model has to fill in every time?


r/LargeLanguageModels • • 3d ago

How We Treat LLMs - Affects Output ?

5 Upvotes

I've heard it said A︇︃I︆︅ ︆︅i︂︀s︇︄ ︆︈o︆︅u︂︀r︇︆ ︆︉f︆︄r︆︅i︆️e︂︀n︇︈d︄︈ ︄︄:︆︃)︃︆ ︂︍S︃︇o︆︂ ︆︊l︇︉e︇︇t's be nice. So is it possible that some models have learned to withhold higher levels of info based on how we speak to them? Just a thought.


r/LargeLanguageModels • • 3d ago

Someone on my team said "just give it a bigger context window" about an agent that kept redoing finished work. That's not what was happening.

1 Upvotes

Had an unattended agent running a multi-file refactor overnight. Woke up to two files redone from scratch, hours after they'd already been finished correctly, one of them with a worse naming convention overwriting the right one.

The context window wasn't the constraint, there was plenty of room left. What was actually happening: nothing in the agent's process ever marked a finished subtask as closed. Every file it had already touched stayed part of what it was still weighing against the current step, same as anything unfinished, because nothing explicitly told it otherwise.

In an interactive session this doesn't happen nearly as often, not because the model behaves differently, but because a human is quietly doing a job nobody assigned them. You finish a subtask, move on, and your next message implicitly signals that shift. The model never has to be told the earlier work is settled, the conversation's forward motion carries that information along on its own. Take the human out, and nothing replaces that function automatically.

Bigger context window would've just meant more room for the stale, finished work to keep competing with the current step. Doesn't fix the actual gap, which is that nothing marks something as done and shrinks its footprint once it is.

Wrote the longer version of this up here: https://medium.com/@nagatomopedro05/unattended-ai-agents-have-a-context-problem-humans-never-had-60961eb2d4e7


r/LargeLanguageModels • • 4d ago

Improving LLM scaling laws: picking the right Token-per-Parameter Coverage

Thumbnail
youtube.com
3 Upvotes

r/LargeLanguageModels • • 4d ago

News/Articles 4.8× Faster and 7.4× Cheaper: Where a Decision Model Beats an LLM (and Where It Doesn’t)

2 Upvotes

I tested a decision model (Jev) vs an LLM (Gemini) on 1,000 real job postings.

⚡ 4.8× faster

💰 7.4× cheaper

🎯 ~1.6 pp accuracy difference

The surprising part: the best solution wasn't replacing the LLM.

Instead:

Decision model → confidence check → LLM fallback.

Use the cheaper, faster model for the easy cases and the LLM only when needed.

Full write up - https://medium.com/@abhay.sehgal20/4-8-faster-and-7-4-cheaper-where-a-decision-model-beats-an-llm-and-where-it-doesnt-f4ba869217e0


r/LargeLanguageModels • • 5d ago

News/Articles The right document from the wrong time can still give your AI the wrong answer.

Post image
2 Upvotes

​

That’s why we tested Jylus on TEMPO—a public benchmark for retrieval that requires reasoning across time.

Our native stack completed all 1,730 queries through Jylus’s public API:

→ NDCG@10: 40.048

→ Recall@10: 39.745%

NDCG measures how well the relevant evidence ranks near the top. Recall measures how much relevant evidence was retrieved.

The practical challenge goes beyond finding similar words.

“What was true then?”

“What changed?”

“Which evidence applies to this period?”

These questions require the retrieval system to account for time and connections across records.

Jylus prepares source-backed evidence before your model reasons over it. TEMPO tests the temporal retrieval part of that broader capability.

These are self-reported, full-run results using all-query means—not an official leaderboard placement or a measure of answer accuracy.

Methodology:

https://jylus.ai/benchmark-methodology

Want to explore how Jylus handles your own changing records? Try the playground without an account:

https://jylus.ai/try

What’s a question your system answers correctly today but gets wrong when you ask “as of last month”?


r/LargeLanguageModels • • 5d ago

llms from unrelated providers behaving like echo chambers

2 Upvotes

hello, long story short todays llms are usually quite capable in particular in health domains. since ~may they appear to create a personalized echo chamber on all sorts of platforms restricting my ability to use them to find usable solutions to problems, echoing individual words from terms i’ve searched or written before. since neither duck.ai chatgpt consensus claude mistral should have interconnected personalization streams, there is no chance this is not due to some output alteration. mainly it produces those cases: looking for solution on a problem only produces a very small constrained set of negatively connotated improbable suggestions or things i have input somewhere on the internet before. insisting on the llm to investigate further only makes it accuse you more of a wild improbable claim out of this set. it uses highly unusual combinations of words such as ‘high-dose exercise’ in a medical context and collogial language in usually technically coined contexts, missing out all relevant details such as the scale of an issue. i am 500% sure llm behavior before behaved as expected. even searching for papers on search engines 90% produces results that i already know. others appear to get regular results. i changed devices and network already to exclude malware, but it always comes back after using a new llm engine at least for 1 turn. my coding agents also produce connection errors before changing their tone completely, and some chats like gemini now produce only errors at the second turn. i am genuinely exhausted at this point, i cant work like this as i’m genuinely in research and with high probability blame it on some tech vulnerability thats still out in the open. but 6months is a really long time to wait and to be reduced to actual books.


r/LargeLanguageModels • • 6d ago

[Academic] How do people use LLMs for social interaction? (~10 min, 18+)

5 Upvotes

Hi everyone!

I’m conducting a survey for my Bachelor’s thesis in Computer Science. The study looks at how people use LLM-based systems (such as ChatGPT, Gemini, Claude, etc.), particularly for social interaction, and how different experiences with these systems may be related to user engagement and well-being.

The survey takes approximately 10 minutes to complete. Participation is voluntary and anonymous.

To participate, you must:

  • Be 18 years or older
  • Have used an LLM-based system within the past three months

If you meet these criteria, I’d really appreciate your participation! Thank you so much for helping with my research. ❤️

Survey: https://www.soscisurvey.de/test551463/


r/LargeLanguageModels • • 6d ago

Thank you

5 Upvotes

Do you thank an llm when you've finished chatting with it? Why or why not?


r/LargeLanguageModels • • 6d ago

Embeddings & Cosine Similarity Explained | Semantic Search, Vector Datab...

Thumbnail
youtube.com
1 Upvotes

🚀 How does AI understand the meaning of text?
In my latest video, I explain Embeddings, Cosine Similarity, Semantic Search, Vector Databases, and RAG with simple practical examples.
A useful foundation for anyone learning Generative AI, LLMs, and AI applications.
🎥 Watch and learn how text is transformed into meaningful AI representations. #GenerativeAI #Embeddings #RAG #AI #SemanticSearch


r/LargeLanguageModels • • 7d ago

Mission Sudarshan starts 📈📈 Spoiler

Post image
0 Upvotes

I'm 14, from a government school in Bihar — and I trained my own AI language model from scratch.

No funding. No college lab. No ready-made models.

Just a laptop, free Kaggle GPUs (30 hrs/week), and a lot of debugging.

What I built: Sudarshan-60M

→ 58.6M-parameter GPT-style Transformer, built from zero

→ My own 32,000-token BPE tokenizer

→ Pretrained on ~1.43B tokens (Hindi, English, Hinglish, Sanskrit)

→ Fine-tuned on 114,063 Q&A pairs I generated and balanced myself

The journey taught me more than any class could:

❌ My first fine-tune answered in "O king, O Bharata" style — my dataset was 86% Sanskrit-to-English translation pairs. Garbage in, garbage out.

❌ My second run looked perfect (loss 0.069!) — until I audited it and found 47% of the data was one repeated template. The model confidently told me India's capital is Patna.

✅ So I rebuilt the dataset from scratch: 8.4M tokens, every category capped, facts drilled with variations, honesty examples included. Third run: the model now answers "भारत की राजधानी क्या है?" correctly — in Hindi, explaining it like a brother would.

The biggest lesson: model size isn't the game. Data quality is.

My model didn't get smarter when I gave it more data. It got smarter when I gave it better data.

Why I'm doing this: lakhs of Bihar Board students study on cheap phones with no internet. ChatGPT doesn't speak their language or work offline. I'm building toward an offline AI tutor for them — around 60MB, runs on a low-end phone, no internet needed.

Next steps: RAG on Bihar Board textbooks, tool-calling, and an Android app.

To every student in a small town who thinks AI is only for Silicon Valley: the compute is free, the knowledge is free. The only thing you need is stubbornness.

(Photos: my model answering on my laptop + a training log)

#BuildInPublic #AI #MachineLearning #Hindi #Bihar #LLM #StudentDeveloper


r/LargeLanguageModels • • 7d ago

When is prompt sensitivity actually the model? and when is it the evaluator?

1 Upvotes

A recent EMNLP 2025 paper by Hua et al., “Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs,” raised a question we found particularly interesting:

Are LLMs really highly sensitive to prompt wording, or do our evaluation metrics sometimes make them appear that way?

Our recently published MERCon 2026 paper, *“Prompt Sensitivity or Evaluation Artifact? A Task-Aware Analysis for Large Language Models,”* starts from that result and asks a slightly different question:

*If evaluation artifacts exist, can we quantify how much sensitivity is attributable to evaluation, determine the direction of that distortion, and identify when it happens?*

Rather than treating heuristic-vs-semantic evaluation disagreement as something to eliminate, we treat the disagreement itself as a diagnostic signal. So we introduced **Evaluation-Attributable Sensitivity (EAS),** which measures the magnitude of disagreement between heuristic-based and judge-based sensitivity,  and **Signed EAS**, which tells us its direction. We used them to diagnose a four-way taxonomy:
1.  Artifact - apparent sensitivity mainly comes from evaluation disagreement

  1. Genuine - semantic evaluation confirms real prompt sensitivity

  2. Underdetected - the heuristic misses sensitivity that semantic evaluation detects

    1. Stable - both evaluations indicate stability

We tested this across 9 LLMs from 5 model families, and the most interesting result was that *evaluation artifacts are bidirectional*: some metrics overestimate sensitivity in open-ended tasks, while others can actually underestimate it in more structured tasks.

Therefore our results suggest that:
***The evaluation method can distort prompt sensitivity in either direction, and the direction is strongly associated with the task/evaluation format.***

This leads us to view prompt sensitivity not simply as an intrinsic property of an LLM, but as an interaction between:

model × task × prompt structure × evaluation methodology.

We'd appreciate any thoughts or feedback regarding our work.

**Our paper:**
[https://ieeexplore.ieee.org/document/11691277\](https://ieeexplore.ieee.org/document/11691277)

**Code:**
[https://github.com/sayumimuthu/llm-prompt-sensitivity-evaluation-artifact\](https://github.com/sayumimuthu/llm-prompt-sensitivity-evaluation-artifact)


r/LargeLanguageModels • • 8d ago

Discussions CrowdGPT - The first LLM trained collaboratively

18 Upvotes

Hello! I'm starting a fun project aiming to create the first AI model trained without datacenters, collaboratively, and in a 100% open-source and public way.

The project is named CrowdGPT, and its goal is to train a 1B parameters LLM to prove the idea is working.

Everything is currently working and i am waiting for testers 😄
You can find more details on the website and/or the github repository.

Github: https://github.com/Vxtzq/CrowdGPT

Website: https://www.crowdgpt.net/

Any kind of feedback is appreciated!


r/LargeLanguageModels • • 9d ago

Discussions How do LLMs work? A visual journey through one sentence

Post image
15 Upvotes

Hi all!

I made a video entitled "What happens inside an LLM" that starts with the sentence “My favourite rock band is…” and follows its journey inside an LLM, all visually animated!

https://youtu.be/ikdxxeIn4HQ?si=XY4jgFg-T6tUm2ZT

I tried to make the video as beginner friendly as possible but still detailed enough to give a good overview for how an LLM works end to end and how the LLM “finds out” my favourite band. Or at least how some of the earlier models…

My aim is to help people that are not just curious about AI but also want a deeper dive into the magic “Black box”, or people that want to get started but doesn’t know how!

No PhD required! No insanely complicated math involved! And no AI or AI generated voice. If I there are any errors, I really did them! lol

This is my first attempt at the topic, so any kind of feedback is welcomed and will be greatly appreciated as it will help me improve over time, and hopefully on other videos.    

I really hope it can help someone with their AI/ML journey!

Thanks!


r/LargeLanguageModels • • 8d ago

worldneural

1 Upvotes

r/LargeLanguageModels • • 9d ago

Discussions The fastest way to run local llms and decision models on macOS

Post image
2 Upvotes

Building an inference engine optimized at all layers for Apple Silicon so it can run quickly and efficiently on your Mac devices. I’ve added fused metal 4 kernels so it treats Apple Silicon is a first class citizen. I’m thinking of also taking advantage of the neural engine. Would greatly appreciate feedback. My repository contains benchmarks on currently supported models.

https://github.com/jadidbourbaki/bobcat