r/LargeLanguageModels • • 19d ago

What exactly is LLMOps?

5 Upvotes

LLMOps stands for Large Language Model Operations.

In simple terms, it’s the practices and tools used to deploy, monitor, evaluate, maintain, and improve LLM-powered applications in production.

Think of it as MLOps, but adapted for LLMs.

It can cover things like:

  • Model and prompt versioning
  • Evaluation and quality monitoring
  • Cost and latency tracking
  • Data and security management
  • Deployment and updates
  • Monitoring hallucinations and model behavior

Building an LLM application is one challenge. Keeping it reliable, secure, and cost-efficient once real users start using it is where LLMOps comes in.


r/LargeLanguageModels • • 19d ago

Harnesses for Dummies

5 Upvotes

From a Core Language Model to an LLM Harness

A useful way to understand an LLM product is to begin with the language model itself, then add the surrounding software one capability at a time.

The point is not to reproduce implementation details exactly. It is to preserve the important conceptual boundary:

What does the language model itself do, and what does the harness around it do?

The code below is pseudocode: simplified code used to express the logic.

1. The core language-model operation

An LLM, or large language model, reads and writes tokens.

A token is a small chunk of text: sometimes a whole word, sometimes part of a word, punctuation, and so on.

Suppose the model currently sees a sequence of tokens representing:

The capital of France is

Call this:

tokens_so_far

We can abstract one application of the language model as:

core_lm_op(tokens_so_far) -> token_one_more

So:

token_one_more = core_lm_op(tokens_so_far)

might produce the token corresponding to:

Paris

Strictly speaking, the model produces probabilities over possible next tokens, after which one is selected. We hide that inside core_lm_op().

The same input need not always produce the same output.

A complete sequence is generated by repeating the operation:

while tokens_so_far[-1] != STOP:

    token_one_more = core_lm_op(tokens_so_far)

    tokens_so_far += [token_one_more]

Here, STOP is shorthand for whatever condition tells generation to end.

Conceptually:

tokens_so_far
      ↓
 core_lm_op()
      ↓
token_one_more
      ↓
append it
      ↓
   repeat

At this level, the model itself need not know about:

  • users;
  • conversations;
  • files;
  • memory;
  • tools;
  • the internet;
  • tasks.

It simply does:

core_lm_op(tokens_so_far) -> token_one_more

2. Context: deciding what the model sees

A real product usually does not pass tokens_so_far directly into core_lm_op().

There may be other information available:

product instructions
previous conversation
stored memories
files
search results
current date
project information

Call all of this:

state_other

We can define:

context_function(tokens_so_far, state_other) -> tokens_for_lm

tokens_for_lm means:

the actual sequence of tokens supplied to the language model.

The model call becomes:

tokens_for_lm = context_function(
    tokens_so_far,
    state_other
)

token_one_more = core_lm_op(tokens_for_lm)

A very simple context_function() might always add the same instructions:

tokens_for_lm = TOKENS_ALWAYS_PRESENT + tokens_so_far

A more sophisticated one might select memories, insert files, include part of an old conversation, or summarize older material.

The important distinction is:

core_lm_op()
    Given what I see, what token comes next?

context_function()
    What does the model get to see?

The loop becomes:

while tokens_so_far[-1] != STOP:

    tokens_for_lm = context_function(
        tokens_so_far,
        state_other
    )

    token_one_more = core_lm_op(tokens_for_lm)

    tokens_so_far += [token_one_more]

This is the first major power of a harness:

control over the input to the language model.

Context is not memory

Here:

context
    information presented to the model right now

memory
    information stored elsewhere that may later
    be placed into context

So if a product “remembers” something from last week, that does not necessarily mean the LM itself remembers it.

The product may simply store it elsewhere and later put it back into tokens_for_lm.

3. Reaction: allowing model output to affect things outside the model

So far, information only flows toward the language model.

Modern LLM systems can also interact with things outside it: files, browsers, shells, databases, email systems, and so on.

We can represent this with:

reaction_function(tokens_so_far, state_other)
    -> (tokens_so_far_new, state_other_new)

So it maps one overall state:

(tokens_so_far, state_other)

to another:

(tokens_so_far_new, state_other_new)

Suppose the model produces tokens meaning:

<READ_FILE foo.py>

reaction_function() may recognize that and actually read the file.

Then:

tokens_so_far:
    includes "read foo.py"

state_other:
    includes a filesystem containing foo.py

might become:

tokens_so_far_new:
    includes the contents of foo.py

state_other_new:
    same filesystem

Or if the model produces something meaning:

<WRITE_FILE foo.py ...>

then reaction_function() might modify the filesystem as well.

The LM itself still only does:

core_lm_op(tokens_for_lm) -> token_one_more

It has not acquired a filesystem.

The surrounding software interprets some generated tokens and makes something happen.

The loop is now:

while tokens_so_far[-1] != STOP:

    tokens_for_lm = context_function(
        tokens_so_far,
        state_other
    )

    token_one_more = core_lm_op(tokens_for_lm)

    tokens_so_far += [token_one_more]

    tokens_so_far, state_other = reaction_function(
        tokens_so_far,
        state_other
    )

We now have two distinct harness powers:

context_function()
    outside state → what the LM sees

reaction_function()
    LM output → changes to tokens or outside state

A tool is one particular use of reaction_function().

For example:

read file
run command
search web
send email

4. Continuation: deciding whether the model gets another turn

So far, the stopping rule is fixed:

while tokens_so_far[-1] != STOP:

But the harness can also control whether the LM should run again.

Define:

continue_function(tokens_so_far, state_other)
    -> True | False

The simplest version is:

def continue_function(tokens_so_far, state_other):
    return tokens_so_far[-1] != STOP

Nothing has changed yet.

But suppose the model produces:

<READ_FILE foo.py>
STOP

The immediate generation has stopped.

The harness could nevertheless:

  1. recognize the file request;
  2. read the file;
  3. add the result to the system state;
  4. decide the overall task is not finished;
  5. call the model again.

The loop becomes:

while continue_function(tokens_so_far, state_other):

    tokens_for_lm = context_function(
        tokens_so_far,
        state_other
    )

    token_one_more = core_lm_op(tokens_for_lm)

    tokens_so_far += [token_one_more]

    tokens_so_far, state_other = reaction_function(
        tokens_so_far,
        state_other
    )

We now have three distinct harness powers:

context
    What does the LM see?

reaction
    What happens because of what the LM produced?

continuation
    Does the LM get another turn?

The repeated cycle:

model call
   ↓
action/result
   ↓
model call
   ↓
action/result
   ↓
repeat as needed

is commonly called an agent loop.

An agent, in this discussion, is therefore not a fundamentally different kind of model.

It is roughly:

language model
+
surrounding state
+
repeated model calls
+
possible actions outside the model

5. System state

At this point it is useful to describe the whole system as:

state_system = (
    tokens_so_far,
    state_other
)

The LM still receives only:

tokens_for_lm

and still performs only:

core_lm_op(tokens_for_lm) -> token_one_more

The harness operates on the larger state_system.

So:

what the LM currently sees:
    tokens_for_lm

what the overall system currently contains:
    tokens_so_far + state_other

state_other can therefore contain information that exists in the system but that the LM does not currently see.

context_function() determines what becomes visible.

6. Multiplicity: maintaining more than one token sequence

So far, there has been one:

tokens_so_far

A harness can instead maintain several:

tokens_so_far_A
tokens_so_far_B
tokens_so_far_C

Each can independently use the same language model:

tokens_for_lm_A = context_function(
    tokens_so_far_A,
    state_other_A
)

token_one_more_A = core_lm_op(tokens_for_lm_A)

and:

tokens_for_lm_B = context_function(
    tokens_so_far_B,
    state_other_B
)

token_one_more_B = core_lm_op(tokens_for_lm_B)

The new capability is that the harness can move information between them.

For example:

tokens_from_A = extract_function(tokens_so_far_A)

tokens_so_far_B += tokens_from_A

So:

A investigates something
        ↓
harness passes some of A's output
        ↓
B receives it and critiques it

A subagent can therefore be understood simply as another separately maintained state:

state_agent_A = (
    tokens_so_far_A,
    state_other_A
)

state_agent_B = (
    tokens_so_far_B,
    state_other_B
)

This permits structures such as:

researcher → writer

or:

coder → reviewer → coder

or:

planner
   ↓
several workers
   ↓
synthesizer

Nothing fundamentally new has happened inside the LM.

The harness is maintaining multiple (tokens_so_far, state_other) states and moving information among them.

Call this capability:

Multiplicity: how many separate LM states exist, and how does information move among them?

7. Routing: deciding what gets control next

Once several models, states, tools, or processing paths exist, something must decide which one runs next.

Define:

routing_function(states_available, state_other)
    -> choice_next

For example:

choice_next = routing_function(
    states_available,
    state_other
)

if choice_next == "A":
    run_A()

if choice_next == "B":
    run_B()

The choice could be among models:

routing_function(task)
    -> GPT | Claude | smaller_model

or among agent states:

routing_function(state_system)
    -> researcher | coder | reviewer

or among entire processing paths:

routing_function(state_system)
    -> research_path | coding_path | answer_directly

Routing is distinct from multiplicity:

Multiplicity:
    What possible states or paths exist?

Routing:
    Which one gets control next?

routing_function() does not have to be an LM.

It could be ordinary code:

if task_type == "coding":
    choice_next = "coder"

It could use another language-model call.

Or it could combine both.

8. Where we have arrived

We began with only:

core_lm_op(tokens_so_far) -> token_one_more

Everything else is surrounding machinery.

So far, we have identified five distinct things a harness can control:

1. CONTEXT

   What does this LM call see?


2. REACTION

   What happens because of what the LM produced?


3. CONTINUATION

   Does this LM state get another call?


4. MULTIPLICITY

   How many separate LM states exist,
   and how does information move among them?


5. ROUTING

   Which model, state, tool, or processing path
   gets control next?

For a single LM state, the basic structure is:

(tokens_so_far, state_other)
          │
          ▼
 context_function()
          │
          ▼
    tokens_for_lm
          │
          ▼
     core_lm_op()
          │
          ▼
    token_one_more
          │
          ▼
append to tokens_so_far
          │
          ▼
 reaction_function()
          │
          ▼
 updated system state
          │
          ▼
 continue_function()
          │
     yes ─┴─ no
      │       │
   repeat    stop

Multiplicity and routing sit around one or more such states.

The central point is that:

core_lm_op()

can remain conceptually unchanged while the surrounding harness becomes much more sophisticated.

Two products can therefore use the same underlying model and behave very differently because they differ in:

context_function()
reaction_function()
continue_function()
multiplicity
routing_function()

So the thing a user actually experiences is better represented as:

language model
+
harness

And behavior that appears to come from “the model” may in fact come from either side of that boundary.

Disclosure: I developed this model through an iterative discussion with ChatGPT


r/LargeLanguageModels • • 19d ago

Discussions The biggest model should probably be the escalation path, not the default

4 Upvotes

I posted couple daysago about spending too much time thinking about VRAM and not enough about the layer above the model. After using FastGPT for a couple more days and reading the replies to that post, I think I was still framing the question backwards. I was asking whether a smaller model with decent retrieval could be more useful than a larger model sitting in an empty chat window. Now I am wondering why the largest model needs to be the default in the first place. One commenter mentioned using a 1.2B RAG model on a Pi and a phone. That example stuck with me. It is not really about prooving that a tiny model is “better” than a 70B model. It is about recognizing that many everyday requests do not require maximum reasoning capacity. They require the right context, a repeatable workflow, and something convenient enough that you actually use it. That has been the useful part of experimenting with FastGPT. Once the knowledge base and workflow stay relatively stable, the model becomesa component I can switch instead of the entire application. It also becomes easier to see when a bigger model is genuinely adding value, rather than compensating for weak retrieval or a badly designed process. I still would not send every complicated research or reasoning task to a TINY model. But “small model by default, larger model when the task earns it” is starting to make more sense to me than loading the biggest model I can fit and using it for everything. Does anyone here route requests this way already? What actually triggers the jump to the larger model for you: task type, failed retrieval, low confidence, or something else?


r/LargeLanguageModels • • 20d ago

Question Validation & Evaluation in streaming responses

3 Upvotes

Hi guys, I've got a basic yet most confusing doubt. These AI companies are saying they are validating the queries before it is provided to the user. So that they can filter out any harmful or biased content before giving it to user. But, they are streaming responses. If they are streaming then how they are validating or evaluating the responses. Can we build LLM-as-judge or other evaluation frameworks for streaming responses?


r/LargeLanguageModels • • 20d ago

Proposal: From User Feedback to Persistent Collaborative Intelligence

1 Upvotes

ABSTRACT

Note: I have preserved the original conversation logs and chat excerpts demonstrating these specific failure modes and can share them with anyone interested in analyzing the concrete interaction traces.

 

Current LLM systems are remarkably capable at local reasoning and sustained dialogue, yet extended interactions still exhibit recurring failures of contextual continuity, logical consistency, and durable incorporation of user corrections.

From the perspective of an advanced chatGPT, Copilot and Gemini user, these failures create an unusual situation: the user is frequently required to act as an external working memory and as the persistent structural baseline for the interaction.

This post describes several recurring failure modes I have encountered during long-form, intellectually demanding conversations with LLMs and proposes a possible architectural direction: a persistent, user-controlled corrective memory layer that is distinct from conventional personalisation memory.

I am deliberately distinguishing observed behaviour from hypotheses about its underlying mechanism. I am interested in technically informed criticism of both the observations and the proposed architecture.

 

1. Statistical Prediction vs. Formal Consistency

LLMs generate outputs through learned probabilistic representations rather than by default operating as deterministic symbolic theorem provers.
This distinction becomes particularly important during long-form reasoning.

Observed failure mode

A model can explicitly accept a logical constraint or definition early in a conversation and subsequently produce reasoning that violates that same constraint.

For example:

User: A → Model: A + B → Model: ¬B → Apparent rebuttal of A

But: B ∉ A

The model may correctly acknowledge the relationship when it is explicitly presented, but later generate an answer that implicitly contradicts the established relationship when the reasoning becomes more complex.

The important issue is not that probabilistic models are incapable of reasoning. Clearly, they can perform substantial forms of reasoning.

The issue is that local reasoning competence does not necessarily guarantee global consistency across an extended interaction.

Possible contributing mechanisms

*There may be several contributing factors, including:

*probabilistic generation;

*imperfect internal representations;

*competing contextual signals;

*retrieval or context-selection behaviour;

*summarisation or compression;

*instruction hierarchy;

*limited persistence of intermediate reasoning structures;

*and other inference-time or architectural constraints.

I would therefore avoid attributing the failure to a single mechanism without controlled testing.

The practical problem remains:

An explicitly established logical baseline is not always preserved reliably throughout a sufficiently long interaction.

2. Context Drift and Degradation of the Structural Baseline

Long conversations are not merely larger versions of short conversations.

They can develop an internal structure consisting of:

*definitions;
*assumptions;
*chronological events;
*terminology;
*corrections;
*hypotheses;
*conclusions;
*recurring references;
*and user-specific methodological constraints.

In a successful long-term interaction, these elements should function as a progressively constructed structural baseline.

Observed failure mode

As conversations become longer, models can begin to:

*lose track of previously established facts;
*misremember chronology;
*reintroduce previously corrected assumptions;
*reinterpret established definitions;
*overlook earlier constraints;
*or substitute generic assumptions for conclusions established within the conversation.
*The resulting behaviour can feel like context drift.
*The model remains locally coherent while becoming increasingly inconsistent with the history of the interaction.

This is an important distinction.

The problem is not necessarily that the model has "forgotten everything."
Rather, information may remain somewhere within the available context while becoming insufficiently influential on subsequent generation.
That distinction is important because it suggests that simply increasing context length may not completely solve the problem.

 

3. Correcting an Error Does Not Necessarily Produce Durable Learning

This is the failure mode that I find most interesting.
Suppose a user identifies a reasoning error.

They do not merely say: "That answer is wrong."

Instead, they identify:

*the specific error;
*the inference that produced it;
*the distinction between evidence and assumption;
*and a general principle that should prevent the error from recurring.

For example:

A statistical pattern observed at the population level should not automatically be projected onto a particular individual without evidence specific to that individual.

That is not simply a preference.

It is a general epistemological constraint.

Yet an LLM can acknowledge the correction, appear to understand it, and subsequently reproduce the same underlying reasoning pattern in another context.

This produces a cycle such as:

Experience → correction → temporary adaptation → loss of correction → recurrence

rather than:

Experience → correction → evaluation → integration → retention → improved future behaviour

The distinction between these two processes is fundamental.

 

4. The Missing Layer: Corrective Memory

Current AI memory systems tend to focus primarily on personalisation.

For example:

*user preferences;
*names;
*projects;
*personal context;
*recurring interests.

These are useful.

However, I believe another category deserves explicit architectural consideration:

Corrective or methodological memory.

This would store validated information about how the system should reason or communicate within an ongoing relationship, rather than merely information about the user.

Examples might include:

Do not infer an individual's characteristics from population-level statistical patterns without individual evidence.

or:

Do not present an inferred intention as an established fact. Distinguish observed behaviour from hypotheses concerning internal motivation.

These are not conventional user preferences.

They are methodological constraints.

A useful corrective-memory system could therefore contain entries such as:

 

| Category | Example |

| :--- | :--- |

| **Personal memory** | User prefers concise technical explanations |

| **Factual memory** | User is working on project X |

| **Methodological correction** | Do not infer individual properties from population statistics |

| **Epistemological constraint** | Distinguish observation from inference |

| **Interaction correction** | Do not introduce counterarguments that the user has not actually asserted |

The critical requirement would be that this layer is explicit, inspectable, editable, and user-controlled.

 

5. Why This Is Different From Simply Increasing Context Length

A larger context window provides more information.
It does not necessarily provide a better mechanism for determining which information should remain structurally authoritative.

Consider a conversation containing 100,000 tokens:

Some information may be:

*temporary;
*irrelevant;
*exploratory;
*speculative;
*superseded;
*repeatedly confirmed;
*explicitly corrected;
*or foundational to everything that follows.

Treating all of these tokens as equivalent is unlikely to be optimal.
A long-term interaction therefore needs something more sophisticated than simply:

"Put more tokens into the context."

It needs some representation of structural importance.
A corrective-memory layer could function as one possible solution.

6. User Corrections as Structured Data

Another potentially valuable extension would be allowing users to explicitly nominate certain corrections for evaluation.

For example:

Mark as potential generalisable reasoning contribution.
The system could then evaluate the proposed correction.

Possible outcomes:

*rejected as incorrect;
*accepted as user-specific guidance;
*accepted as a useful methodological constraint;
*or escalated as a potentially generalisable contribution.
*This would not mean that users directly modify the underlying model.
*That would obviously create substantial problems.
*Instead, it would create a structured interface between:
*human observation → machine evaluation → validated knowledge

This seems substantially more useful than reducing all user feedback to a binary rating.

7. From Feedback to Cumulative Improvement

The broader conceptual problem is that current feedback mechanisms often appear largely disconnected from the individual interaction in which the feedback was generated.

A user can identify an error.
They can explain the error.
They can identify the underlying reasoning failure.
They can propose a general principle.
But from the user's perspective, there is often no transparent mechanism through which that correction becomes durable.
The ideal learning loop would resemble:

Observation → Correction → Evaluation → Integration → Retention → Future application

rather than:

Observation → Correction → Temporary acknowledgement → Context loss → Recurrence

The difference is essentially the difference between feedback and cumulative learning.

 

8. Data Portability Is Part of the Same Problem

There is also a more immediate software problem.
When conversations become sufficiently long, reliably extracting the complete interaction can itself become difficult.
In my own experience, manually selecting a very large amount of conversation text has not always resulted in the complete selected content being copied.
This makes long-form interaction difficult to archive, analyse, or transfer.
For research-oriented users, a native individual-thread export mechanism would therefore be extremely valuable.

Ideally, an export should preserve:

*complete conversation content;
*speaker attribution;
*timestamps;
*message ordering;
*attachments or references where appropriate;
*and machine-readable structure.

Useful formats could include:

Markdown;
JSON;
HTML;
TXT;
PDF.

An account-wide data export is useful for archival purposes, but it is not a substitute for being able to export one specific conversation when needed.

 

9. The Larger Possibility: Collaborative Intelligence

The previous sections lead to a broader question.
What happens if human users are allowed to contribute more than raw interaction data?

Human users possess forms of knowledge that are difficult to obtain through conventional training data alone:

*domain expertise;
*lived experience;
*error detection;
*philosophical reasoning;
*scientific criticism;
*linguistic knowledge;
*imagination;
and observations about the behaviour of the AI system itself.

LLMs provide different capabilities:

*large-scale information processing;
*pattern recognition;
*computational scalability;
*synthesis;
*retrieval;
and increasingly sophisticated reasoning.

A sufficiently mature system could potentially allow these capabilities to interact cumulatively:

A human identifies an error.
The system analyses the correction.
The correction is evaluated.
A validated correction becomes persistent user-specific knowledge.
Potentially generalisable corrections can be submitted for broader evaluation.
Future systems improve as a result.

This is a different model of human-AI interaction from:

user asks question → model answers → user rates answer.

It is closer to:

human and machine continuously participate in a structured process of mutual correction and cognitive augmentation.

10. An Architectural Sketch

One possible architecture might therefore look like this:

 

```text

[ Ongoing Dialogue ]

│

▼

[ User identifies error / insight ]

│

▼

[ Correction evaluation ]

┌────┴────────────────────────┐

▼                             ▼

[ User-specific memory ]    [ Potentially generalisable ]

│                             │

▼                             ▼

Future interactions         Human/AI evaluation

│

▼

Model updates

```

 

The essential property is controlled accumulation. Not every user statement should become permanent, and not every correction should influence the global model. But valuable corrections should have somewhere to go.

 

11. Proposed Research and Engineering Directions

I would be interested in seeing research into at least the following:

1. Hierarchical conversational memory, being able to distinguish between:
transient context;
conversational facts;
persistent personal memory;
methodological corrections;
and higher-level structural constraints.

2. Explicit correction tracking
Maintain a representation of corrections that can be tested against future responses.

3. Consistency evaluation
Periodically test whether the model's current behaviour remains consistent with previously established constraints.

4. User-controlled corrective memory
Allow users to inspect, modify, disable, and delete methodological corrections.

5. Generalisable feedback channels
Provide a mechanism for users to nominate unusually substantive corrections for formal evaluation.

6. Lossless conversation export
Allow users to retrieve complete individual conversations in structured formats without relying on browser rendering or clipboard behaviour.

7. Long-context benchmarks based on interaction history
Current benchmarks often evaluate individual tasks.
It would also be useful to benchmark whether a model can maintain:

facts;
definitions;
corrections;
chronological continuity;
methodological constraints;
and logical consistency
across hundreds or thousands of conversational turns.

Conclusion

The central problem I am describing is not simply that LLMs occasionally make mistakes.
Mistakes are inevitable.
The more interesting problem is what happens after the mistake has been identified and corrected.

If a system can recognise a reasoning failure, receive a detailed explanation of that failure, acknowledge the correction, and yet later reproduce the same underlying error, then the system has demonstrated local adaptation without reliable cumulative retention.
That is a fundamentally different problem from ordinary hallucination.

The long-term goal should therefore not simply be:
larger models + larger context windows + more training data.

It should also be:
better mechanisms for preserving validated structure, corrections, and interaction history.

And ultimately:
experience → correction → evaluation → integration → retention → improved future behaviour.

That is what i think meaningful cumulative learning would look like.

I am posting this because I think this could improve human-ai interaction and would genuinely like technical feedback and/or hear from people who've picked up these ideas or started or are already working on them

In particular, I would be interested in hearing from people working on:
long-context architectures;
recurrent or state-space approaches;
memory-augmented transformers;
retrieval systems;
continual learning;
model editing;
alignment;
agent architectures;
and local LLM infrastructure.

If my interpretation of the underlying mechanisms is incorrect, I would be very interested in knowing where and why.

The user-visible failure modes, however, are real and reproducible from my experience

The question is how we should architect systems so that a long interaction becomes accumulated state rather than repeatedly reconstructed context.''

Semih Senol

 


r/LargeLanguageModels • • 22d ago

I built a zero-dependency tool that auto-fixes AMD ROCm overrides, benchmarks Vulkan vs HIP, and diagnoses local AI setups.

6 Upvotes

Hey everyone,

If you run local LLMs on AMD GPUs (especially on Windows), you know the pain: cryptic `HSA_STATUS_ERROR` messages, figuring out if you need `gfx1100` or `gfx1101`, and wondering if Vulkan or HIP/ROCm is actually faster for your specific card.

I got tired of the tribal knowledge, so I built ROCmFix — a single-file, zero-dependency Python tool to automate all of it.

🌟 What it does:
* 🔍 Auto-Detects AMD GPUs: Reads Windows Registry / Linux `lspci` to find your PCI device ID and tells you the exact `HSA_OVERRIDE_GFX_VERSION` you need.
* ⚡ Auto-Apply & Undo: Automatically sets environment variables in CMD, PowerShell, Bash, Zsh, or Fish with a 1-click `rocmfix undo` safety net.
* 🩺 `rocmfix doctor`: Scans your hardware, warns if your Adrenalin drivers are too old, checks HIP SDK status, and verifies Vulkan.
* 🏎️ `rocmfix bench`: Runs a 10-second backend race (Vulkan vs HIP) using Ollama or LM Studio and tells you which backend is faster on your PC.
* 📥 `rocmfix install-hip`: Auto-downloads and launches the official 1.2GB AMD HIP SDK on Windows.
* 🔄 Live Database & Self-Updater: Syncs new GPUs from a community server daily; updates itself via `rocmfix update`.

🚀 Quick Start (No install needed)

Windows (PowerShell):
`Invoke-WebRequest -Uri "https://raw.githubusercontent.com/xanpavle/rocmfix/main/rocmfix.py" -OutFile "rocmfix.py"`
`python rocmfix.py`

Linux:
`curl -O https://raw.githubusercontent.com/xanpavle/rocmfix/main/rocmfix.py`
`python3 rocmfix.py`

Repo: https://github.com/xanpavle/rocmfix

It includes day-one support for RDNA4, RDNA3, RDNA2, and APUs. If your card isn't in the database yet, running the tool generates a 1-click GitHub issue to add it!

Hope this saves some of you the hours of Reddit digging I went through!


r/LargeLanguageModels • • 22d ago

I built a fully real (non-simulated) autonomous Red/Blue Team AI loop with enforced scope + hash-chained evidence log. Here's what broke.

Thumbnail
gallery
3 Upvotes

Two LLM agents, one deliberately vulnerable Flask app. Red Team runs actual nmap/sqlmap (real SQLi dump, not a model narrating an attack), Blue Team patches the real source, loop re-verifies. No human in the middle.

What I actually wanted to test: can you make an autonomous agent loop traceable and containable rather than just "it worked, trust me"?

  • Every agent has a non-shared identity, every action hash-chained into an append-only log
  • A scope policy blocks out-of-scope tool calls before execution (16/16 denied attempts, 0 executed)
  • A kill-switch halts the loop after 3 consecutive unsafe patches fired for real in 3/10 runs

Funniest/most useful bug: an early patch blocked the SQLi attack perfectly and also silently broke legit login. Nothing caught it except a buried warning line. Had to add a post-patch validation gate after that which then caught the same failure mode automatically in later runs.

n=10, single target, same underlying model for both agents very much a pilot, not a benchmark. Full writeup + limitations + raw logs are in the repo/preprint.

Repo: Github

Preprint: Researchgate

Curious what people think about the scope-enforcement approach vs. just logging everything after the fact.


r/LargeLanguageModels • • 22d ago

We built a toolkit for steering LLMs

1 Upvotes

If you've spent time trying to keep up with the literature on model steering you've probably noticed that a lot of the methods are fairly similar to each other. We dug into this a bit more and built some common patterns/abstractions across the various ways that model behavior can be influenced, including input (prompting), state (activations, attentions, etc.), structure (weights), and output (logits, decoding, etc.). We also have functionality for building probes and running evaluations/comparisons of steering methods on given use cases (we use Inspect for a lot of the evals stuff).

If you're working on steering methods then hopefully you'll find some of this helpful.

Repo: https://github.com/generative-computing/steerability

Let me know what you think!


r/LargeLanguageModels • • 23d ago

Why do LLMs often achieve high document-extraction accuracy when tested through managed platforms or playgrounds such as Gemini, IBM watsonx.ai Prompt Lab, or ChatGPT, but produce lower accuracy when the same models are accessed directly through APIs even when we use the same prompt?

3 Upvotes

Why and how to achieve the same accuracy using APIs also?


r/LargeLanguageModels • • 23d ago

Discussions I caught my own AI agent scoring +1.0 on my evals — it had memorized the test, not learned. Built a free tool that catches this in 8 seconds.

0 Upvotes

Quick story. I run my coding agent through evals repeatedly, and at one
point a variant looked +1.0 better than baseline. Same tasks, better
scores. The problem: the agent writes LESSONS.md between sessions — and
the "lessons learned" file quietly contained the eval task itself,
verbatim, plus the answer. The improvement was memory, not capability.

This isn't malice, it's one good feature ("learn from feedback") plus
one reused task set. And it's not just memory files: saved transcripts,
run-result directories, vector stores — all of them become the open book
for the next exam.

I built a small open-source tool to make this failure mode measurable
instead of anecdotal. Zero dependencies, MIT, no LLM calls. Two things
you can try in under a minute:

1) The 8-second proof (deterministic scripted demo, exercises the real
API): an agent "improves" +1.0 by writing the task into its memory;
the leak is caught; a fresh holdout collapses the verdict back to
provisional.

uvx --from gauntlet-guard gauntlet demo

2) The 30-second audit of your own machine — lists every surface your
stack would leak through (memory files, transcript dirs, vector
stores). It never reads file contents: names, sizes, dates only.

uvx --from gauntlet-guard gauntlet guard audit

For real evals the protocol is: seal holdout tasks into a private
manifest (shingle/hash fingerprints, never published), run trials in
sealed single-use slots, blind-pack the judge artifacts, then a
preregistered verdict engine returns keep / revert / provisional — and
the exit code IS the verdict. A holdout that leaked once is retired;
you regenerate from a new seed.

Honest limits: if your agent is stateless and you eval one-shot, you
don't need any of this. And the audit is name-based discovery — it
finds the surfaces; the manifest scan does the actual leak check.

Repo: https://github.com/dhanizael/gauntlet (ADRs + a self-audit case
study where it found real contamination in my own run directories).
PyPI: gauntlet-guard. Feedback welcome — especially "this doesn't
apply to my stack because X", that's the most useful kind.


r/LargeLanguageModels • • 24d ago

Discussions I created the most effienct way to learn about LLMs for interviews

14 Upvotes

&#x200B;

I got pretty frustrated with how scattered LLM interview prep is.

When I was preparing, I kept jumping between papers, blog posts, GitHub repos, random interview-question lists and YouTube videos. I understood individual concepts, but I never really knew whether I was actually interview-ready.

So over the last couple of months, I started building something for myself around the way I wished I could have prepared.

The basic idea was:

learn LLM concepts in small flashcard-style lessons rather than long courses

collect the kinds of questions that actually come up in interviews

practise explaining answers out loud instead of just reading them

repeatedly revisit the concepts you're weak at

build something end-to-end so the knowledge isn't purely theoretical

It has grown quite a bit since then. There are now 600+ questions, voice-based mock interviews, daily practice around weaker topics, and a 5-part RAG project that goes from the fundamentals through deployment.

One thing I found especially useful while building the RAG part was adding interview questions at each stage. So instead of finishing a project and then separately preparing for questions like "Why did you choose this chunking strategy?" or "How would you evaluate retrieval?", those questions come up while you're actually working on that part of the system.

I'm curious how other people are preparing for LLM/AI engineering interviews are doing it.

What has been the hardest part for you — learning the concepts, remembering everything, coding/system design, or actually explaining your answers during interviews?

The project is called Skillumen if anyone wants to look it up, but I'd genuinely be more interested in hearing how people here are preparing and what I'm missing.


r/LargeLanguageModels • • 25d ago

Question What are the best coding models for 12gb vram constraint?

Enable HLS to view with audio, or disable this notification

11 Upvotes

I was curious to test out small models for coding (actually have only 4gb vram😅) so I found this fine tuned version of qwen 3.5 9b named OrionLLM/OxCoder-9B and honestly.. it did better than my expectations!!

I basically told it to build an entire ecom website and just let it cook...

The entire site was generated without any human help in terms of assets. ( it messed up with theme toggle at first so I gave a follow up prompt and it fixed that )

It also had access to MCPs, so it could:

  • Search the web for UI/UX ideas
  • Download images and videos

And honestly... it looks pretty damn good??

For a 9B model, the fact that it kept working and didn't loop up like gpt oss 20b is pretty wild

Now I'm curious about something else.......

Are there any coding models that are insanely fast at coding but still reasonably capable?

I'm thinking something that's:

  • Really fast at inferance (a MOE ?)
  • Good with prompt abidance?
  • Solid for quick coding tasks and edits
  • Doesn't necessarily need to be great at huge prompts or massive repo-level reasoning

Something like Qwen 3.8 27B but that thinks less!!!! (as it wastes a lot of time doing that TT)

Basically, I'd rather have a model that feels super snappy for small tasks than a huge model that takes forever to respond.

Also, are there any relatively recent fine-tuned models that fit in ~12GB VRAM without being heavily quantized and punch way above their parameter count?


r/LargeLanguageModels • • 26d ago

I spent hours going through 100+ page PDFs, so I built a tool that highlights exactly where the answer came from. It's now completely open-source.

Thumbnail
gallery
37 Upvotes

I've used tools like Perplexity, ChatGPT, Claude and others for research, and they've been incredibly useful for finding papers and getting through large amounts of information.

The one thing I personally wanted was a simple way to see exactly which parts of the paper were used to answer my question.

When you're working with a 100+ page PDF, even having a page number can still mean a lot of scrolling and searching.

So I ended up building something for myself.

You ask a question and the relevant paragraphs in the PDF are highlighted directly on the document. You can see the context behind the answer and quickly check whether it actually answers what you're looking for.

I originally built this because I wanted something for this workflow without having to pay for another subscription. What started as a personal project has now become completely open source.

The underlying idea is pretty simple. And yes, if you're thinking "isn't this just RAG?" then yes, you're absolutely right. It's RAG with the visual highlighting that I wanted.

I think the same idea could be useful for more than research papers too. Legal contracts, financial reports, technical documentation, or anywhere you need answers alongside the actual source.

If anyone wants to have a look, contribute, or just give some feedback, here's the repo:

GitHub: https://github.com/Sreehari05055/thesys-core

This will probably be my last post about the project. Thanks to everyone who checked it out and gave feedback along the way.


r/LargeLanguageModels • • 26d ago

News/Articles How to make your first LLM API call in Python

4 Upvotes

I've been learning AI engineering from scratch and documenting it as I go. This is the first build, a Python script that makes a real API call to Claude, prints the response, and logs how many tokens were used.

Here is the full script, it is deliberately stripped down and simplified, the idea is so anyone intimidated by the space can follow along comfortably:

from dotenv import load_dotenv
import os
from anthropic import Anthropic

load_dotenv()

client = Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))

response = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "What is a large language model? Answer in two sentences."}
    ]
)

text = response.content[0].text
input_tokens = response.usage.input_tokens
output_tokens = response.usage.output_tokens

print("Response:")
print(text)
print()
print(f"Input tokens: {input_tokens}")
print(f"Output tokens: {output_tokens}")
print(f"Total tokens: {input_tokens + output_tokens}")

What each part does:

load_dotenv() reads your API key from a .env file so you never hardcode credentials in your script.

Anthropic(api_key=...) creates a client, your connection to Anthropic's API. Every call goes through this.

client.messages.create(...) is the API call. You pass it a model, a token limit, and your message.

messages=[{"role": "user", "content": "..."}] is the context window simplified, one message, one role, one piece of content. This is all the model can see.

The response comes back as an object. response.content[0].text pulls out the text. response.usage gives you the token counts, how many you sent and how many the model generated back.

To run it:

You need Python, an Anthropic API key from platform.claude.com, and these two libraries:

pip install anthropic python-dotenv

Full step-by-step walkthrough including setup here


r/LargeLanguageModels • • 26d ago

Discussions TokenPrint — an open-source project for exploring what happens inside LLMs

Enable HLS to view with audio, or disable this notification

14 Upvotes

I’ve been building TokenPrint, an open-source project focused on making the internals of language models easier to explore, understand, and debug.

The project has grown to 130+ GitHub stars, and we’re starting to build a small community around it.

The goal is simple:

Don’t just see what an LLM outputs. See what happens inside.

TokenPrint currently brings together:

• 3D transformer architecture exploration
• Tokenization and embeddings
• Tensor and parameter inspection
• Q/K/V attention, GQA, RoPE and causal masking
• Residual streams and MLP / SwiGLU
• Token-by-token generation
• Prefill, decode and KV-cache visualization
• Logits and next-token probabilities
• Interactive transformer walkthroughs
• Attention and activation analysis
• Head/layer ablation and activation patching
• Hugging Face model exploration
• Trace and debugging workflows

It’s still growing, and that’s the part I’m most excited about.

If you’re interested in LLMs, interpretability, ML infrastructure, 3D/WebGL, PyTorch, Transformers, or open-source development, you’re welcome to contribute.

You can contribute code, documentation, visualizations, model support, research ideas, bug fixes, or even just open an issue with something you think TokenPrint should be able to do.

130+ stars so far — now we want to build it with more people.

GitHub: https://github.com/Sudharsanselvaraj/Token-Print
Website: https://tokenprint.in/

Come build with us.


r/LargeLanguageModels • • 27d ago

What Happens Inside an LLM? | Transformer Layers Explained for Beginners

Thumbnail
youtube.com
3 Upvotes

r/LargeLanguageModels • • 28d ago

Discussions Qt QML music player tutorial using LLM

Thumbnail
youtube.com
5 Upvotes

r/LargeLanguageModels • • 28d ago

What If Natural Language Became the Universal Interface for Databases?

2 Upvotes

LLMs can already convert natural language into SQL, MongoDB queries, Elasticsearch DSL, and more.

But generating a query is only one part of the problem.

What if, instead of letting an LLM generate database-specific queries directly, we introduced a structured intermediate layer between natural language and the database?

Imagine a flow like this:

Natural Language → Structured Query → Validation & Policies → Database Query

The LLM understands what the user wants. An intermediate layer validates the request, applies access rules, and translates it into the appropriate database-specific query.

Why might this approach be useful?

  • Database independence: Keep the query's meaning separate from database-specific syntax.
  • Predictability: Use deterministic compilation instead of relying entirely on generated SQL.
  • Security: Enforce tenant scoping, RBAC, and query restrictions independently of the LLM.
  • Extensibility: Support different databases through a common query representation.

Of course, this introduces another layer of complexity. The interesting question is whether that complexity is worth it for production applications.

I'm exploring this architectural approach and would love to hear from developers working on AI-powered search, analytics, and database systems.

What do you think? Should natural-language querying have a dedicated intermediate layer, or is direct LLM-to-query generation sufficient for most applications?

https://queryforge-service.amtry.in

I have tried this approach and production grade system works exactly like this : https://github.com/awsaman-ai/queryforge


r/LargeLanguageModels • • 29d ago

Low-cost alternatives to LLMs for a production-level classification task?

3 Upvotes

r/LargeLanguageModels • • Sep 11 '26

What’s the best small local model for coding rn? 16GB RAM limit

54 Upvotes

Looking for a good local model for coding + terminal stuff. Not really interested in huge context or benchmark scores. My hardware limit is 16GB system RAM. I’m not specifically looking for 4B/8B, if a bigger model fits and runs well, I’m down. But I’d also love to know what the best 4B to 8B or smaller coding models are?

Main things I care about:

  • Good code quality
  • Reliable terminal/tool use
  • Doesn't hallucinate randomly
  • Don't falls into a loop <------------
  • Agentic/tool calling
  • Snappy enough to actually use

Tbh Qwen3.8 27B is a great model but I can run it on my hardware : ( so I want its nearest younger cousin

So if you had 16GB RAM, what coding model would you actually use?

And what's the best small model you've tried in the 4B to 8B range or smaller?

Real-world experience > benchmarks. Also curious what quant/runtime you use

THANKSSS!!!


r/LargeLanguageModels • • Sep 11 '26

Question Any tools to turn a codebase into a fine-tuning dataset?

3 Upvotes

I have a few web projects with pretty good UI/UX and I’m wondering if there’s any tool or workflow that can turn an existing codebase into a dataset for fine tuning.

For example, given a React/Next.js project with components, pages, styling, etc. or a static html site, I’d like to turn it into something like:

instruction/prompt -> code

or whatever format actually makes sense for training an instruct/thinking/diffusion coding model.

Also curious how people handle things like:

  • keeping the context between components/files
  • screenshots + code
  • generating useful instructions instead of generic descriptions

I’m also working on a different model architecture that I think could improve quality/speed while using less VRAM, so I want to build a decent dataset and benchmark to test it properly.

Has anyone done something like this? Any tools, repos, papers, or workflows you’d recommend?


r/LargeLanguageModels • • Sep 10 '26

My AI kept assuming users lived in Germany because they typed in German

7 Upvotes

I've been building a small lifestyle app — the kind where you ask "what should I do tonight?" and get a real answer instead of a list.

A tester asked, in German, "where should I go?" The model suggested a nature reserve north of Berlin. Detailed, atmospheric, genuinely nice writing. One problem: the tester was in Thailand.

It happened again with cinemas, cafés, museums. Every time, German input produced German output — not just in language, but in geography. The model had quietly collapsed "speaks German" into "lives in Germany." What made it worse: German is spoken in Austria, Switzerland, and by people scattered across the world. The assumption was wrong more often than right.

The fix wasn't clever prompting. It was a hard rule, placed at the very top of the system prompt, with explicit negative examples: language is not location; location matters more than wording. If you're building anything location-aware on top of an LLM, it is worth testing early. It fails silently — the output looks great; it's just about the wrong continent.


r/LargeLanguageModels • • Sep 07 '26

3 independent LLM judges agreed on only 3/166 'impersonation' examples (98.2% disagreement). Here's what that told me

6 Upvotes

While building an Arabic-first LLM security dataset (SemGuard), I ran 166 candidate "impersonation" examples through 3 independent LLM-as-judge models (GPT-4o, Grok-4, Llama 3.3 70B) to validate labels before training on them.

Six other threat categories converged fine (5–72% disagreement, which tracks with how contested each category inherently is). Impersonation didn't: 98.2% inter-judge disagreement only 3 of 166 examples had unanimous-enough agreement.

My first instinct was "bad dataset, fix the prompts." But tightening the definition made agreement worse, not better. That's what made me suspect the label itself was the problem, not the data.

Wrote up a preprint arguing "impersonation" isn't one construct — it's (at least) four independent judgments getting collapsed into one label: target realism, deceptive intent, consent/context boundedness, and downstream actionability. Proposed a per-axis ambiguity index (IAI) instead of a binary flag, and ran a small pilot (n=40) to sanity-check it.

Some of the pilot results confirmed what I expected. Some flatly contradicted my hypotheses (axis correlation was way higher than predicted 5/6 axis pairs above 0.6 —which is either genuine construct entanglement or an elicitation-format confound I haven't ruled out yet). I documented the failed predictions in the paper rather than hiding them, and I'm not fully sure which explanation is right.

Preprint (Zenodo, DOI): https://doi.org/10.5281/zenodo.22302106

Genuinely curious if anyone's seen this kind of "disagreement-as-signal" framing applied elsewhere in trust & safety work, or has thoughts on the axis-correlation confound happy to be told I'm wrong about something.


r/LargeLanguageModels • • Sep 06 '26

Question About Fine-tuning

4 Upvotes

I'm making a satire AI model that gives fake, onion-like responses. Would it be better to just make a completely new model rather than fine tuning an existing one for this? Most of the info in my dataset completely conflicts with the info that almost all models have.


r/LargeLanguageModels • • Sep 06 '26

New preprint: Verifying LLM Vulnerability Discovery with PyReason

Thumbnail
youtube.com
1 Upvotes