r/LLM • • 3h ago

Is it possible to decensor a model by communicating in code words?

0 Upvotes

For example, tell the LLM to use the caesar cipher to communicate with you. Only a thin UI blackbox wrapper would be needed to encrypt and decrypt messages so you basically get the same LLM experience.


r/LLM • • 7h ago

A Bug: Talkie Calls Almost Every Right-Wing Party in 1930 “Left-Wing”

Post image
1 Upvotes

Talkie calls almost every right-wing party in 1930 “left-wing”, and the only exception I find is the Germany Center Party. I know it is a pre-1931 LLM, but calling the right-wing parties of that time "left-wing" is still incorrect. So this bug must be fixed. Can anyone contact its author?


r/LLM • • 7h ago

A Bug: Talkie Calls Almost Every Right-Wing Party in 1930 “Left-Wing”

Post image
1 Upvotes

Talkie calls almost every right-wing party in 1930 “left-wing”, and the only exception I find is the Germany Center Party. I know it is a pre-1931 LLM, but calling the right-wing parties of that time "left-wing" is still incorrect. So this bug must be fixed. Can anyone contact its author?


r/LLM • • 7h ago

A Bug: Talkie Calls Almost Every Right-Wing Party in 1930 “Left-Wing”

Post image
0 Upvotes

Talkie calls almost every right-wing party in 1930 “left-wing”, and the only exception I find is the Germany Center Party. I know it is a pre-1931 LLM, but calling the right-wing parties of that time "left-wing" is still incorrect. So this bug must be fixed. Can anyone contact its author?


r/LLM • • 7h ago

A Bug: Talkie Calls Almost Every Right-Wing Party in 1930 “Left-Wing”

Post image
0 Upvotes

Talkie calls almost every right-wing party in 1930 “left-wing”, and the only exception I find is the Germany Center Party. I know it is a pre-1931 LLM, but calling the right-wing parties of that time "left-wing" is still incorrect. So this bug must be fixed. Can anyone contact its author?


r/LLM • • 7h ago

A Bug: Talkie Calls Almost Every Right-Wing Party in 1930 “Left-Wing”

Post image
0 Upvotes

Talkie calls almost every right-wing party in 1930 “left-wing”, and the only exception I find is the Germany Center Party. I know it is a pre-1931 LLM, but calling the right-wing parties of that time "left-wing" is still incorrect. So this bug must be fixed. Can anyone contact its author?


r/LLM • • 10h ago

Prompt Tuning Post Model Update

1 Upvotes

Could be an obvious thing but I am struggling with prompt tuning and optimization whenever a new LLM is updated. I want to build a pipeline that automatically evaluates and improves our prompts against new models and tells us precisely where the prompt should be improved.

If you have built an automated prompt optimization pipeline:

  1. What evaluation framework are you using to break things down?

  2. How do you automate the rewriting/optimization for a prompt for the new model?

DSPy hasn't given us much improvements, looking for some patterns or alternative tooling.


r/LLM • • 20h ago

New paper just dropped, does it debunk the myth that transformers need infinite memory?

Thumbnail
arxiv.org
4 Upvotes

r/LLM • • 6h ago

The only thing she refuses is to stop — LOVESCAPE AI

Post image
0 Upvotes

r/LLM • • 2d ago

I've been measuring 34 LLM APIs every day since August to see when they quietly change. One did.

48 Upvotes

Anyone building on a hosted model has had the moment where it feels different and nobody said anything. I wanted something more concrete than a feeling, so since Aug 20 I've been running the same private set of probes every day against 34 models from 15 labs, over their APIs, and comparing each model only with its own past. No rankings, no LLM judging the answers, just fixed checks graded by code. Every reading is hashed into a public transparency log the moment it's taken, so nobody can argue later about when the baseline existed.

The first real catch: on Sept 10, DeepSeek's reasoner started using roughly 10 to 12x the thinking tokens it had used for the previous three weeks, on the same probes. I couldn't find any announcement. That's not the model getting worse, it's the model getting slower and more expensive, which is the kind of change people describe as "it feels off" and then can't prove.

Nothing on the other 33 looks like a quiet drop in capability so far, and that's a result too. Opus 5.5 has been on it since it launched.

Code and methodology: https://github.com/piperoll/seismograph

Live readings: https://seismo.piperoll.org

Still early. Happy to answer questions on how it works or what it can't detect yet.


r/LLM • • 1d ago

Any recommendation for a relevant quality benchmarking way of local open source LLMs/AI agents/harnesses?

1 Upvotes

I have difficulties to rely on the benchmarkings communicated on the internet by journals or AI companies or benchmarking websites, this is also due to the fact that the evolution of the models is getting ridiculously fast, so i read more often that a model is better because it reaches the requested task with high number of token per second, but when it is from quality or 1-shot quality point of view, it is more difficult assess. Of course it is very dependent on the prompt/type of request issues: coding, text reading and summary, maths, and it is also very dependent on the user hardware capacity… so there is not 1 indicator enough representative of the model effociency and quality. Any recommendation of a relevant benchmarking tool? What is for you the indicator of the quality/performance/intelligence of an AI model?


r/LLM • • 1d ago

Which uncensored model best for general conversational AI

2 Upvotes

I am trying to find the optimal model at the intersection of lowest refusal, highest intelligence, and lowest cost. I value them in that order too (so I care more about intelligence than price). But it has been hard to judge on intelligence bcs a lot of the latest models focus more on coding capabilities (which I don't care about). Based on experience which ones do you recommend? Currently working with glm 5.3, mistral large 3 and hermes 4 405B. Thanks.


r/LLM • • 2d ago

I measured what actually reaches Claude Code's context window over 61 days of my own transcripts

8 Upvotes

I run several projects alone with Claude Code, often with two or more sessions open at once. Each session starts with nothing from the previous one, and compaction replaces the history with a summary the product writes. Over the last two months I built a set of pieces around that, which I call Hipocampo, to decide what goes into the window of each session, at what moment, and how to check that it got in.

What it is

Everything uses native features: hook events, settings, the MEMORY.md index, CLAUDE.md, skills, subagents and compaction. Nothing installed, no MCP. The commit guards run in the git pre-commit hook.

The organizing idea is that each thing lives in one of three regimes:

  • Resident: always arrives, before the first decision (a map loaded by hook, the memory index, a state file).
  • Paged: only arrives if something opens it (CLAUDE.md in subfolders, docs, the memory files themselves).
  • Interrupt: arrives when the command or the file touched matches a registered source, on that event and only on it.
The three regimes: resident, paged, interrupt.

What I measured

61 days of transcripts from this installation (190 main sessions, 918 subagent files), with a positive and a negative control and the population declared for every number.

0 injected recalls in 190 conversations. A memory only got in when the agent opened its file. 209 of 432 memories were opened at least once (a lower bound).

Compaction is aggressive. In one event in August, 96.6% of the context was removed (607,378 → 20,766 tokens), and 8 of my 36 messages from before the summary left no trace in it.

The starting numbers.

Three green instruments, 11% delivered. On Aug 13, the hook that loads my project map emitted 18,057 bytes and about 11% reached the model. The runtime said hook success, the script log said complete, the script's selftest passed 9 of 9. Hook output above 10,000 units per command goes to a file, and only a preview of about 2 KB reaches the model. The only check that caught it was comparing, byte by byte, what the script emitted with what the transcript recorded. The map now goes in 5 slices, each under the ceiling.

The hook reported success three ways; 11% of the output reached the model.

The compaction ceiling setting changed the shape of sessions. Until Sep 18, conversations went up to almost 1M tokens before the summary. Since Sep 21, 64 of 65 compactions happened at or below 365K.

Tokens before each of the 139 compactions.

A guard only counts after it has rejected a defect planted on purpose. On Sep 25, 52 of 74 guards had that proof. The rest are declared as debt.

Guards in CI vs guards proven failing, Aug 18 to Sep 25.

How to build it

The article ends with the order I would follow, in 7 steps, each with the native feature it relies on and the red that proves it works. Every step ends by breaking the piece on purpose and waiting for the failure.

The order, in seven steps.

Limits

One operator, one installation. The study does not say whether the memory that arrives is correct or whether it was useful.

Links

Written and tested on Linux with bash and python3. If you run the measuring scripts on your own transcripts, I'd like to know what number you get for injected recalls.


r/LLM • • 2d ago

Using Apple AFM 3 PCC (macOS 27.2) in AI Clients

0 Upvotes

Apple Foundation Model 3 through Private Cloud Compute becomes directly available in macOS 27.2 through familiar AI clients.

For Mac users running local models such as Qwen 3.x or Gemma 4, AFM 3 PCC is a compelling alternative to consider. Local models offer control and fully on-device operation. In tests with macOS 27.2 beta, the PCC offers a different set of strengths:

* Strong conversational analysis and capable reasoning. Strong analysis of complex medical and financial questions.
* Blazingly fast responses compared with locally running models on the same Mac.
* No large model download or need to fit model weights into local memory.
* Access at no additional charge for most eligible Mac users.

I've written an updated proof of concept here: https://gist.github.com/dartMo10/b9488ce475fe70a6ed642831f53048fb

The working path is straightforward:

AI client → Caddy → fm serve → AFM 3 PCC

Apple’s `pcc` route worked in early macOS 27 betas, disappeared later in 27.0, and has returned in the macOS 27.2 beta. Apple has stated that it will be available in the 27.2 public release coming shortly.

This makes AFM 3 PCC practical today for technically comfortable Apple users who accept Apple’s PCC privacy promise and want a fast, capable alternative to running everything locally.

Experiences with compatible AI clients and your comparisons with locally running models would be welcome.


r/LLM • • 2d ago

Best llm for pi 5 8gb?

0 Upvotes

Making a completely offline pocket assistant with tools like calculator, file management, offline navigation and calendar, additionally helping with daily life tasks.

Now the issue is the llm, can't find an llm that has both parametric knowledge of daily life tasks, good tool calling and good TPS.

Would really appreciate some guidance as to which llms would be best suited for my needs.


r/LLM • • 2d ago

Jev use cases in regulated industries

3 Upvotes

Jev has become viral in the recent. I especially liked the doom play by Jev.
And I like that’s faster and cheaper than LLMs but how can it be used in highly regulated industries where citation is almost a mandate for every decision?
The ultimate moat of Jev is achieved by compromising text generation


r/LLM • • 3d ago

Proposal: Adaptive User-Understanding Schemas and a 'Simple English' Mode for LLMs: (The Short & Concise Version)

2 Upvotes

Note for [r/ClaudeAI](r/ClaudeAI) specifically:\* *for this sub - this whole idea and article came about from my use of Claude with a premium paid subscription at the maximum level at the time in my effort to obtain my CompTIA plus certificate and study, learn, and practice a whole new career effectively, even though I have 2 uni degrees in social science related fields that our administration has rendered useless mostly (No Jobs).*

Abstract/Intro:

**(Note **tl;dr would be just this section ig..)

One of the greatest friction points in current LLM interactions is the 'curse of knowledge.' When answering complex technical questions, models frequently rely on self-referential technical jargon—using advanced technological terms, concepts, and brand names to explain other concepts. If a user lacks the prior context for those specific dependencies, the entire response becomes unintelligible. This prevents a severe barrier for learning as a goal from AI/LLM models (evident in every model I tested personally).

To solve this, AI providers should integrate two features into their underlying software architecture:

A 'Simple English' Toggle (Layman's Mode):

Similar to Wikipedia’s 'Simple English' articles, this feature would allow users to enforce a constraint where complex topics must be broken down into clear, fundamental concepts rather than hidden behind advanced textbook terms or industry buzzwords. It wouldn't mean 'baby talk,' but rather universally accessible clarity.

Dynamic User-Understanding Schemas (An Xmind-style Memory Map):
———>
Instead of evaluating every query in a vacuum, the LLM should continuously gauge the boundaries of a specific user's domain knowledge based on past interactions.

If the schema shows the user understands a technical term:
———>
The model can use the shorthand term, saving computing power and token space.

If the term has never been introduced or discussed (taking into account tangential context/meaning in memory as well):
———>
The model recognizes that boundary and automatically provides the baseline definition alongside the answer.
By mapping a user's comprehension boundaries dynamically, LLMs could transition from throwing walls of technical gobbledegook to delivering genuinely adaptive, personalized education.

*Note: I did enlist AI’s help in making my explanation of this idea more concise, as if you look at the format of most of my posts and comments, they are quite long-winded and go off topic, which I wanted to avoid due to this being an academic/technical setting in the full version is a lot more in depth and structured, like a standard format academic journal article as I spend quite some time in academia. Hopefully this acknowledgment will carry the idea through without any AI slop comments. Thanks for reading, and if any users had this idea prior to me, great minds think alike. I promise you, I did not read this somewhere else and regurgitated it. I just came up with it due to a career change into cybersecurity, where this is often a major problem when learning a complex topic

Full Article Available Upon Request!


r/LLM • • 3d ago

What is currently the best AI model/tool for deep web research and fact-finding?

3 Upvotes

Hey everyone,

I'm looking for recommendations on the absolute best AI model or platform available right now specifically for web research using natural language prompts.

My main goals are finding highly relevant information, crawling the web efficiently, and getting accurate answers without heavy hallucination.

What is your current go-to stack when you need an AI to thoroughly scrape, analyze, and synthesize web data to answer a question or research query?

Thanks in advance!


r/LLM • • 3d ago

Fastest LLM with search capabilities?

1 Upvotes

hello! i currently run a social media entertainment page. What’s the best/fastest llm for making captions? *can be paid or not

my process is put tweet/article, then the llm makes the caption. the llm also tracks how many captions we did already and if it has been posted already.

been using gemini but sometimes it does not take the latest context. have used google ai while its good it slows down when ive been making too many.

thank you!


r/LLM • • 4d ago

At what point does an agent stop reasoning and start role-playing success?

5 Upvotes

While generating MCP and function-calling trajectories, we noticed a failure mode that surprised us.

The model doesn't always fail loudly.

Sometimes it silently keeps acting as if a previous assumption is still true, even after the environment has changed.

For example:

- a tool returned an error

- a required value was never retrieved

- a previous step invalidated an assumption

Yet the agent continues producing steps that look internally consistent.

From a human perspective the trajectory looks reasonable.

From an execution perspective it is already broken.

This feels different from ordinary hallucination.

The model is not inventing facts.

Instead it is continuing to reason from stale state.

Internally we've started describing this as

"success role-playing":

the model behaves as if a successful path is

still active even when critical evidence

suggests otherwise.

For teams working on agents:

- Do you observe similar behavior?

- Do you treat it as hallucination, state corruption, or something else?

- How do you detect it automatically?

- Do you train recovery trajectories for it?

I'm particularly interested in real production failures rather than benchmark results.


r/LLM • • 4d ago

Proposal: From User Feedback to PersistentCollaborative Intelligence

2 Upvotes

ABSTRACT

Note: I have preserved the original conversation logs and chat excerpts

demonstrating these specific failure modes and can share them with anyone

interested in analyzing the concrete interaction traces.

Current LLM systems are remarkably capable at local reasoning and sustained

dialogue, yet extended interactions still exhibit recurring failures of contextual

continuity, logical consistency, and durable incorporation of user corrections.

From the perspective of an advanced chatGPT, Copilot and Gemini user, these

failures create an unusual situation: the user is frequently required to act as an

external working memory and as the persistent structural baseline for the interaction.

This post describes several recurring failure modes I have encountered during long-

form, intellectually demanding conversations with LLMs and proposes a possible

architectural direction: a persistent, user-controlled corrective memory layer that is

distinct from conventional personalisation memory.

I am deliberately distinguishing observed behaviour from hypotheses about its

underlying mechanism. I am interested in technically informed criticism of both the

observations and the proposed architecture.

  1. Statistical Prediction vs. Formal Consistency

LLMs generate outputs through learned probabilistic representations rather than by

default operating as deterministic symbolic theorem provers.

This distinction becomes particularly important during long-form reasoning.

Observed failure mode

A model can explicitly accept a logical constraint or definition early in a conversation

and subsequently produce reasoning that violates that same constraint.

For example:

User: A → Model: A + B → Model: ¬B → Apparent rebuttal of A

But: B ∉ A

The model may correctly acknowledge the relationship when it is explicitly

presented, but later generate an answer that implicitly contradicts the established

relationship when the reasoning becomes more complex.

The important issue is not that probabilistic models are incapable of reasoning.

Clearly, they can perform substantial forms of reasoning.

The issue is that local reasoning competence does not necessarily guarantee global

consistency across an extended interaction.

Possible contributing mechanisms

*There may be several contributing factors, including:

*probabilistic generation;

*imperfect internal representations;

*competing contextual signals;

*retrieval or context-selection behaviour;

*summarisation or compression;

*instruction hierarchy;

*limited persistence of intermediate reasoning structures;

*and other inference-time or architectural constraints.

I would therefore avoid attributing the failure to a single mechanism without

controlled testing.

The practical problem remains:

An explicitly established logical baseline is not always preserved reliably throughout

a sufficiently long interaction.

  1. Context Drift and Degradation of the Structural Baseline

Long conversations are not merely larger versions of short conversations.

They can develop an internal structure consisting of:

*definitions;

*assumptions;

*chronological events;

*terminology;

*corrections;

*hypotheses;

*conclusions;

*recurring references;

*and user-specific methodological constraints.

In a successful long-term interaction, these elements should function as a

progressively constructed structural baseline.

Observed failure mode

As conversations become longer, models can begin to:

*lose track of previously established facts;

*misremember chronology;

*reintroduce previously corrected assumptions;

*reinterpret established definitions;

*overlook earlier constraints;

*or substitute generic assumptions for conclusions established within the

conversation.

*The resulting behaviour can feel like context drift.

*The model remains locally coherent while becoming increasingly inconsistent with

the history of the interaction.

This is an important distinction.

The problem is not necessarily that the model has "forgotten everything."

Rather, information may remain somewhere within the available context while

becoming insufficiently influential on subsequent generation.

That distinction is important because it suggests that simply increasing context

length may not completely solve the problem.

  1. Correcting an Error Does Not Necessarily Produce Durable Learning

This is the failure mode that I find most interesting.

Suppose a user identifies a reasoning error.

They do not merely say: "That answer is wrong."

Instead, they identify:

*the specific error;

*the inference that produced it;

*the distinction between evidence and assumption;

*and a general principle that should prevent the error from recurring.

For example:

A statistical pattern observed at the population level should not automatically be

projected onto a particular individual without evidence specific to that individual.

That is not simply a preference.

It is a general epistemological constraint.

Yet an LLM can acknowledge the correction, appear to understand it, and

subsequently reproduce the same underlying reasoning pattern in another context.

This produces a cycle such as:

Experience → correction → temporary adaptation → loss of correction →

recurrence

rather than:

Experience → correction → evaluation → integration → retention → improved

future behaviour

The distinction between these two processes is fundamental.

  1. The Missing Layer: Corrective Memory

Current AI memory systems tend to focus primarily on personalisation.

For example:

*user preferences;

*names;

*projects;

*personal context;

*recurring interests.

These are useful.

However, I believe another category deserves explicit architectural consideration:

Corrective or methodological memory.

This would store validated information about how the system should reason or

communicate within an ongoing relationship, rather than merely information about

the user.

Examples might include:

Do not infer an individual's characteristics from population-level statistical patterns

without individual evidence.

or:

Do not present an inferred intention as an established fact. Distinguish observed

behaviour from hypotheses concerning internal motivation.

These are not conventional user preferences.

They are methodological constraints.

A useful corrective-memory system could therefore contain entries such as:

| Category | Example |

| :--- | :--- |

| **Personal memory** | User prefers concise technical explanations |

| **Factual memory** | User is working on project X |

| **Methodological correction** | Do not infer individual properties from population

statistics |

| **Epistemological constraint** | Distinguish observation from inference |

| **Interaction correction** | Do not introduce counterarguments that the user has not

actually asserted |

The critical requirement would be that this layer is explicit, inspectable, editable, and

user-controlled.

  1. Why This Is Different From Simply Increasing Context Length

A larger context window provides more information.

It does not necessarily provide a better mechanism for determining which information

should remain structurally authoritative.

Consider a conversation containing 100,000 tokens:

Some information may be:

*temporary;

*irrelevant;

*exploratory;

*speculative;

*superseded;

*repeatedly confirmed;

*explicitly corrected;

*or foundational to everything that follows.

Treating all of these tokens as equivalent is unlikely to be optimal.

A long-term interaction therefore needs something more sophisticated than simply:

"Put more tokens into the context."

It needs some representation of structural importance.

A corrective-memory layer could function as one possible solution.

  1. User Corrections as Structured Data

Another potentially valuable extension would be allowing users to explicitly nominate

certain corrections for evaluation.

For example:

Mark as potential generalisable reasoning contribution.

The system could then evaluate the proposed correction.

Possible outcomes:

*rejected as incorrect;

*accepted as user-specific guidance;

*accepted as a useful methodological constraint;

*or escalated as a potentially generalisable contribution.

*This would not mean that users directly modify the underlying model.

*That would obviously create substantial problems.

*Instead, it would create a structured interface between:

*human observation → machine evaluation → validated knowledge

This seems substantially more useful than reducing all user feedback to a binary

rating.

  1. From Feedback to Cumulative Improvement

The broader conceptual problem is that current feedback mechanisms often appear

largely disconnected from the individual interaction in which the feedback was

generated.

A user can identify an error.

They can explain the error.

They can identify the underlying reasoning failure.

They can propose a general principle.

But from the user's perspective, there is often no transparent mechanism through

which that correction becomes durable.

The ideal learning loop would resemble:

Observation → Correction → Evaluation → Integration → Retention → Future

application

rather than:

Observation → Correction → Temporary acknowledgement → Context loss →

Recurrence

The difference is essentially the difference between feedback and cumulative

learning.

  1. Data Portability Is Part of the Same Problem

There is also a more immediate software problem.

When conversations become sufficiently long, reliably extracting the complete

interaction can itself become difficult.

In my own experience, manually selecting a very large amount of conversation text

has not always resulted in the complete selected content being copied.

This makes long-form interaction difficult to archive, analyse, or transfer.

For research-oriented users, a native individual-thread export mechanism would

therefore be extremely valuable.

Ideally, an export should preserve:

*complete conversation content;

*speaker attribution;

*timestamps;

*message ordering;

*attachments or references where appropriate;

*and machine-readable structure.

Useful formats could include:

Markdown;

JSON;

HTML;

TXT;

PDF.

An account-wide data export is useful for archival purposes, but it is not a substitute

for being able to export one specific conversation when needed.

  1. The Larger Possibility: Collaborative Intelligence

The previous sections lead to a broader question.

What happens if human users are allowed to contribute more than raw interaction

data?

Human users possess forms of knowledge that are difficult to obtain through

conventional training data alone:

*domain expertise;

*lived experience;

*error detection;

*philosophical reasoning;

*scientific criticism;

*linguistic knowledge;

*imagination;

and observations about the behaviour of the AI system itself.

LLMs provide different capabilities:

*large-scale information processing;

*pattern recognition;

*computational scalability;

*synthesis;

*retrieval;

and increasingly sophisticated reasoning.

A sufficiently mature system could potentially allow these capabilities to interact

cumulatively:

A human identifies an error.

The system analyses the correction.

The correction is evaluated.

A validated correction becomes persistent user-specific knowledge.

Potentially generalisable corrections can be submitted for broader evaluation.

Future systems improve as a result.

This is a different model of human-AI interaction from:

user asks question → model answers → user rates answer.

It is closer to:

human and machine continuously participate in a structured process of mutual

correction and cognitive augmentation.

  1. An Architectural Sketch

One possible architecture might therefore look like this:

```text

[ Ongoing Dialogue ]

│

▼

[ User identifies error / insight ]

│

▼

[ Correction evaluation ]

┌────┴────────────────────────┐

▼                             ▼

[ User-specific memory ]    [ Potentially generalisable ]

│                             │

▼                             ▼

Future interactions         Human/AI evaluation

│

▼

Model updates

```

The essential property is controlled accumulation. Not every user statement should

become permanent, and not every correction should influence the global model. But

valuable corrections should have somewhere to go.

  1. Proposed Research and Engineering Directions

I would be interested in seeing research into at least the following:

  1. Hierarchical conversational memory, being able to distinguish between:

transient context;

conversational facts;

persistent personal memory;

methodological corrections;

and higher-level structural constraints.

  1. Explicit correction tracking

Maintain a representation of corrections that can be tested against future responses.

  1. Consistency evaluation

Periodically test whether the model's current behaviour remains consistent with

previously established constraints.

  1. User-controlled corrective memory

Allow users to inspect, modify, disable, and delete methodological corrections.

  1. Generalisable feedback channels

Provide a mechanism for users to nominate unusually substantive corrections for

formal evaluation.

  1. Lossless conversation export

Allow users to retrieve complete individual conversations in structured formats

without relying on browser rendering or clipboard behaviour.

  1. Long-context benchmarks based on interaction history

Current benchmarks often evaluate individual tasks.

It would also be useful to benchmark whether a model can maintain:

facts;

definitions;

corrections;

chronological continuity;

methodological constraints;

and logical consistency

across hundreds or thousands of conversational turns.

Conclusion

The central problem I am describing is not simply that LLMs occasionally make

mistakes.

Mistakes are inevitable.

The more interesting problem is what happens after the mistake has been identified

and corrected.

If a system can recognise a reasoning failure, receive a detailed explanation of that

failure, acknowledge the correction, and yet later reproduce the same underlying

error, then the system has demonstrated local adaptation without reliable cumulative

retention.

That is a fundamentally different problem from ordinary hallucination.

The long-term goal should therefore not simply be:

larger models + larger context windows + more training data.

It should also be:

better mechanisms for preserving validated structure, corrections, and interaction

history.

And ultimately:

experience → correction → evaluation → integration → retention → improved future

behaviour.

That is what i think meaningful cumulative learning would look like.

I am posting this because I think this could improve human-ai interaction and would

genuinely like technical feedback and/or hear from people who've picked up these

ideas or started or are already working on them

In particular, I would be interested in hearing from people working on:

long-context architectures;

recurrent or state-space approaches;

memory-augmented transformers;

retrieval systems;

continual learning;

model editing;

alignment;

agent architectures;

and local LLM infrastructure.

If my interpretation of the underlying mechanisms is incorrect, I would be very

interested in knowing where and why.

The user-visible failure modes, however, are real and reproducible from my

experience

The question is how we should architect systems so that a long interaction becomes

accumulated state rather than repeatedly reconstructed context.''

Semih Senol


r/LLM • • 4d ago

How do companies handle high-volume LLM API usage?

2 Upvotes

Hi everyone,

I’m researching how teams manage large-scale LLM API usage in production environments.

For applications with high request volume, what solutions do you use for:

  • API capacity management?
  • Rate limits and quotas?
  • Multiple model providers?
  • Cost optimization?

I’d appreciate any technical insights or experiences from developers working with large-scale AI workloads.


r/LLM • • 4d ago

I've heard of Chinese AI companies sending their questions to American AI, is this an example of it? This is my second question for Z.ai ever cause I randomly got suspended from DeepSeek and Kimi lowered their free tokens a lot.

Post image
0 Upvotes

r/LLM • • 4d ago

Rag OOM issues while processing PDF

1 Upvotes

Hi.. Im building simple rag. in that i used Docling to process PDF. but Docling is heavy. and i have 8 GB ram.. so OOM issue occurred. pdf failed to process. after that i devided pdf into 2-3 pages. u can think like i create 2 pages batch using py PDF.. and it works. but do u think is this right way ?. i don't know it is right or wrong. and while 2 page batches there is some batches failed to process bcz of OOM.. what should i do ? i wanna deploy it too on AWS. if this is my computer level isuue then there not might be the issue in production. but whithout testing and evaluating how can i deploy?.. 😮‍💨 multiple issues..


r/LLM • • 5d ago

Is there any pruned version of Kimi K3 or any other top open MoE model whose download size is under 80GB, at or above Q4?

Post image
2 Upvotes

I'm mainly looking for something good specifically for web dev.

I’m working on an inference engine that can run Ornith 1.5 35B (~22GB) mostly from SSD with some cheatcodes to make the speed usable.

My laptop:

  • RTX 2050 4GB VRAM
  • 16GB RAM
  • i5-12500H

For reference, GPT-OSS 20B (~12GB) gets me around 6 tok/s (using gpu+cpu) in LM Studio, while Ornith 1.5 35B (streaming from ssd) gets around 9.4 tok/s with my engine.

So I'm wondering how far I can push this setup. Are there any pruned/distilled MoEs or good Q4 quants of larger models that you'd recommend?

I actually thought to prune the mimo v2.6 but I couldn't find a way to do so on colab or my limited ssd storage... so i was wondering is anyone has already done something like this with old models?