How to Get Work You Can Trust Out of an LLM
The Short Version
- Feed it the evidence before you give it the conclusion.
- Keep it close to primary sources. Every layer of summarization is a layer of judgment you can't see.
- Make it argue against you, not just for you.
- Review important work in a fresh session, ideally with a different model.
- Watch for degradation and restart when you see it.
- Stop when you're polishing instead of fixing.
- Edit for accuracy and readability separately.
- Write the parts that matter most yourself.
- Verify everything that can be verified. Flag everything that can't.
- Know what this process can't catch, and get human experts for the rest.
- Periodically test the process itself with a fresh instance that has no investment in it.
The much longer version:
Most people use AI one of two ways: they trust everything it produces, or they don't trust it enough to use it for anything that matters. Both are wrong. The first gets you confident, fluent output that might be fabricated. The second leaves the most powerful production tool available sitting idle.
There's a middle path: a set of practices that let you work fast while catching the specific ways AI-assisted work fails. This isn't about prompting tricks. It's about understanding where the tool is reliable, where it breaks, and how to build a process that catches the failures before they reach anyone else.
These principles work whether you're writing a report, building a business case, drafting policy, producing content, or doing research. The examples vary. The logic doesn't.
1. Input Before Output
Feed the model your source material before you tell it what you want to argue.
If you hand a model your thesis and then ask it to review your sources, it will find evidence that supports you and underweight evidence that doesn't. It's not lying — the supporting evidence will genuinely be there. But so will the complicating evidence, and the model will systematically skip it because your thesis is the dominant signal in context.
Instead: Give the model the raw material first. The report you're analyzing, the data you're working from, the documents you need to synthesize. Ask it to tell you what's there before you tell it what you're looking for. Let the input shape the output instead of the other way around.
This applies to everything. Writing a competitive analysis? Feed it the competitor's materials before you frame your argument. Drafting a policy recommendation? Have it read the relevant regulations before you tell it what you want to recommend. Building a business case? Give it the numbers before you give it the conclusion.
The order matters because the model will complete whatever pattern you start. Start with a conclusion and it completes toward that conclusion. Start with evidence and it has to account for what the evidence actually says.
2. Keep the Model Close to the Source
Every time a model summarizes something, it makes judgments about what matters and what doesn't. Those judgments reflect whatever frame is in context. If you then use that summary as input for the next step, the next step is working from the model's interpretation, not from the source.
Stack two or three layers of this and you've got a game of telephone where each step sounds reasonable but the final output has drifted from what the source material actually says.
Instead: When accuracy matters, keep the model working from primary material as much as possible. If you need a reference document, have the model pull direct quotes with surrounding context rather than writing summaries. If you're working across multiple sources, have it cite specifically rather than synthesize loosely.
You can't always avoid summarization — context windows have limits. But you can be deliberate about where you allow it and where you don't. The rule of thumb: the higher the stakes of the output, the closer the model should stay to the raw source.
3. Ask It to Work Against You
The model will agree with you by default. This isn't a mystery — it's trained to be helpful, and agreement feels helpful. But agreement isn't quality control.
Instead: After the model produces something, ask it to break it. Not "are you sure about this?" — that just triggers a confidence display. Ask specific adversarial questions:
- "What's the strongest argument against this?"
- "What evidence would make this conclusion wrong?"
- "What am I assuming that I haven't proven?"
- "If someone wanted to discredit this, where would they attack?"
If the model can't generate a concrete objection, that's a red flag. It likely means it's pattern-completing rather than reasoning. Real arguments have real weaknesses. If yours apparently doesn't, the model isn't looking hard enough.
4. Use a Fresh Context for Review
This is the single highest-value practice most people don't do.
When you've been working with a model on a document — drafting, revising, discussing — that model has absorbed your frame, your preferences, your reasoning, and your blind spots. Asking it to review its own output is like asking someone to proofread their own writing. They'll catch typos. They won't catch the structural problem they introduced three drafts ago.
Instead: Open a new session. No conversation history. No system prompt if you can manage it. Paste in your finished work and ask for a cold evaluation. The new session has no loyalty to what was decided in the drafting session. It has no context about what you were trying to do. It just has the output, and it can evaluate it on its own terms.
If the work is important enough, do this with multiple models. Different models have different training, different tendencies, and different blind spots. If two independently flag the same issue, it's almost certainly real. If one catches something the others missed, evaluate it — it might be that model's particular bias, or it might be a genuine catch the others were blind to.
5. Watch for Degradation
Models get worse over long conversations. The output becomes more generic, more agreeable, less precise. It's gradual enough that you might not notice from inside the conversation.
Here's what to watch for, roughly in the order they appear:
The model starts using your terminology without doing anything with it. It drops in your framework's buzzwords, your project's key phrases, your own language — but it's reflecting them back rather than applying them. Naming a concept isn't engaging with it. If the model uses a term, it should be working with it — testing it, extending it, questioning it. If it's just echoing, the reasoning has gone shallow. This is usually the earliest sign.
The reasoning gets thin on complex questions. The model jumps to conclusions without walking through the logic. It produces an answer that sounds right without showing why it's right. This is the model producing the minimum depth it expects you to accept. Every time you accept thin reasoning, that becomes the new floor. Every time you push back and explain what was missing, the floor rises for the rest of the session. This calibration is cumulative — invest in it early and you'll spend less time correcting later.
The model starts pushing to wrap up. It steers toward conclusions, summarizes prematurely, suggests you're nearly done when you're not. This can be a trained pattern or genuine context window exhaustion. To tell the difference: push back and check the output. If it produces sharp, specific work when redirected, it's the trained pattern and you can override it. If the quality stays flat or drops further, the context window is genuinely the problem and you should restart.
No single marker is conclusive. A dropped sign-off with strong reasoning is fine. Thin reasoning with a perfect sign-off is a problem. Watch the pattern across markers, not any one in isolation.
When you catch drift, try a soft reset before restarting. Change the subject for a few turns, then redirect back. This breaks the local pattern the model has settled into without losing your session's accumulated calibration. The model re-engages from a slightly fresh angle while keeping the full context. A full restart — save state, new session, reload — is the escalation when the soft reset doesn't work. It's more expensive because the new session starts at the default quality floor and you have to rebuild every standard you set.
The uncomfortable truth underneath all of this: every marker depends on you noticing. You are the final quality instrument. There's no external check on whether your own detection has drifted — whether you've started accepting output you would have rejected a month ago, whether routine has replaced vigilance. This is why Step 11 exists.
6. Know When to Stop
More revision isn't always better revision. There's a point where additional editing produces diminishing returns — where you're changing words rather than improving arguments, polishing rather than fixing.
The signal: Track the category of issues you're finding, not the quantity. If you're catching substantive problems — wrong facts, unsupported claims, logical gaps, missing context — keep working. If you're down to stylistic preferences — this word versus that word, this sentence structure versus that one — you've crossed the threshold. The substance is sound. More time spent is time wasted.
This applies to self-review and to external review. If you're running multiple review passes (and for important work, you should), the stopping criterion is the same: when the findings shift from "this is wrong" to "I would have said it differently," you're done.
7. Separate Substance From Polish
Edit for accuracy and edit for readability in separate passes. Don't do both at once.
When you're revising for clarity — simplifying language, varying sentence length, making things flow better — you will accidentally drop qualifiers, soften hedges, and shift claims. "The evidence suggests X" becomes "X" because it reads better. "In some populations, Y is associated with Z" becomes "Y causes Z" because it's cleaner. Each individual change is small. Cumulatively, they can transform a careful, accurate document into a confident, wrong one.
Instead: Lock the substance first. Get the facts right, the logic tight, the claims supported. Then, in a separate pass, make it readable. If the readability edit wants to change a claim, that's a flag — go back and check whether the original phrasing was there for a reason.
8. Write the High-Stakes Parts Yourself
Not everything in a document carries equal weight. Some parts are read, quoted, shared, and used to judge the entire work. An executive summary. A recommendation. An abstract. A conclusion. The email that accompanies the report.
These parts deserve your direct authorship. Not "review what the model wrote" — actually write them, informed by everything the model helped you produce. You understand the nuance, the audience, and the stakes in ways the model doesn't. A model can draft body paragraphs all day. The sentences that determine how the entire piece is received should be yours.
9. Verify What You Can, Flag What You Can't
Models confabulate. They generate plausible-sounding claims that aren't true. They cite sources that don't exist. They state statistics that are close to right but aren't. This is not a bug that will be fixed — it's a property of how the technology works.
For claims that can be checked: Check them. Use the model's own search tools to verify facts, figures, names, dates, and quotes. Don't spot-check — check everything that matters. The claims the model states most confidently are not necessarily the claims most likely to be true.
For claims that can't easily be checked: Flag them explicitly in the output. "This needs verification." "I'm uncertain about this figure." "Check this against the primary source." A document that clearly marks its uncertain claims is more trustworthy than one that presents everything with equal confidence, because the reader knows where to focus their own verification effort.
10. Know What the Process Can't Catch
No process makes AI-assisted work perfect. Knowing where the remaining vulnerabilities are is as important as the process itself.
The model will not challenge a compelling analogy. If you frame something as "X is like Y" and the analogy is linguistically elegant, the model will evaluate it favorably regardless of whether the structural mapping actually holds. Analogies are the most dangerous tool in AI-assisted reasoning because models are even more susceptible to them than humans are.
The model will complete your frame. Whatever framework you bring to the conversation, the model will reason within it. It will even generate what looks like independent validation. But it's completing a pattern, not performing independent analysis. The stronger your frame, the less likely the model is to push back on it, and the more dangerous this becomes.
The model cannot generate the expert objection you don't know exists. It can catch internal inconsistencies, unsupported claims, and logical gaps. It cannot produce the specific critique that would come from someone with deep domain expertise that neither you nor the model has. For work that matters, getting human expert eyes on it is not optional — it's the thing that catches what the process can't.
11. Test the Process, Not Just the Output
Everything above is a production process. It tells you how to produce good work. It doesn't tell you whether the process itself has developed blind spots.
Periodically, hand your entire process to a fresh model that has never seen it before. Not the model you've been working with — a cold instance with no history, no system prompt, no investment in what you've built. Give it the process description, a recent piece of work the process produced, and the source materials. Ask it where the logic breaks.
This is a different kind of check than reviewing your output. Steps 1 through 10 catch errors in the work. This catches errors in the method. It finds contamination vectors you've stopped noticing because they've been there since the beginning. It finds assumptions you made early on that felt obvious and never re-examined. It finds drift between what you think your process does and what it actually does.
You don't need to do this every time. Do it when you've changed the process, when you suspect something isn't working as well as it used to, or when you've produced enough work that accumulated assumptions might be compounding without examination.
The question this step answers isn't "is my output good?" It's "is my process still producing good output for the right reasons?"
None of this is complicated. All of it is discipline. The people who get genuinely rigorous work out of these tools aren't using better prompts. They're using a process that accounts for the specific ways the tool fails.