How Language Models Work
You do not need the math. You need a mental model accurate enough to predict when Claude will be brilliant and when it will confidently make something up. Without that model, people make one of two mistakes: they treat Claude like a search engine with perfect recall, or they catch one wrong answer and dismiss it entirely. Both waste the opportunity.
What Claude actually does
Claude reads everything currently in front of it — your instructions, the conversation so far, any files you provided — and generates the most plausible continuation, word by word. It is not looking up answers in a database. It is producing what a knowledgeable response would look like, drawing on patterns learned from enormous amounts of text.
Give it the sentence "The invoice is 45 days past due, so the next step is..." and it continues the way an experienced accounts-receivable person plausibly would. That is the whole trick — and it is remarkably powerful, because most knowledge work is producing the right continuation: the summary of this document, the reply to this email, the analysis of this spreadsheet.
Everything Claude can currently see is called its context. Think of it as a whiteboard: what is on the whiteboard shapes the answer; what is not on the whiteboard effectively does not exist, no matter how important it is to you.
The open-book rule
Here is the single most useful consequence. Ask Claude a factual question with nothing on the whiteboard, and it answers from training patterns — a closed-book exam. It will usually be right about broad, stable knowledge and shakiest on specifics: exact figures, dates, names, citations, anything recent, anything internal to your company. When the true answer is not available, the shape of a plausible answer still is — which is why a fabricated statistic or citation looks so convincing. This is what people mean by hallucination. It is not lying; it is pattern completion without a source.
Paste in the actual contract, report, or policy, and the exam becomes open book. Claude works from your document, and you can check every claim against the page it came from. Most of the reliability you will ever get from Claude comes from this one move: supply the source.


Placeholder: replace with a diagram showing instructions and sources entering the model, followed by a draft that passes through human verification.
Four behaviors this explains
- Context matters. Missing background produces generic output or wrong assumptions — Claude fills gaps with the most typical case, which may not be yours.
- Wording matters. The task description steers the continuation. "Summarize this" and "List the three commitments we made and their deadlines" produce very different documents.
- Outputs vary. The same request can yield different phrasing or structure on different runs. Useful for options; a reason to keep formats explicit when consistency matters.
- Confidence is not proof. Fluency is how Claude says everything, including its mistakes. Never use tone as evidence.
Where Claude is strong — and where to verify
Claude excels at transforming material you supply: summarizing, comparing, extracting, classifying, drafting, explaining, and proposing options. It is also genuinely useful for structuring an unfamiliar problem and surfacing the questions you should be asking.
Treat these as unverified until checked:
- Dates, names, figures, quotations, and citations
- Claims not grounded in the sources you provided
- Legal, financial, employment, security, or safety conclusions
- Anything that would change records, permissions, commitments, or customer outcomes
Match the checking effort to the consequence. Brainstorming team-offsite ideas needs a glance. A claim going into a customer proposal needs source-by-source verification. A decision affecting a person's pay, contract, or data needs a qualified human owner — with Claude as preparation, never as the decider.
Practice
Give Claude a short document from your own work and ask for a five-point summary with a source reference beside each point. Check every reference against the document. Record:
- What was accurate?
- What was incomplete?
- What sounded plausible but was not actually supported?
- What instruction would improve the next attempt?
Ten minutes of this builds more calibrated trust than a month of casual use.
Ready to move on
You can explain the open-book rule to a colleague, say why fluent output is not verified output, and choose a checking level that matches the consequence of the task.