EverProduct
AI

Stage 04 · Everyday Work

Documents and Research

A hundred-page PDF and three questions you need answered. This is the model working in its most reliable mode — provided you ask the document rather than the model.

A contract, a report, a folder of interview transcripts, a specification nobody has read end to end. You need three things out of it and you have twenty minutes.

This is the best-case configuration for everything in this sphere. The material is on the desk, so the job is transformation rather than recall. The answer is anchored in text you possess, so verification is comparison rather than faith. Every strength from article 9 lines up, and every weakness is avoidable — if you keep one distinction.

Ask the document, not the model.

The difference is entirely in whether the answer is required to come from the text. "What does this contract say about late payment?" invites recall to fill any gaps. "Quote every clause in this contract that bears on late payment, with the section number; if there are none, say so" makes fabrication visible.

The working method

Demand traceability, always. Every claim gets the exact sentence it came from. This one habit converts the whole activity from trusting to checking — you're comparing quotes against a document you have, which takes seconds per atom.

Summarise for a purpose, never in general. "Summarise this" produces the median of the document: technically accurate, uniformly useless. "Extract everything relevant to payment terms and penalties, as a table with section references" produces something you can act on. The question does the work; without one you get compression without judgement.

Ask what isn't there. The most under-used question in the set: "What would you expect a document of this type to address that this one doesn't?" Gaps, missing clauses, unstated assumptions, the section that quietly avoids a topic. This is where an outside reader earns its keep.

Compare rather than read. Two versions of a contract, three vendor proposals, last quarter's report against this one: "list every substantive difference, ignoring formatting". Diffing is mechanical work at which it is excellent and you are slow.

Hunt for contradictions. "Where does this document contradict itself, or contradict the second document?"

Prepare your questions, not your conclusions. "Based on this, what are the ten questions I should ask the author before signing?" You get leverage without delegating the judgement.

The failure that doesn't announce itself

Hallucination is not the main risk here. Omission is.

A model that misses a clause doesn't say so — it produces a confident, complete-looking answer covering everything it did find. Absence of mention is not evidence of absence in the document, and the difference is invisible from the answer alone. Two structural causes: the middle of a long context is the weakest position (Lost in the Middle, article 2), and a summary is by definition a decision about what to drop.

So separate the two kinds of question and treat them differently. For interpretation — what does this mean, what follows from it — the model is strong. For completeness — are these all of them, is it mentioned anywhere — trust nothing without a mechanical check:

  • ask for an exhaustive list with locations, then verify with a plain text search for the key term;
  • process long documents in explicit sections rather than in one pass;
  • ask the same completeness question in a fresh chat and compare the lists.

Anything you'd stake a decision on gets read by you at the location the model pointed to. It found the place; you read the words.

Search, and its own failure mode

Turning on web search changes the mode: the model now answers from pages it retrieved rather than from the landscape. That fixes the cutoff problem and reduces invention substantially. It introduces three new ones.

A real link can sit under a claim it doesn't support. Grounding cuts fabrication, not misreading. Open the links — the ones that matter, at least.

Sources agree because they copy each other. Three articles repeating one press release look like corroboration and are one source. Ask where a claim originates, not just who repeats it.

Retrieval is a filter with its own bias. What surfaced is what was findable and popular, which is not the same as what's true or best. On anything contested, ask explicitly for the strongest opposing position and who holds it.

Research is not summarising

A tempting workflow: forty papers in, forty summaries out, a synthesis at the end. It feels like a literature review. It isn't one, and the reason is the illusion of competence from How to Learn: you'll hold the shape of the field with none of the detail, and — worse — the confidence of someone who has read it.

A version that works keeps you in the loop where it matters:

  1. Map the territory with the model: what are the main positions, the vocabulary, the standard objections, who's cited constantly. It's genuinely strong here, and this is the part that used to take a week.
  2. Find candidates by search — then check that they exist and say what was claimed.
  3. Read the important ones yourself. Non-negotiable if the topic is yours. Use the model while reading: explain this paragraph, what does this term mean, how does this relate to that other paper.
  4. Synthesise after reading, using it as an opponent: what am I missing, what would someone disagree with, where am I over-reading a single study.

The rule underneath: use it to decide what deserves your reading, not to replace the reading. Screening is where the leverage is; the reading is where the understanding is.

Numbers in documents

One reminder from article 8: figures pulled out of a table by eye are generated, not computed. For anything you'll rely on, ask for a script or a formula that produces the number from the data, and run it. Extraction into a structured table is a legitimate use; arithmetic in prose is not.

In practice

Always demand exact quotes with locations.

Summarise toward a question, never in general.

Ask what a document of this type should contain and doesn't.

Treat completeness questions as untrusted — exhaustive list, then a text search to confirm.

Open the links that matter, and ask where a claim originates.

Read the key sources yourself; use the model to decide which ones those are.

Check yourself

Close the article and answer in your own words:

  1. What's the difference between asking the document and asking the model, and how do you enforce it?
  2. Why is "summarise this" a weak instruction, and what replaces it?
  3. Why is omission a bigger risk here than invention, and what are its two structural causes?
  4. How do you handle a completeness question differently from an interpretation question?
  5. What three problems does search introduce even though it reduces fabrication?
  6. Which steps of a research workflow can be delegated and which can't?

In short

  • Documents are the model's best configuration: material on the desk, transformation rather than recall, verification by comparison.
  • Demand exact quotes with locations — it turns trusting into checking.
  • Summarise toward a specific question; a general summary is the median of the document.
  • The signature failure is omission, not invention: a missed clause produces a confident, complete-looking answer. Causes: weak middle of long contexts, and summarising being a decision about what to drop.
  • Interpretation questions are safe; completeness questions need an exhaustive list plus a mechanical text search.
  • Search fixes the cutoff but adds: real links under unsupported claims, false corroboration from sources copying each other, and retrieval bias toward the findable.
  • Use it to decide what deserves reading, not to replace reading. Summaries of forty papers are not a literature review.
  • Ask the document, not the model.