The Dialogue Is the Method
The first answer is rarely the one you want — and it was never supposed to be. It's a measurement that tells you exactly which part of the brief was missing.
You send a carefully written request. What comes back is close, but wrong in a way that's hard to name.
At this point the conversation forks. One person deletes everything, rewrites the prompt from scratch and fires again — and again, hunting for the magic phrasing. Another reads the answer for what it reveals, changes one thing, and lands it on the second or third try. Same model. The difference is entirely in what they do with a first answer.
The first answer isn't the result. It's a measurement.
Read the answer as a report
Before judging quality, read the reply as evidence of what the model understood. It's the only window you have into how your brief was actually parsed, and it usually names its own problem.
Three failures, three different repairs:
It's generic — true but weightless. The material slot was empty. It's answering from the average because you gave it nothing specific to stand on. Fix: put your documents, numbers and examples on the desk. Rephrasing will not help.
It's specific but not what you want. The bar was missing. It made a reasonable choice you disagree with. Fix: state the criterion you were carrying silently — "each section needs a concrete example", "no recommendation without a cost".
It answered a different question. The task itself was ambiguous. This is the one case where rewriting from scratch is right, because the request genuinely didn't say what you meant.
Nine times out of ten it's the first two, and both are repaired by adding, not by rewording.
Fix by address
"Make it better" is the least useful instruction available. Better along which axis? You're asking a ball to roll somewhere unspecified — you'll get a different answer, not a better one, and often the same failing with new words.
Address the change: "Sections 2 and 4 are fine, leave them. The intro is too long — cut it to three sentences and start with the conclusion. Everywhere you make a claim in section 3, add a number or delete the claim."
Two habits make this dramatically more effective:
Say what to keep. Without it, an edit often costs you the parts that were already right. "Keep everything except X" is a cheap sentence with a large effect.
Change one axis at a time. Fixing tone, length and structure in one message gives you a new answer with no way to tell which instruction did what. Sequential edits converge faster than one big rewrite request — and they leave you with a rule you can reuse next time.
And when words fail, fall back to the previous article: show, don't tell. Paste a rewritten paragraph and say "this register, throughout." One demonstration beats four adjectives.
Iterate on cheap artefacts
The most common waste is iterating on a finished text. Producing three pages, hating them, producing three pages again — each round is expensive to make and expensive to read.
Move the iteration upstream. Ask for the outline first, or the ten bullet points, or the approach in a paragraph. Argue with that — it's a tenth of the size and it's where the real decisions live. Only when the skeleton is right do you ask for the full text.
The same logic applies to anything long: agree on the plan, then execute the plan. Article 8 makes this a full method for hard tasks; here it's simply the cheapest place to iterate.
Can it critique itself?
Half-yes, and the boundary matters.
Asking the model to review its own output does help when there's a real external signal to check against — code that runs and fails, a test, a checklist you supplied, a document to verify claims against. This is roughly what the Self-Refine work (Madaan et al., 2023) showed: generate, critique, revise, and quality improves.
Without such a signal, the self-check is far weaker than it looks. In Large Language Models Cannot Self-Correct Reasoning Yet (Huang et al., 2023), models asked to review their own reasoning with no external feedback frequently changed correct answers to incorrect ones — self-critique produced motion, not accuracy. Which is the same lesson as article 4, from a different angle: a check that stays inside the loop isn't a check.
So use it deliberately. "Critique this against the checklist, then rewrite" is a good move — it separates the critic from the author, and the critic in a fresh turn doesn't have to defend what was just written. But treat its criticism as a list of candidates to consider, not as verification. The judgement stays with you.
When to stop
Iteration has a floor, and people blow past it constantly. Three rounds with no real progress means one of two things: the desk is now full of failed attempts and is dragging every new answer down (article 2 — start a fresh chat and carry over only the best fragment), or this particular task is faster to finish by hand.
That second one deserves saying plainly, because nobody wants to admit it after twenty minutes of investment: sometimes the last 20% is quicker to type yourself. Getting the model to nail a specific closing sentence can cost more time than writing the sentence. Take the 80% and finish it. That's not defeat — using it for the bulk and doing the final polish yourself is often the fastest possible route.
Keep what worked
A request that produced an excellent result is not a message — it's a tool you now own. Copy it somewhere. A short personal file of "the brief that works for meeting notes / for code review / for client emails" compounds faster than any prompting trick, and it's the raw material for the permanent setup in article 18.
One last connection to How to Learn: practice without feedback doesn't produce improvement — it just entrenches what you already do. Iterating with a model is an unusually fast feedback loop, but the feedback is supplied by you. Vague feedback teaches nothing to either party; specific feedback improves the answer now and sharpens your own sense of what "good" means here. The loop is only as good as what you put into it.
In practice
Read the first answer as a diagnosis. Which slot was empty — material, bar, or the task itself?
Never write "make it better". Name the axis, the location and the target.
Say what to keep before saying what to change.
Iterate on the outline, not the finished text. Cheap artefacts, cheap rounds.
Three failed rounds means restart or finish by hand. Not "try harder".
Save the prompts that worked. They're assets, not chat history.
Check yourself
Close the article and answer in your own words:
- Why is the first answer more useful as information than as a product?
- Three ways an answer can be wrong — how do you tell them apart, and what does each one require?
- Why is "make it better" a bad instruction, and what replaces it?
- Why change only one axis per round?
- When is asking the model to critique itself genuinely useful, and when is it worse than useless?
- What are the two possible meanings of three failed rounds in a row?
In short
- The first answer is a measurement. Read it for what the model understood before judging what it produced.
- Generic answer → missing material. Wrong choices → missing bar. Different question → ambiguous task, and the only case where rewriting from scratch is right.
- Edit by address: name what to keep, what to change, where, and to what. One axis per round.
- When wording fails, demonstrate — one rewritten paragraph beats four adjectives.
- Iterate on outlines and plans, not on finished text: cheap artefacts make cheap rounds.
- Self-critique helps against an external signal (Self-Refine, 2023) and misleads without one (LLMs Cannot Self-Correct Reasoning Yet, 2023). Treat its critique as candidates, not verification.
- Three rounds without progress: restart the chat or finish it yourself. The last 20% is often faster by hand.
- Prompts that worked are assets — keep them.
- The first answer isn't the result. It's a measurement.