Common prompt engineering mistakes and how to fix them
Published 2 October 2026
Most bad output is not a model problem. It is a prompt that never said who the answer is for, what shape it should take, or how you would know it was good. Below are the mistakes that cost the most time, each with the fix, plus two pieces of widely repeated advice that Anthropic's own current documentation has demoted or dropped.
The prompt is too broad to fail
"Write something about our onboarding flow" cannot fail, because nothing was specified that could be missed. You get a plausible paragraph, you feel vaguely unsatisfied, and you have no way to say what was wrong.
The fix is to state the thing you would have complained about. Anthropic's prompting best practices page frames this as treating the model like a brilliant new employee who lacks context on your norms and workflows, and says that if you want behavior that goes above and beyond, you should explicitly request it rather than hoping it gets inferred. Its golden rule is worth stealing: show the prompt to a colleague who has minimal context and ask them to follow it. If they would be confused, the model will be too.
No stated audience and no stated tone
A prompt with no reader in it produces writing aimed at nobody, and one sentence fixes it. Anthropic's page says setting a role in the system prompt focuses behavior and tone, and that even a single sentence makes a difference. Its worked example is not elaborate: "You are a helpful coding assistant specializing in Python."
The more useful half is what comes with the audience. Anthropic's page advises giving the motivation behind an instruction rather than the bare instruction: telling the model the response will be read aloud by a text to speech engine works better than a flat prohibition like never use ellipses. Constraints with a reason attached survive cases you did not anticipate. Bare prohibitions do not.
Several unrelated jobs stuffed into one prompt
Summarize this, then rewrite it for executives, then draft the follow up email, then list open risks. Each task degrades the others, and when the result is wrong you cannot tell which instruction the model dropped.
Two fixes, and they stack. First, separate inputs from instructions structurally. Anthropic's page recommends XML tags precisely for prompts that mix instructions, context, examples and variable inputs, with consistent descriptive tag names. Second, split genuinely separate jobs into separate calls. There is a caveat: the page says current models handle most multistep reasoning internally through adaptive thinking and subagent orchestration, so explicit chaining earns its keep mainly when you need to inspect intermediate outputs or enforce a specific pipeline structure. It names self correction as the most common chaining pattern: draft, review against criteria, refine, each step a separate call.
This is not Anthropic-specific advice. Google's prompting strategies guide for the Gemini API gives the same rule under the heading Break down prompts into components: "Instead of having many instructions in one prompt, create one prompt per instruction."
No output format, so you get a different shape every time
If the prompt never says what the answer should look like, you re-read a differently organized answer on every run and nothing downstream can consume it.
Anthropic's page lists four ways to steer format:
- Say what to do instead of what not to do
- Use XML format indicators
- Match your prompt style to the output you want
- Be detailed about specific formatting preferences
It also notes that removing markdown from your prompt can reduce the amount of markdown in the output.
Examples do this job better than description does. Anthropic's page calls examples one of the most reliable ways to steer output format, tone, and structure, and recommends including three to five, relevant to your actual use case and diverse enough to cover edge cases so the model does not pick up a pattern you did not intend.
Over-specifying, the opposite failure
Beginners get told to be more specific, so they keep going, and end up hand-writing a step-by-step procedure the model then follows literally past the point where it made sense. The popular technique lists rarely warn about this direction.
Anthropic's context engineering post advises writing system prompts at what it calls the right altitude, a Goldilocks zone between two common failure modes, one on each side of the target. Too vague and you get the broad-prompt problem above. Too prescriptive and you have replaced the model's judgment with your own guess at a procedure.
The best practices page arrives from another direction. The first of its thinking and reasoning guidelines is to prefer general instructions over prescriptive steps: an instruction to think thoroughly often produces better reasoning than a hand-written plan. Its blunt version: "Claude's reasoning frequently exceeds what a human would prescribe."
The context engineering post also explains why over-stuffing hurts. It describes an attention budget models draw on when parsing large volumes of context, treats context as a finite resource with diminishing marginal returns, and sets the goal as the smallest possible set of high-signal tokens that maximize the likelihood of the outcome you want.
Copying advice the documentation has demoted or removed
Two techniques on the popular list have changed status on Anthropic's current page: chain of thought has been demoted, and prefilling the assistant turn is documented as removed. Chain of thought is not a headline technique there, and prefilled responses on the last assistant turn are no longer supported starting with Claude 4.6 models and Claude Mythos Preview.
Prompt technique lists age badly, and two items on the popular list no longer mean what you think if you are citing Anthropic's current page as your authority.
Chain of thought is not a headline technique there. The page has no chain-of-thought section, and manual chain-of-thought appears once, as a fallback for when thinking is turned off. Its thinking guidance runs the other way, toward general instructions rather than prescribed steps.
Prefilling the assistant turn is documented as removed, not as a trick. Anthropic's page states that starting with Claude 4.6 models and Claude Mythos Preview, prefilled responses on the last assistant turn are no longer supported: "Requests with prefilled assistant messages to these models return a 400 error." It gives migration paths instead, including:
- Structured Outputs or a tool with an enum field where you need a machine-readable shape
- Direct system-prompt instructions for removing preambles
- Moving a continuation into the user message
- Hydrating context through tools or compaction
The general lesson matters more than either specific: that page is versioned against named models, and it tells you to re-check a model-specific technique against your own evaluations before applying it elsewhere.
Never testing the prompt more than once
One good output is not evidence. The failure you care about is usually waiting on an input you have not tried yet.
Run the prompt on several inputs, including the awkward ones, before you trust it. Anthropic's page suggests two cheap checks along the way: ask the model to evaluate your examples for relevance and diversity or to generate more, and ask it to check its own answer before finishing. Neither replaces reading real outputs yourself.
Google's prompting strategies guide treats that iteration as the normal case rather than a sign something went wrong: "Prompt design can sometimes require a few iterations before you consistently get the response you're looking for."
Where this fits if you use Bespoke Prompting
The Prompt build type emits a fixed set of sections in a fixed order: ROLE, CONTEXT, CONSTRAINTS, SUCCESS CRITERIA, APPROACH, TASK. That ordering does two of the fixes above by construction. Role and context land before the task rather than being remembered afterwards, and SUCCESS CRITERIA forces you to write down what a good answer looks like before you write the request, which is the thing a too-broad prompt leaves out.
Sources
- Anthropic, Claude Developer Platform documentation, "Prompting best practices": https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
- Anthropic Applied AI team, "Effective context engineering for AI agents": https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Google, Gemini API documentation, "Prompting strategies": https://ai.google.dev/gemini-api/docs/prompting-strategies