Guide

How to write a system prompt that actually works

Most system prompts fail in one of two directions. Either they are too vague and the model fills the gaps with its own assumptions, or they are a rigid procedure that holds up until reality stops matching step three. The shape that survives contact with real inputs has four parts: an identity, an output contract, constraints with reasons attached, and a small set of examples.

The altitude problem

Anthropic's Applied AI team uses the phrase right altitude for this in Effective context engineering for AI agents, and calls it the Goldilocks zone between two common failure modes. The post describes both. At one extreme it puts engineers "hardcoding complex, brittle logic in their prompts to elicit exact agentic behavior", which it says creates fragility and increases maintenance complexity over time. At the other it puts guidance that is vague and high level, and "fails to give the LLM concrete signals for desired outputs or falsely assumes shared context". Both produce the same symptom: the model guesses.

Structure is the other half, and it comes from a different Anthropic page. Prompting best practices says XML tags help Claude parse a prompt unambiguously when that prompt mixes instructions, context, examples and variable inputs, and it recommends consistent descriptive tag names with nesting for natural hierarchies. The context engineering post's related advice is about volume rather than length: good context engineering, it says, means finding "the smallest possible set of high-signal tokens" that maximize the likelihood of the outcome you want.

Identity: narrower than "helpful assistant"

A role that could describe any assistant steers nothing. Name the domain, the seniority, and the thing this prompt does not do.

Anthropic's Prompting best practices page makes a modest claim about roles: setting one in the system prompt focuses behavior and tone, and even a single sentence makes a difference. Its worked example is a one line system string describing a helpful coding assistant specializing in Python.

The word doing the work is specializing. The page's golden rule is the test: show the prompt to a colleague with minimal context, and if they would be confused, the model will be too.

The instinct to put this in the system prompt rather than the request is not Anthropic-specific. Google's prompting strategies guide for the Gemini API gives the same placement rule under Core prompting principles: "Place essential behavioral constraints, role definitions (persona), and output format requirements in the System Instruction or at the very beginning of the user prompt."

Output contract: "return a summary" is not an instruction

This is the part that decides whether anything downstream can consume the result. A contract worth the name states the fields, the type of each, the allowed values where the set is closed, what to emit when a value is unknown, and whether anything may appear outside the structure. If you do not say no preamble, you will get a preamble.

Anthropic's page lists four ways to steer output format:

  • tell the model what to do rather than what not to do
  • use XML format indicators
  • match your prompt style to the output you want
  • use detailed prompts for specific formatting preferences

It adds a detail that catches people out: removing markdown from your own prompt can reduce the volume of markdown in the output.

Constraints: give the reason, and watch the volume

Providing the motivation behind an instruction works better than a bare rule, and Anthropic's example is telling the model that the response will be read aloud by a text to speech engine rather than only forbidding ellipses. A constraint with a reason generalizes to cases you did not list. A bare prohibition does not.

There is also a volume warning. The page says Claude Opus 4.5 and Claude Opus 4.6 are more responsive to the system prompt than previous models, so prompts written to fix undertriggering may now overtrigger, and it advises dialling back language like "CRITICAL: You MUST use this tool when..." to a plain "Use this tool when...".

Examples: three to five, relevant and diverse

Anthropic calls examples one of the most reliable ways to steer output format, tone and structure, and recommends three to five for best results. Its criteria are that examples be:

  • relevant, mirroring your actual use case
  • diverse enough to cover edge cases so the model does not pick up an unintended pattern
  • structured, wrapped in example tags with multiple examples nested inside an outer tag

Google's guide is blunter about whether to include any at all, under the heading Zero-shot vs few-shot prompts: "We recommend to always include few-shot examples in your prompts. Prompts without few-shot examples are likely to be less effective."

If your prompt carries bulk data, order matters: the page says to put longform data above the query, instructions and examples, and reports that queries at the end can improve response quality by up to 30 percent in tests.

Before and after, on one ticket triage prompt

The first version below leaves the output unspecified and hands the model a step list. The second drops the step list for a definition of done, a closed set of values, and the decisions the business actually cares about, each with the reason it exists, and leaves how to read the ticket to the model.

Here is a made up but realistic case. A support team wants inbound tickets triaged into a queue, and the first version reads like five minutes of work.

You are a helpful support assistant. Read the customer's ticket and
summarize it for the team. Be thorough and professional.

Steps:
1. Read the ticket.
2. Work out what the customer wants.
3. Check whether they seem upset.
4. Decide how urgent it is.
5. Write the summary.

Never be rude.

Both failure modes are present at once. The output is unspecified, so every ticket comes back in a slightly different shape and nothing downstream can parse it. The steps describe reading comprehension, not your business, and the model will follow them straight past any ticket they do not fit.

<identity>
You triage inbound tickets for a B2B billing product. You handle one
ticket per call and produce a record for the on-call queue. You never
reply to the customer.
</identity>

<output_contract>
Return one JSON object and nothing else. No preamble, no code fence.

  summary       string, 30 words or fewer, the problem in the
                customer's own terms
  category      one of: billing, provisioning, auth, export, other
  severity      integer 1 to 4, where 1 means money is not moving
                right now
  account_id    string, or null if the ticket does not state one
  needs_human   boolean
  human_reason  string when needs_human is true, otherwise null
</output_contract>

<constraints>
Never infer account_id from a company name, because a wrong id attaches
this ticket to a different customer's billing history.
Set severity 1 only when the ticket states that a payment or payout is
currently failing, because severity 1 pages someone at night.
When the ticket is ambiguous, set needs_human true and say why in
human_reason. A slow ticket costs less than a misrouted one.
</constraints>

<examples>
<example>
Ticket: "Payouts stuck in pending since Friday. We are Northwind,
account starts with NW I think."
{"summary": "Payouts stuck in pending since Friday",
 "category": "billing", "severity": 1, "account_id": null,
 "needs_human": true,
 "human_reason": "Customer guessed the account id rather than stating it"}
</example>
</examples>

The second version is longer and less prescriptive at once. The step list is gone, replaced by a definition of done, a closed set of values, and the decisions the business actually cares about, each with the reason it exists. How to read the ticket is left to the model, which is the part it is good at.

Three claims this page no longer supports

Three widely repeated claims no longer match Anthropic's Prompting best practices page as it stands. It does not recommend chain of thought prompting, it does not tell you to prefill the assistant turn, and it demotes prompt chaining.

It does not recommend chain of thought prompting. There is no such section, and manual chain of thought appears once, as a fallback for when thinking is off. The reasoning guidance leads with the opposite instruction: prefer general instructions over prescriptive steps, because telling the model to think thoroughly often beats a hand written step by step plan.

It does not tell you to prefill the assistant turn. Anthropic states that starting with Claude 4.6 models and Claude Mythos Preview, prefilled responses on the last assistant turn are no longer supported, and that such requests return a 400 error. The page gives five migration paths, including Structured Outputs for formatting and system prompt instructions for suppressing preambles.

And it demotes prompt chaining. The section still exists, but it opens by saying the models now handle most multistep reasoning internally through adaptive thinking and subagent orchestration, and that explicit chaining remains useful when you need to inspect intermediate outputs or enforce a particular pipeline structure.

One caveat across all of it: the page is model dated, and much of its advice is scoped to a named model. Attribute claims to the model they were measured on and re-check against your own evaluations before applying them elsewhere.

Where this fits if you use Bespoke Prompting

Bespoke Prompting's Prompt build type emits a fixed set of sections in a fixed order: ROLE, CONTEXT, CONSTRAINTS, SUCCESS CRITERIA, APPROACH, TASK. Two of the four parts above are pinned there by name, identity as ROLE and constraints as a required section of their own, and SUCCESS CRITERIA does the work an output contract does: it forces you to say how you would recognize a correct output before you have one in front of you. The fixed order is most of the value: you cannot quietly skip a section because you were in a hurry.

Sources

Suggested internal links