Skip to content New: the Claude Skills catalog — what each skill does and how to install it →
LRN-04 · Manual · chapter 04 · rev. 10/2026

Few-shot prompting

A few-shot prompt shows the model a handful of finished input and output pairs before the real input. Anthropic calls examples one of the most reliable ways to steer format, tone and structure, and nothing about the model has to be retrained.

LevelBeginner to intermediate
Works inChatGPT, Claude, Gemini, APIs
Examples3 to 5 (Anthropic)
CheckedOct 2026

Few-shot prompting means putting a few worked examples into your prompt so the model can copy the pattern. You show a few inputs with the answers you want (Anthropic suggests three to five), then give it a new input. The model learns the task from the prompt alone, and nothing about the model itself changes. This chapter covers where the term comes from, how to build a few-shot prompt, and what research says about choosing the examples.

What few-shot prompting is

The term comes from the 2020 GPT-3 paper, "Language Models are Few-Shot Learners" by Brown et al. The authors used "few-shot" for the setting where the model "is given a few demonstrations of the task at inference time as conditioning", with no weight updates allowed. Each demonstration is a context and a desired completion, such as an English sentence and its French translation. After K of them comes one more context, and the model writes the completion.

The paper also defines one-shot (exactly one demonstration plus a task description) and zero-shot (an instruction with no demonstrations). Its abstract stresses that GPT-3 was applied "without any gradient updates or fine-tuning", with tasks and demonstrations "specified purely via text interaction with the model". That is why the technique is often called in-context learning: the learning happens inside the prompt.

OpenAI's current API guide describes it the same way: few-shot learning "lets you steer a large language model toward a new task by including a handful of input/output examples in the prompt, rather than fine-tuning the model."

Anatomy of a few-shot prompt

A few-shot prompt has four parts: an instruction, the examples, a clear boundary between examples and the real input, and the real input with an empty slot for the answer. Here is one for sorting customer feedback into your own categories.

Try it
Sort each piece of customer feedback into one category: [category 1], [category 2] or [category 3]. Reply with the category name only.

<examples>
<example>
Feedback: [a real feedback line]
Category: [the right category]
</example>
<example>
Feedback: [a second line, different in length or tone]
Category: [the right category]
</example>
<example>
Feedback: [a tricky edge case your team argues about]
Category: [the category you decided on]
</example>
</examples>

Feedback: [the new feedback to sort]
Category:

The XML-style tags come from Anthropic's guide, which recommends wrapping examples in <example> tags, and several of them in <examples>, so Claude can tell them apart from the instructions. The layout carries over: OpenAI's guide also suggests XML tags to mark where each part of a prompt begins and ends, and Google asks for the same structure and formatting in every example.

The same pattern works for rewriting in a house tone, extracting fields into a fixed table, or writing product descriptions to a template. Our ChatGPT prompt library and Claude prompts contain many prompts you can turn into few-shot versions by adding two or three of your own past outputs.

How many examples to use

There is no single number, but the vendors give useful anchors.

  • Anthropic: "Include 3–5 examples for best results." The same guide suggests asking Claude to check your examples for relevance and diversity, or to generate more from your first set.
  • Google: models can often pick up a pattern from a few examples, but "if you include too many examples, the model may start to overfit the response to the examples." Google advises experimenting with the count.
  • The GPT-3 paper: the authors typically used 10 to 100 examples, because that is how many fit in GPT-3's 2,048-token context window. Those were benchmark runs.

A practical rule: start with three, add one for each failure you see, and stop when new examples stop changing the output.

What the examples teach the model

Two research findings change how you should pick examples.

Format matters more than you would expect. Min et al. (2022), "Rethinking the Role of Demonstrations", tested 12 models including GPT-3 on classification and multiple-choice tasks. Randomly replacing the correct labels in the demonstrations "barely hurts performance". What helped was the demonstrations showing the set of possible labels, the kind of input text, and the overall format. In other words, the model reads your examples mostly as a specification of shape.

Order and selection can swing results. Zhao et al. (2021), "Calibrate Before Use", found that "the choice of prompt format, training examples, and even the order of the training examples" can move GPT-3's accuracy from near chance to near state-of-the-art. They trace this to biases toward answers that appear near the end of the prompt or that are common in training data.

Correct labels still matter for your reader, so keep them right. The lesson is to spend your effort on consistent formatting and a balanced mix of examples.

How to choose good examples

  1. Relevant. Anthropic: mirror your actual use case closely. Use real inputs from your own work.
  2. Diverse. Anthropic asks for examples that "cover edge cases" and vary enough that the model does not pick up unintended patterns. OpenAI says to show "a diverse range of possible inputs with the desired outputs".
  3. Consistent. Google: keep the structure and formatting of every example the same, since showing the response format is one of the main reasons to add examples.
  4. Balanced. Given the bias toward recent and frequent answers that Zhao et al. describe, do not put three examples of the same label in a row, and do not end every prompt on the same label.
  5. Separated. Mark where the examples end and the real input begins, with tags or a clear heading.

In API work, OpenAI notes that examples typically go in the developer message. In chat apps, the same examples can live in project instructions or a reusable prompt, which the system prompts guide covers.

Zero-shot vs few-shot

Zero-shotFew-shot
What you giveAn instructionAn instruction plus 3 to 5 worked examples
Time to writeSecondsMinutes, once
Format consistencyVaries with how well you describe itStrong, the examples define it
Custom labels or house toneHard to convey in wordsShown directly
Prompt lengthShortLonger, uses more context
RiskAmbiguityCopying quirks of the examples

Start with zero-shot prompting and move to few-shot when the output is inconsistent. The two are not exclusive: a good few-shot prompt still opens with a clear instruction, written the way how to write prompts describes.

Few-shot examples with reasoning

Examples can also teach a method. Anthropic's guide notes that multishot examples work with thinking: worked examples "shape how Claude approaches similar problems". It suggests presenting each one as the problem, the method to apply, and the expected answer. This is the bridge to the next chapter, chain of thought prompting, where the examples show the steps as well as the result.

Try it
Estimate effort for each task as S, M or L. Show the method, then the answer.

Task: [a past task]
Method: [how you judged it, in one or two sentences]
Answer: [S / M / L]

Task: [another past task of a different size]
Method: [how you judged it]
Answer: [S / M / L]

Task: [the new task]
Method:

FAQ

What is a few-shot prompt?

A prompt that includes a few input and output examples before the real input. The GPT-3 paper defines the few-shot setting as giving the model K demonstrations at inference time with no weight updates.

Is few-shot prompting the same as few-shot learning?

Few-shot prompting is the way you apply it. Brown et al. use "few-shot learning" for a model adapting to a new task from a few demonstrations in its context, and note that it is related to few-shot learning as used elsewhere in machine learning.

How many examples should a few-shot prompt have?

Anthropic recommends 3 to 5. Google warns that too many can make the model overfit to them and suggests testing different counts.

Do the example answers have to be correct?

Keep them correct, but know that format carries much of the effect. Min et al. (2022) found that random labels in demonstrations barely hurt accuracy on the tasks they tested, while format and the label set mattered.

What is one-shot prompting?

The same idea with exactly one example plus a task description. Brown et al. say it most closely matches the way some tasks are explained to people.

Sources

  1. Language Models are Few-Shot Learners (Brown et al., 2020) — arXiv, accessed October 2026
  2. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? (Min et al., 2022) — arXiv, accessed October 2026
  3. Calibrate Before Use: Improving Few-Shot Performance of Language Models (Zhao et al., 2021) — arXiv, accessed October 2026
  4. Prompting best practices — Anthropic, accessed October 2026
  5. Prompt design strategies — Google AI for Developers, accessed October 2026
  6. Prompt engineering — OpenAI API docs, accessed October 2026
Weekly

New prompts in your inbox

One email a week: the best new prompts and one short guide. Unsubscribe any time.