Chain of thought prompting
Chain of thought prompting asks a model to write out intermediate reasoning steps before its answer. It made older models much better at math and logic; current thinking models do this on their own, so the prompt now looks different.
Chain of thought (CoT) prompting is a technique where the prompt gets the model to produce a series of intermediate reasoning steps before the final answer. It was introduced in a 2022 paper and is one of the core techniques in prompt engineering. This chapter covers what the research showed, how to write a CoT prompt, and what changed once models started reasoning by default.
What chain of thought prompting is
The term comes from the paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" by Jason Wei and colleagues, first posted to arXiv in January 2022. The authors define a chain of thought as "a series of intermediate reasoning steps". Their method was few-shot: the prompt included a few worked examples where each answer showed its reasoning, and the model copied that pattern on new questions.
The reported gains were large. In the paper's abstract, a 540B-parameter model prompted with eight chain of thought exemplars reached state of the art accuracy on GSM8K, a benchmark of math word problems, "surpassing even finetuned GPT-3 with a verifier". The authors also wrote that this reasoning ability emerges "in sufficiently large language models", which means small models gained less.
Two follow-ups worth knowing
Zero-shot chain of thought
A few months later, Kojima and colleagues showed that worked examples were not always needed. In "Large Language Models are Zero-Shot Reasoners" (May 2022), adding the phrase "Let's think step by step" raised InstructGPT (text-davinci-002) from 17.7% to 78.7% accuracy on the MultiArith benchmark, and from 10.4% to 40.7% on GSM8K. For more on prompting without examples, see zero-shot prompting.
Self-consistency
Wang and colleagues proposed self-consistency in March 2022. Instead of taking one reasoning path, the method samples several different paths and picks the answer that most of them agree on. The paper reports a 17.9% improvement on GSM8K over standard chain of thought prompting. In a chat app you can approximate it by asking for several independent solutions and a comparison, as in the second prompt below.
When it works and when it does not
Chain of thought helps on tasks with several dependent steps, where one early slip ruins the answer: arithmetic word problems, logic puzzles, multi-step planning, and checking a calculation in a spreadsheet. The original paper tested arithmetic, commonsense and symbolic reasoning.
It helps little, or costs you, in three situations:
- Simple lookups and rewrites. A translation or a summary has no chain to reason through, so extra steps only add length.
- Small models. The 2022 paper tied the effect to model scale.
- Models that already think. This is the biggest change, covered in the next section.
How reasoning models changed the practice
Current models from OpenAI, Anthropic and Google can reason internally before they answer. Their documentation, as of October 2026, changes the advice in clear ways:
- OpenAI. The reasoning best practices guide says to avoid chain-of-thought prompts: since these models reason internally, telling them to "think step by step" or "explain your reasoning" is unnecessary. It also advises trying zero-shot first.
- Anthropic. Claude's guide recommends general instructions over prescriptive steps: a prompt like "think thoroughly" often produces better reasoning than a hand-written step-by-step plan. Manual chain of thought is kept as a fallback for when thinking is off. On the newest Claude models, such as Claude Opus 5.5, thinking is always on, and Anthropic notes that on several current models a prompt asking the model to write out its reasoning in tags may be declined. Worked examples still help, because they shape how Claude approaches similar problems in its own thinking.
- Google. Gemini 3 and 2.5 series models think by default, and the API exposes a
thinking_levelsetting. Google suggests low thinking for fact retrieval or classification and the highest level for advanced coding, math or multi-step planning.
In practice: with a thinking model, describe the goal and how you will judge the answer, and let the model plan. Use explicit CoT with faster, non-reasoning models, or when you need the steps on the page for someone to check.
Which approach for which model
The same task calls for a different prompt depending on whether the model thinks on its own. This summary follows the vendor guides cited below.
| Situation | What to write |
|---|---|
| Reasoning model (OpenAI reasoning models, Claude with thinking on, Gemini 3) | The goal, the constraints and how you will check the answer. No "step by step" instruction. Raise the thinking or effort setting for hard problems if the app or API offers one. |
| Non-reasoning or fast model | An explicit instruction to work through the steps before answering, with the final answer on a separate, labeled line. |
| You need to audit the logic | Ask for the working in the visible answer, in numbered steps, so a person can check each one. On several of the newest Claude models, Anthropic notes that such a request may be declined. |
| A wrong answer is costly | Several independent solutions and a comparison, in the spirit of self-consistency. |
| The model keeps using the wrong method | Two or three worked examples that show the method you want. |
Prompts to copy
The first prompt is classic chain of thought, suited to a non-reasoning model or to any case where you want to read the working. Replace the bracketed blank with your problem.
Solve the problem below. First list the facts given in the problem. Then work through the steps one at a time, showing each calculation. Check the result against the facts. Finish with one line that starts with "Answer:" and contains only the final answer. Problem: [paste the problem]
The second prompt imitates self-consistency in a chat window. It is useful when a wrong answer is expensive, such as a pricing or scheduling calculation.
Solve this problem three times, each time using a different method. Keep the three solutions independent of each other. Then compare the three final answers. If they agree, state the answer. If they differ, find the step where they diverge, decide which method is correct and explain why in two sentences. Problem: [paste the problem]
For practice material, the data prompts and education prompts in our library include calculation and tutoring tasks where showing the working is the point.
Chain of thought with examples
The original method was few-shot. You can still use it: write two or three solved problems where each answer shows its steps, then add the new problem. Anthropic suggests presenting each example as a problem, the method to apply, and the expected answer. How to choose good examples is covered in the few-shot prompting chapter. When the task depends on long documents rather than reasoning, the bottleneck is usually what the model can see, which is the subject of context engineering.
FAQ
What is chain of thought in AI?
It is a model's written sequence of intermediate reasoning steps before a final answer. Chain of thought prompting is any prompt that gets the model to produce those steps, either with worked examples or with an instruction.
Does "let's think step by step" still work?
On non-reasoning models it is still a valid zero-shot technique, as shown by Kojima et al. On reasoning models, OpenAI calls it unnecessary, and Anthropic recommends general instructions such as "think thoroughly" over prescriptive steps.
What is the difference between chain of thought and self-consistency?
Chain of thought produces one reasoning path. Self-consistency samples several paths and takes the most common answer, which the 2022 paper reports as a further accuracy gain on math benchmarks.
Who wrote the chain of thought paper?
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter and co-authors. It was first submitted to arXiv on January 28, 2022, with the latest version dated January 10, 2023.
Sources
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Wei et al., arXiv, accessed October 2026
- Large Language Models are Zero-Shot Reasoners — Kojima et al., arXiv, accessed October 2026
- Self-Consistency Improves Chain of Thought Reasoning in Language Models — Wang et al., arXiv, accessed October 2026
- Reasoning best practices — OpenAI, accessed October 2026
- Prompting best practices — Anthropic, accessed October 2026
- Gemini thinking — Google AI for Developers, accessed October 2026