Skip to content New: the Claude Skills catalog — what each skill does and how to install it →
LRN-06 · Manual · chapter 06 · rev. 10/2026

Context engineering

Context engineering is choosing everything a model sees when it answers: instructions, examples, documents, tools, memory and conversation history. More context is not automatically better, so the job is to include what the task needs and leave out the rest.

LevelAdvanced
Works inChatGPT, Claude, Gemini, agents
Reading time7 min
CheckedOct 2026

Context engineering is the practice of deciding what information a language model has in front of it at the moment it answers. The prompt you type is one part. The rest is the system prompt, examples, attached documents, tool definitions, tool results, memory and the conversation so far. This chapter explains each part, how context engineering relates to prompt engineering, and the techniques that keep long chats and agents accurate.

What context engineering means

Two definitions are useful. Harrison Chase of LangChain wrote in June 2025 that "context engineering is building dynamic systems to provide the right information and tools in the right format such that the LLM can plausibly accomplish the task." Anthropic's Applied AI team, in a September 2025 engineering post, described it as the set of strategies "for curating and maintaining the optimal set of tokens (information) during LLM inference."

Both point at the same shift. When a model fails, Chase argues, the cause is either a model that is not capable enough or a model that "was not passed the appropriate context to make a good output", and he expects the second cause to grow as models improve.

Context engineering vs prompt engineering

Prompt engineeringContext engineering
Main questionHow should the instructions be worded?What should the model see at all?
Unit of workOne prompt, one answerEvery turn of a conversation or agent run
Typical toolsRole, task, rules, format, examplesRetrieval, memory, tool design, summaries, sub-agents
Where it matters mostChat apps, single tasksLong chats, assistants, coding agents

Chase treats prompt engineering as a subset of context engineering. Anthropic writes that it views context engineering "as the natural progression of prompt engineering." The skills from earlier chapters still apply: a well-written system prompt is part of good context.

What goes into the context window

Anthropic's documentation describes the context window as "all the text a language model can reference when generating a response, including the response itself", a working memory separate from training data. For the Claude API it lists what counts: the system prompt, every message (including tool results, images and documents), the tool definitions, and the output, including thinking. As of October 2026, Anthropic lists a 1M-token context window for its current models and 200k tokens for some others, such as Claude Sonnet 4.5. Google's long context page describes Gemini context windows of 1 million tokens or more.

The parts, and what Anthropic's engineering post advises for each:

  • Instructions. The post asks for a system prompt at the right altitude: specific enough to guide behavior, flexible enough to give the model strong heuristics. It warns against hardcoding brittle logic and against vague, high-level guidance. The system prompts guide covers how to write one.
  • Examples. A few diverse, canonical examples instead of a list of every edge case. The post calls examples the "pictures" worth a thousand words. See few-shot prompting.
  • Documents. Load data when it is needed. Agents can keep "lightweight identifiers" such as file paths or links and fetch the content at runtime, instead of loading everything up front.
  • Tools. Each tool should be self-contained, robust to error and clear about its purpose, with no overlap between tools that would leave the model unsure which to call.
  • Memory and history. Conversation history accumulates turn by turn. LangChain's examples include short-term memory (a summary of the conversation) and long-term memory (user preferences retrieved from earlier sessions).

Why more context can make answers worse

Anthropic's documentation names the problem context rot: "As token count grows, accuracy and recall degrade". The engineering post adds that models have an "attention budget", and every token spends part of it. The post links this to the transformer architecture, where every token relates to every other token, so a longer context spreads attention thinner.

One layout rule helps when the context is long. Anthropic recommends putting long documents at the top and the question at the end, and reports that this improved response quality by up to 30 percent in its tests on complex, multi-document inputs. Google's long context guide gives the same advice for Gemini.

Techniques for long tasks

Anthropic's post describes three techniques for work that outgrows one context window:

  1. Compaction. Summarize the conversation as it nears the limit and continue from the summary. The Claude API offers server-side compaction for this, in beta for Claude 4.6 and later models.
  2. Structured note-taking. The agent writes notes to a file outside the context window, such as a progress log, and reads them back when needed.
  3. Sub-agents. Focused agents work in their own clean context windows and return a condensed summary, "often 1,000-2,000 tokens", to a coordinating agent.

Claude Skills apply the same principle to instructions. Anthropic's documentation says only a skill's name and description (about 100 tokens) sit in context until it is triggered, and its full instructions load only when a request matches. The Claude Skills chapter explains how they work, and the Claude Code cheat sheet is a one-page reference for the coding agent where long sessions are most common.

Context engineering in a chat app

You do not need an API to apply this. Chat apps give you some of the same controls. In Claude, Projects hold their own chat history, uploaded documents and project instructions. Anthropic's help center says free accounts can create up to five projects, and on paid plans a project switches to retrieval when its knowledge approaches the context limit, expanding capacity "by up to 10x".

When a chat gets long and the answers start to drift, do compaction by hand. Ask for a handoff note, start a new chat, and paste the note in.

Try it: handoff note
We are going to continue this work in a new chat. Write a handoff note I can paste there.

Include:
- The goal of the work in one sentence.
- Decisions we made, each with the reason.
- Facts, figures and names the next chat must keep exactly.
- What is finished and what is still open.
- The next step.

Leave out greetings, dead ends and anything we later replaced. Keep it under [word limit, e.g. 300] words.

For a fresh task, assemble the context deliberately instead of pasting everything you have. The template below follows the order the vendor guides recommend: material first, instructions and question last.

Try it: context pack
<documents>
[paste only the documents this task needs]
</documents>

<background>
Who I am: [your role]
Who the result is for: [audience]
What I already know or decided: [key facts]
</background>

<task>
[the one thing you want done, with a verb]
Use only the documents above. If they do not contain the answer, say so.
Format: [output shape]
</task>

Anthropic recommends XML tags like these for Claude, and OpenAI's prompt engineering guide also suggests XML tags to mark where each part of a prompt begins and ends. To practice on real tasks, try the business prompts with your own documents attached.

FAQ

What is context engineering in AI?

It is the work of choosing and arranging everything a model sees during a task: instructions, examples, documents, tools, memory and history. LangChain defines it as building systems that give the model the right information and tools in the right format.

Is context engineering replacing prompt engineering?

It extends it. Anthropic describes context engineering as the natural progression of prompt engineering, and LangChain counts prompt engineering as one part of it. Clear instructions still matter, and they now share the window with documents, tools and history.

Does a bigger context window remove the need for context engineering?

No. Anthropic's documentation says more context is not automatically better, because accuracy and recall degrade as the token count grows. A 1M-token window raises the ceiling, while curation still decides the quality.

Where should I put the question in a long prompt?

At the end, after the documents. Both Anthropic and Google recommend this layout for long inputs.

Sources

  1. Effective context engineering for AI agents — Anthropic, accessed October 2026
  2. The rise of "context engineering" — LangChain, accessed October 2026
  3. Context windows — Anthropic, accessed October 2026
  4. Prompting best practices — Anthropic, accessed October 2026
  5. Agent Skills — Anthropic, accessed October 2026
  6. What are projects? — Claude Help Center, accessed October 2026
  7. Prompt engineering — OpenAI, accessed October 2026
  8. Long context — Google AI for Developers, accessed October 2026
Weekly

New prompts in your inbox

One email a week: the best new prompts and one short guide. Unsubscribe any time.