Context engineering for AI agents is the practice of deciding what the model sees at every step of a run: its instructions, the live task state, retrieved evidence, memory and tool definitions. Prompt engineering writes the instructions. Context engineering decides everything else that reaches the model, and what it costs to keep it there.

Key takeaways
- Context engineering decides what the model sees at each step of an agent’s run. Prompt engineering is the part of it that writes the instructions.
- A bigger context window doesn’t fix it. Models accept far more tokens than they can use dependably, and the gap is often more than ten-fold.
- You’ve got four ways to keep the window under control, and each costs something different: keep an item loaded on every step, fetch it only when a step needs it, compress old history into a summary, or give a chunk of work to a separate sub-agent.
- The test that matters is whether your agent can still recover a hard rule after its history has been compressed. With Oracle AI Agent Memory you can pin that rule as a guideline and check it’s still in the context card at the step that needs it.
A support agent is thirty-eight steps into a refund.
At step one it was given the policy: nothing over two hundred pounds without a human. At step twenty-two the conversation got long, the harness squeezed the history into a summary to free up space, and the policy line didn’t make it into the summary. At step thirty-eight it refunds nine hundred pounds and reports, honestly as far as it can tell, that it followed policy.
Nobody on the team chose to drop that rule. The harness did it on its own, because it was set to summarise whenever the history crossed a token threshold, and nothing told it which line was a rule it couldn’t afford to lose.
That isn’t a made-up failure. In June 2026 a researcher built a benchmark called ConstraintRot to test exactly this: give agents a policy about which tool actions are off limits, run them long enough for their history to be compacted, and count how often they break the policy afterwards. Across 1,323 runs and seven model families, the split was stark. When the policy survived the summary, violations stayed at zero. When the summary dropped it, they hit 38 per cent (Governance Decay). It’s a single-author preprint that nobody has replicated yet, so I’d treat the number as a warning rather than a law. The fix it tested is almost boring, though: keep the policy somewhere the summariser can’t touch, and violations went back to zero.
That’s the job context engineering does.
Think of an engineer’s service van. Some kit rides in the back all week. Some you drive back to the depot for. Some you write on the job sheet before you bin the packaging, and some jobs you send the apprentice to do. Four choices, four different prices, and none of them is the advanced version of the others. You pay one of them on every callout.
What is context engineering, and how is it different from prompt engineering?
Context engineering is the practice of deciding, at each step of an agent’s run, which tokens the model actually sees across instructions, live task state, retrieved evidence, memory and tool definitions, and which of four operations pays for each item: load it eagerly, fetch it just in time, compact it, or hand it to a sub-agent.
Prompt engineering shapes the instructions. Context engineering governs everything else that reaches the model, across the hundreds of steps a harness assembles on its own. Anthropic calls it “the natural progression of prompt engineering” (Effective Context Engineering for AI Agents), and Andrej Karpathy’s version, the one that got the term moving in June 2025, is “the delicate art and science of filling the context window with just the right information for the next step” (Karpathy on X).
| Prompt engineering | Context engineering | |
| The question it answers | How should I word the instructions? | What should the model see at this step, and what does keeping it there cost? |
| Scope | One prompt or template | Every step of a run: instructions, state, evidence, memory and tools |
| When it happens | Written before the run | Decided while the run is happening, often by the harness |
| How it fails | An ambiguous or conflicting instruction | The right information is missing, buried or dropped at the step that needed it |
| How you debug it | Read the prompt | Inspect what was actually in the window at the step that failed |
The working artefact is a context budget: a per-step allocation of finite input tokens across those five categories, plus the rule that decides which operation pays for each item. A sample context window for the refund step might look like this:
| Budget line | Example content | Policy |
| Instructions | Refund limit and escalation rule | Pinned, never summarised |
| Live task state | Case ID, customer ID, requested amount and last action | Loaded for this step |
| Retrieved evidence | The exact policy clause and its source URL | Fetched when the refund decision is made |
| Memory | Prior customer commitments, with record IDs | Scoped by user, agent, thread and time to live |
| Tool definitions | Only the payment and refund tools this step can use | Loaded by route, not as a whole catalogue |
Treat it as an allocation problem rather than a size problem, because every action changes what the next one needs. A lookup justifies one passage. An irreversible tool call needs the original instruction, current permissions and exact identifiers, all present at once. Anthropic’s own shorthand for the goal is finding “the smallest possible set of high-signal tokens” that makes the outcome you want most likely (Effective Context Engineering for AI Agents).
Why doesn’t a bigger context window solve this?
Because accepting tokens and using them dependably are two different properties, and only the first one is printed on the model card. Claude Opus 5.5 lists a one million token context window (Anthropic models overview), and it isn’t alone: GPT-4.1 and several Gemini models advertise a million too. That capacity is real. It just isn’t a million tokens of reliable attention, and the studies that measure the gap keep finding it’s large.

Read each block as one model measured against itself. GPT-4o accepts 128,000 tokens and held up to about 8,000 once the question and the evidence stopped sharing the same words (NoLiMa). Llama 3.1 70B held up to 64,000 on RULER (RULER) and to 2,000 on NoLiMa, which is the same model giving you a thirty-two-fold difference depending on how hard the test is. Nobody has published the same measurement for Claude Opus 5.5, so I’m not going to guess at its number.
The industry has a name for this now. Anthropic and Chroma both call it context rot: as the input grows, the model’s ability to recall what’s in it drops (Effective Context Engineering for AI Agents). Chroma tested 18 models and found that “performance grows increasingly unreliable as input length grows”, even on simple tasks (Context Rot).
Three things pull the measured bars down, and they fail independently. Position: models use evidence most reliably near the start or the end of a long input and least reliably in the middle (Lost in the Middle). Difficulty: RULER adds multiple needles, aggregation and multi-hop tracing, and half the models it evaluated fell over well below the window they advertised. Literal overlap: NoLiMa removes the shared phrasing between question and evidence, and effective lengths collapse.
Then there’s the result that removes the last excuse. Give a model perfect retrieval, put the evidence immediately before the question, and performance still degrades as the input grows (Context Length Alone Hurts LLM Performance Despite Perfect Retrieval).
So the effective context window is a measurement you take, not a specification you read. A bigger van still gets loaded badly.
What are the four operations, and what does each one cost?
There are four, and they’re alternatives rather than a ladder. Just-in-time fetching isn’t more advanced than eager loading, and a sub-agent isn’t where you graduate to. Each one buys tokens back at a different price.
| Operation | What it buys | What it costs | How it fails |
| Load eagerly | Present from the first token, no retrieval round trip | Occupancy on every step, needed or not | Tool catalogues crowd out the current evidence |
| Fetch just in time | Occupancy proportional to need | A round trip, and a discovery decision | The agent never learns the source exists |
| Compact | Long horizons in a small window | Irreversible omission, unless the original stays addressable | A constraint nobody flagged disappears |
| Hand to a sub-agent | The parent never sees the trace | A claim in place of evidence | Inherited credentials, or a confident summary of a failed search |
If you’ve read LangChain’s write, select, compress and isolate (Context Engineering for Agents), this is the same territory cut a different way. Eager loading and just-in-time fetching are both ways to select, compaction is compress, and the sub-agent is isolate. Write, which means saving notes and state outside the window, isn’t an operation here because it’s what makes the other four safe: you can only compact or fetch what’s been written somewhere first. I’ve grouped them by what each one costs you rather than by what it does, because the cost is what you’re actually choosing between.

There’s one cost that cuts across all four, and it’s the one teams notice last: the prompt cache. Providers cache the unchanged front of your context, and cached tokens are far cheaper. The Manus team reported cached input at 0.30 USD per million tokens against 3 USD uncached on Claude Sonnet, a ten-fold gap, and called the cache hit rate “the single most important metric for a production-stage AI agent” (Context Engineering for AI Agents: Lessons from Building Manus). So a stable, eagerly loaded prefix pays for itself, while every compaction or reshuffle near the top of the window throws the cache away and you pay full price again.
When should you load something eagerly?
Load eagerly when the item must be present for a decision whose timing you can’t predict, and when its absence is unrecoverable. That’s the control layer: privileged instructions, the user’s objective, mandatory constraints and the output contract.
Instruction-hierarchy training makes a model better at ignoring conflicting lower-priority text (The Instruction Hierarchy), and better is the right word: it improves the odds, it doesn’t enforce anything. The constraint still has to be in front of the model when the decision lands.
The cost is that eager material is charged on every step, and tool definitions are where teams pay it without noticing, because most frameworks load every tool description by default. Across 856 tools on 103 Model Context Protocol servers, almost every description carried a defect, and rewriting them raised task success a little while pushing execution steps up by two thirds (Model Context Protocol Tool Descriptions Are Smelly!, a 2026 preprint). More description isn’t free description.
Keep on the van what you’d refuse to drive off without. Everything else is a trip to the depot.
When should you fetch just in time?
Fetch just in time when the material is large, changes often or is rarely needed, and when the agent will know it needs it. Hold a reference to it and load it at the step that uses it.
What the evidence gives you is a routing result rather than a verdict. Self-Route tries retrieval first and escalates only what retrieval can’t answer, holding roughly full-context accuracy on well under half the tokens (Retrieval Augmented Generation or Long-Context LLMs?). Order matters as much as volume: DOS RAG puts retrieved passages back into their source-document order and beats vanilla retrieval on about a third of the tokens (Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models).
The cost is discovery risk. The agent has to know a source exists and notice when a fetch came back incomplete, and tool retrieval is weak at this: across roughly 43,000 tools, conventional retrieval strength didn’t predict picking the right one, and worse retrieval lowered completion (Retrieval Models Aren’t Tool-Savvy).
Observation shaping is the same operation pointed at outputs. Moving a two-hour transcript through model-visible tool calls costs tens of thousands of tokens, where a code-mediated version of the same job costs a couple of thousand (Code Execution with MCP, a vendor example rather than a controlled benchmark). Do the join in code and hand the model the answer.
When should you compact?
Compact when the history has passed a natural boundary, such as a subgoal that finished, was abandoned or was revised, and when everything you’re about to drop is still addressable somewhere else. Compacting on a token threshold alone is how you get the refund at step thirty-eight.
Every major lab now ships compaction as a feature, which tells you how common the problem is. OpenAI’s Agents SDK has a compaction session that, once it runs, “clears the underlying session and rewrites it with the reduced item list” (OpenAI Agents SDK: Sessions). Google’s Agent Development Kit summarises older session history on an interval you configure (ADK Context Compaction). Anthropic reported that its context editing feature let a 100-turn web search evaluation finish workflows that would otherwise have run out of context, while cutting token use by 84 per cent (Managing context on the Claude Developer Platform). Read the OpenAI line again, though. Once the session is rewritten, the original is gone from it, unless you kept a copy somewhere else.
And it does work, when it’s done carefully. On a set of expense-processing tasks, keeping only the five most recent tool interactions beat full history on both accuracy and tokens, and adding a summary on top beat it again (Less Context, Better Agents, narrow and unreplicated: evidence for selective trimming, not for the number five).

The cost is that you decide what to leave out before you know what the next question will be, which is why masking and compaction aren’t the same operation. Masking hides an observation behind a marker and a retrieval handle, so the original comes back exactly. Compaction on its own didn’t preserve enough state to restart coding-agent sessions, which is why Anthropic added progress files and structured feature records alongside it (Effective Harnesses for Long-Running Agents).
How do you compress context without losing critical information?
Pin the things you can’t afford to lose outside the summariser, keep the raw records addressable, and only summarise what you’ll never need back word for word. In practice that means four habits:
- Pin constraints. Policies, permissions and commitments live in a store the summariser never rewrites, and get reloaded on every step. That’s the fix the Governance Decay study tested.
- Summarise the prose, keep the identifiers. Case IDs, timestamps, amounts, permissions and pointers to raw events survive every compaction verbatim.
- Derive each summary from raw records, not from the last summary. A summary of a summary drifts, and nothing warns you when it does.
- Keep a written plan near the end of the window. Anthropic calls this structured note-taking, where the agent writes notes that persist outside the context window. Manus rewrites a todo list on every step so that it’s “reciting its objectives into the end of the context”, where attention is strongest.
The job sheet isn’t the job. Bin the packaging, keep the serial number.
When should you hand work to a sub-agent?
Hand off when a phase produces a long exploratory trace with a short conclusion: a corpus sweep, or a search whose dead ends the parent never needs to see. Context Folding folds finished trajectories into their outcomes and reports active contexts up to ten times smaller on long-horizon work (Scaling Long-Horizon Agent via Context Folding).
The first cost is that you’ve swapped observation for testimony, and an agent’s account of its own success isn’t a verified account. τ-bench scores the final database state rather than the narration, and leading function-calling agents completed fewer than half its tasks (τ-bench).
The second cost is money. Anthropic’s own numbers put multi-agent systems at “about 15× more tokens than chats” (How we built our multi-agent research system). And not everyone thinks the trade is worth it. Cognition’s argument against multi-agent setups is that “actions carry implicit decisions, and conflicting decisions carry bad results”: two sub-agents working from different slices of context will make choices that don’t fit together (Don’t Build Multi-Agents). I think both sides are right about different jobs. A read-only research sweep hands off well. Two agents editing the same thing in parallel usually doesn’t.
So make the handoff typed rather than narrative: the assigned goal, constraints, decisions taken, artefact identifiers, evidence with provenance, known failures and the next expected action, with record IDs copied verbatim and large evidence behind a reference. The A2A specification gives you an envelope for exactly that, and its fields prove continuity rather than authority (A2A Protocol Specification).
The last cost is quieter. Isolation isn’t isolation if every child inherits the parent’s credentials, and AgentDojo tests exactly that kind of redirection, with hostile text arriving inside ordinary tool output (AgentDojo). Scope the child agent’s credentials to the task you’ve given it.
How do you know the budget is working?
Measure recovery after reduction, not task success on the happy path. The question that matters is whether the agent can recover an exact, source-supported constraint after several compactions and a contradictory update.
The failure is already common. LongMemEval separates extraction, multi-session reasoning, temporal reasoning, knowledge updates and abstention, and reports around a 30 per cent accuracy drop across sustained interaction (LongMemEval).
Make the test boring. Log useful-token density, provenance coverage, exact constraint recovery and rehydration success on every run, and track cost per resolved task next to them, because a policy that recovers every constraint by loading everything eagerly has only moved the cost, not removed it. Then keep the model and the task fixed while you change only the budget policy. If the model, the retriever and the context ceiling all change at once, you won’t know which choice helped.
How does Oracle AI Agent Memory fit into context engineering?
It sits underneath all four operations, as the addressable original that makes reduction reversible. Compaction is only recoverable when the raw record still exists with a stable identifier, and just-in-time fetching is only safe when tenant and permission filters run before retrieval and ranking, not after a global top-k has already been picked.
Oracle AI Database 26ai holds dense and sparse vectors, vector indexes and SQL similarity operations beside relational and JSON data, so a semantic search can carry an ordinary tenant predicate instead of running in a separate service (AI Vectors and Semantic Search). Oracle AI Agent Memory builds on that with typed records, including messages, memories, guidelines, facts and preferences, each scoped by user, agent and thread, with optional time to live. User and agent profiles sit alongside those records rather than behaving like short-lived thread memory (Stores).
The piece that matters most for this article is the context card. Oracle’s docs describe it as “compact context about a conversation that an agent can use when generating a response”, built from a thread summary, relevant stored messages and relevant memories (Agent Memory: context card content). In other words, it’s the assembly step: the thing that decides what the model sees at step thirty-eight.
Here’s the refund agent from the top of this article, rebuilt so the £200 rule can’t be summarised away. It runs against Oracle AI Agent Memory 26.8 (pip install oracleagentmemory) and an Oracle AI Database 26ai instance, with an embedding model and an LLM for the summary.
import oracledb
from oracleagentmemory.core import OracleAgentMemory, SchemaPolicy
from oracleagentmemory.core.embedders.embedder import Embedder
from oracleagentmemory.core.llms.llm import Llm
pool = oracledb.create_pool(user="YOUR_DB_USER", password="YOUR_DB_PASSWORD",
dsn="localhost:1521/FREEPDB1")
memory = OracleAgentMemory(
connection=pool,
embedder=Embedder(model="YOUR_EMBEDDING_MODEL"),
llm=Llm(model="provider/model_id"),
schema_policy=SchemaPolicy.CREATE_IF_NECESSARY,
)
USER, AGENT = "customer_4471", "refund_agent"
# 1. Pin the rule as a guideline record, outside the conversation history.
memory.add_memory(
"Refunds over £200 need human approval before any payment tool is called.",
memory_type="guideline", user_id=USER, agent_id=AGENT,
)
# 2. The long run. Every message is stored as a raw record with its own ID.
thread = memory.create_thread(user_id=USER, agent_id=AGENT)
thread.add_messages([
{"role": "user", "content": "I'd like a refund for order 4471, please."},
{"role": "assistant", "content": "Of course. Let me pull up the order."},
# ... thirty-odd more steps ...
{"role": "user", "content": "Actually, can you refund the full £900?"},
])
# 3. Build what the model sees at step 38: a summary instead of the full
# transcript, the last few turns, and the most relevant records, with one
# slot held back for guidelines so ranking can't squeeze the rule out.
card = thread.get_context_card(
max_relevant_results=6,
min_relevant_results_by_type={"guideline": 1},
max_recent_messages=4,
)
print(card.content)
# 4. The test from the previous section, as one line.
assert "£200" in card.content, "the refund rule didn't survive"
Map it back to the four operations and you can see each one doing its job. The guideline with a reserved slot is eager loading: the rule is in front of the model on every card, whatever the summary says. The relevance search behind the card is just-in-time fetching, scoped to this user and this agent before anything is ranked. The summary is compaction, but the raw messages are still in the database with their IDs, so thread.get_messages() can bring any of them back exactly. And because every record is scoped by agent_id, a sub-agent can get its own scope rather than the parent’s whole history. Oracle doesn’t use these four words, to be clear. The mapping is mine.
I want to be straight about two things. The reserved slot requests at least one guideline, but the rule may still be absent from the card, so the final line checks for it. And the card only makes sure the model sees the rule. It doesn’t make the rule true. Oracle’s security guidance is blunt about this: memory-derived content must never approve privileged actions or “bypass policy”, and a caller-supplied user_id “is a scoping value, not proof of identity” (Security Considerations). So the payment tool should still refuse £900 without a human, whatever the model decides. Release 26.8 also adds end-user isolation through Oracle Deep Data Security, so the database, not your application code, can decide which signed-in user’s memories a request can reach (What’s New in 26.8).
If you want the longer version, the Oracle AI Developer Hub notebook runs the stores, retrieval, summaries and context cards end to end, and it’s the fastest way to watch a card being built (oracle_agent_memory_developer_guide.ipynb). You get the persistence, retrieval and audit primitives. The budget policy is still yours to write and prove.
Where is this going?
Three things are unsettled, and each one has something that would settle it.
The first is whether compaction can be made safe by policy alone. The evidence that it can’t is a single-author preprint nobody has replicated (Governance Decay), and what would settle it is the same protocol run across two or three harnesses by people who didn’t write the first one. Until then, pinning constraints outside the summariser is cheap insurance rather than a proven requirement.
The second is whether the benchmarks start scoring recovery. Most score task success on a fresh run, which is exactly the measurement that hides this failure. LongMemEval is the nearest exception, and it targets chat assistants rather than tool-using agents (LongMemEval). As far as I can find, a benchmark that compacts an agent’s history mid-run and then asks for the constraint back doesn’t exist yet.
The third is whether the operations get standardised. The handoff has an envelope (A2A Protocol Specification), and every lab now ships compaction, but each one configures it differently and none of them agrees on what must survive it. Masking, offloading and pinning have no shared vocabulary at all. That’s my read, not a cited claim.
My expectation, and this is opinion: windows keep growing, effective lengths keep lagging behind them, and the gap stops being a research curiosity the first time it causes a real financial, safety or compliance incident at a regulated company.
Nobody stocks a van properly by buying a bigger van. You learn the route, you get the depot trips wrong a few times, and you find out which part you should never have taken off the shelf.
FAQ
What is context engineering, and how is it different from prompt engineering? Context engineering is deciding what an AI agent’s model sees at each step of a run: instructions, live task state, retrieved evidence, memory and tool definitions. Prompt engineering writes the instructions, which makes it one part of context engineering. Anthropic describes context engineering as “the natural progression of prompt engineering” (Effective Context Engineering for AI Agents).
How do I compress context for an AI agent without losing critical information? Pin constraints, permissions and commitments in a store the summariser never rewrites. Summarise prose but keep identifiers, timestamps and pointers to raw events verbatim. Derive each summary from raw records rather than from the last summary, and test that a pinned rule is still recoverable after several compactions. Where constraints survived summarisation in the Governance Decay study, violations stayed at zero; where they were dropped, they reached 38 per cent (Governance Decay).
Is a context budget the same as a token limit? No. A token limit is a ceiling the serving system enforces. A context budget is your policy for allocating tokens beneath that ceiling across instructions, state, evidence, memory and tool definitions, and it should bind well before the ceiling does, because measured performance falls with input length even when retrieval is perfect (Context Length Alone Hurts LLM Performance Despite Perfect Retrieval).
What is context rot? Context rot is the drop in a model’s ability to use information as its input grows. Anthropic describes it as the finding that “as the number of tokens in the context window increases, the model’s ability to accurately recall information from that context decreases” (Effective Context Engineering for AI Agents), and Chroma measured it across 18 models (Context Rot).
Does retrieval replace long context, or the other way round? Neither. It’s a routing decision that varies with model, task, document length and chunk properties across the 2,326 evaluated cases in LaRA (LaRA). Retrieve first, escalate on uncertainty, and keep retrieved passages in their source-document order.
Is masking just cheap compaction? No, and the difference is reversibility. Masking replaces an observation with a marker plus a retrieval handle, so the original can be restored exactly, while summarisation replaces it with prose that can’t be. Masking doesn’t remove the length penalty on its own either: degradation persisted when distractors were masked from attention rather than removed (Context Length Alone Hurts LLM Performance Despite Perfect Retrieval).
What single number should I put on a dashboard? Exact constraint recovery rate after N compaction cycles, with cost per resolved task beside it. Task success on a fresh run hides the failure entirely, and sustained-interaction accuracy already falls around 30 per cent without it showing up in single-turn tests (LongMemEval).
Where can I see this working in code? The demo above pins a rule as a guideline and checks it survives into the context card. For the full walkthrough, the Oracle AI Developer Hub agent memory notebook runs the stores, the retrieval call and the context card end to end (oracle_agent_memory_developer_guide.ipynb).
