By Komodor
Picture a Monday at 03:10 UTC. An agent investigating an out-of-memory alert on payment-service works out the pattern: it runs out of memory every Monday between 03:00 and 04:00, and the spike lines up with the batch reconciliation job. The next Monday the alert fires again. A different agent picks it up and starts from zero, because nothing it can see holds what the first one learned.
Most writing about context engineering for AI agents focuses on one agent’s context window. In production operations, the harder problems sit between agents and across time: what another agent already learned, whether this one can use it, and whether it’s still true. This guide covers what to feed an agent, what to leave out, and why operations teams should treat context as shared infrastructure.
Key Takeaways
- More context isn’t better. Accuracy degrades as input grows, and irrelevant or stale context actively misleads.
- Reliable agents keep a small, stable core in context and fetch everything else on demand through tools, search and skills.
- In operations, context is shared infrastructure. What one agent learns should be reviewed, made available to the agents that need it, and retired when it goes stale.
- Specialist agents work best on shared context. They read from and write to the same layer, and hand off findings rather than transcripts.
What Is Context Engineering for AI Agents?
Context engineering is the discipline of deciding what information an AI agent receives, when it receives it, and what it never sees. Prompt engineering shapes the instruction. Context engineering shapes everything around it: the tools, documents, memories and live data the model reasons over at each step of a task.
Anthropic’s engineering team describes the goal as finding “the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome.”
| Prompt engineering | Context engineering | |
|---|---|---|
| Scope | The instruction | Everything the model sees |
| When | Written once | Decided at every step |
| Typical failure | An ambiguous instruction | The right instruction with the wrong, missing or stale information |
| Operations example | “Investigate this alert.” | Which logs, which runbook, which past incident, which recent change |
What Actually Counts as Context
For an operations agent, context is anything the model reasons over beyond its own training. That covers its instructions, the skills and tools it can use, curated knowledge, memory from past runs, live output from the systems it queries, and findings handed over by other agents. Each comes from a different place and goes stale at a different rate.
- Instructions: the agent’s job, scope and output format. Small, stable, always present.
- Environment rules: constraints manifests don’t show, such as a workload holding EU customer data that must never fail over to a US cluster.
- Skills: reusable procedures, such as how to investigate a failed rollout.
- Knowledge: curated runbooks, postmortems, service docs and policies.
- Memory: what previous runs learned, like the payment-service pattern above.
- Live tool output: logs, metrics, events and deploy history pulled during the run. Usually the largest and most perishable.
- Handoffs: findings passed along by other agents in the same workflow.
Operations adds one more dimension, time: what was true at 14:00 may not be true at 14:05, so change history belongs in the picture too.
Who Decides What Gets Loaded, and When
Three parties load context: the system, the model and humans. The system loads a small, stable core before the run starts, the model fetches more on demand through tools, search and skills, and humans add questions, curated documents and approvals. Good designs keep the first set small and make the second cheap.
| Who loads it | What | Example |
|---|---|---|
| The system | Instructions, hard environment rules, an index of available skills, relevant memories at planning time | An orchestrator starts an investigation from prior runs, proven causes and the fixes that held |
| The model | Tool calls, knowledge search, the full content of a skill | Midway through a run, the agent decides it needs the rollback procedure and opens it |
| Humans | Questions in chat, the curated knowledge base, approvals | An on-call engineer asks whether the latency started after the 14:02 deploy |
On the Komodor Agentic Operations Platform, an agent gets a short index of its skills plus the content of the most relevant few, and opens the rest when it decides it needs them. Knowledge is fetched only on demand: the agent calls a search tool when it needs background and gets back up to ten passages, each with its section heading, so it can cite what it used.
The principle for agent context engineering: anything always loaded is paid for at every step, so reserve it for what must never be missed. A knowledge base answers “How do I restart the payment service?” An always-present rule answers “Do not restart the payment service during business hours.” We’ve written before about the difference between always-present context and retrieved context.
Why More Context Makes Agents Worse
Beyond a point, adding context makes an agent less reliable. Models use long inputs unevenly, irrelevant material distracts them, stale material misleads them, and every token adds cost and latency at each step. Bigger context windows raise the ceiling, but they don’t change the curve.
Chroma’s research on “context rot” tested 18 models and found performance varies significantly with input length, even on simple tasks, and that “even a single distractor reduces performance relative to the baseline.” Komodor Co-Founder and CTO Itiel Shwartz has made the same point about Kubernetes investigations: “For complicated domains, you can’t really reach that close to the context window because your LLM will start to hallucinate.”
The failure modes look like this in practice:
- Distraction: unrelated history, such as last month’s OOM on a different service, pulls the reasoning off course.
- Staleness: as the Komodor docs put it, “A stale runbook is worse than a missing one, because it retrieves.”
- Confident wrong answers: the model reasons plausibly over the wrong evidence.
Here’s how that shows up in one example run graded by the platform’s LLM-as-a-judge evals. The investigation scored 4.6 out of 5 for root-cause quality and 4.4 for evidence discipline, and still failed: a logs.search call at step 22 pulled 9.8 MB over a 24-hour window. Tool-use efficiency scored 1.8, dragging the weighted total to 3.5 against a 3.75 threshold. The agent found the answer, but fed itself far more than it needed.
Core Techniques for Managing Context
Effective agent context management comes down to seven habits: filter data before it reaches the model, retrieve narrowly, load procedures only when needed, split work across specialists that share findings, store memory as facts, review and expire what’s stored, and measure what each tool call costs.
1. Filter before the model. Raw logs and events are too big and noisy to pass to an LLM directly. Cut them down first with traditional ML that filters, clusters and correlates them, so the LLM works on the curated result.
2. Retrieve narrowly. Narrow by metadata first (cluster, service, environment), then search within what’s left.
3. Load on demand. Give the agent an index of what’s available and let it open the full content when needed. Descriptions matter: in the docs’ words, “a vague description means a skill that never gets opened.”
4. Use specialists that share what they find. An orchestrator can dispatch Datadog, Kubernetes, AWS and Grafana specialists in parallel, each with a clean, narrow context. They hold together because every specialist reads from the same context layer, returns its findings rather than a transcript, and works to a step that declares what it needs and what it produces. On shared data, specialists coordinate rather than duplicate each other’s work.
5. Store facts, not transcripts. Distill each run into discrete, environment-specific facts, such as “Service A depends on Service B,” rather than replaying past investigations word for word.
6. Review and expire. Review what gets written before other agents can read it, check for contradictions, and retire facts that newer evidence has replaced.
7. Measure context cost per tool. Track token efficiency per tool call, and strip duplicate input before it reaches the model.
How Komodor Builds Context for Operations Agents
On the Komodor Agentic Operations Platform, context is a shared layer, not a per-agent configuration. Every agent, whether Komodor’s own or one a team builds or imports, reasons from the same knowledge graph and shares memory through spaces, so a reviewed finding from one investigation is available to the next.
The shared memory and context layer has four parts:
- Knowledge graph: a living map of services, owners, dependencies, deploys and data stores. It stays current from operational activity.
- Knowledge base: the runbooks, policies and tribal knowledge a team curates, searchable and cited.
- Agent memory: incident and remediation memory, recall across runs, and on-call handoff.
- Context layer: the shared context every agent inherits automatically.
Change history sits alongside these, because a rollback recommendation is only safe if the previous revision is known to be good and the difference is understood.
Memories live in shared spaces, because, as Komodor’s memory docs put it, “What one agent learned is invisible to another until they share a space.” A reviewer agent checks every write before any agent can read it, and a sweep every six hours catches contradictions and outdated memories. Approvals, outcomes and grades feed back in, so later runs build on what worked.
That’s what operations teams need from shared context: it stays current, nothing is trusted before it’s reviewed, and anything that’s no longer true gets retired. Every agent running on the platform inherits it from its first run. Back to that Monday: once the first agent’s finding passes review, the second agent starts from it.
Context Engineering FAQs
RAG is one technique within context engineering: it retrieves relevant documents and adds them to the prompt. Context engineering is the broader discipline. It also covers what’s always loaded, what’s remembered between runs, how findings pass between agents, and what’s deliberately left out.
Knowledge and memory should be shared, scoped by permission and reviewed before use, so one team’s hard-won finding isn’t rediscovered by another. Instructions usually stay specific to each agent. The goal is one reviewed source of operational truth that every agent reads from.
No. Chroma’s context-rot research found performance changes significantly with input length, even on simple tasks, and that a single distractor makes it worse. A larger window also adds cost and latency to every step. Bigger windows help with occasional large inputs, but they don’t replace deciding what belongs in context.
Accuracy improves because the model reasons over relevant, current evidence instead of noise. Scalability improves because shared memory and knowledge let many agents reuse what one has learned. Cost falls because smaller, targeted context means fewer tokens per step, and cheaper models can handle narrower jobs.
Memory is what an agent carries forward from past runs, like a recurring failure pattern or a fix that worked. A knowledge base holds documents a team curates on purpose, such as runbooks and postmortems. Memory is learned from operations, and knowledge is written by people.
See how the shared memory and context layer works across every agent on the Komodor Agentic Operations Platform.
