Context Engineering for AI Agents: RAG, Memory, Tools, and Production Architecture
A practical guide to context engineering for production AI agents: RAG, memory, live tools, workflow state, security, evaluation, and cost control.
By BoundLayer Engineering Team
BoundLayer is a senior engineering partner for SaaS, fintech, AI automation, cloud infrastructure, legacy modernization, Web3, IoT, GPU computing, and data systems.
When an AI agent gives a poor answer, teams often try to fix the prompt or switch to a larger model.
In production systems, the model is frequently not the main problem. The agent may have received stale documents, irrelevant conversation history, missing permissions, an incomplete customer record, ambiguous tool descriptions, or too much context to identify what matters.
Context engineering is the discipline of giving an AI system the right information, instructions, state, and capabilities at the right moment.
It includes prompts, but it is much broader than prompt engineering. A production context layer may combine retrieval-augmented generation, structured application state, short-term session history, long-term memory, tool results, user identity, business rules, and evidence from live systems.
At BoundLayer, we design context as production infrastructure. It must be relevant, secure, traceable, cost-efficient, and measurable. This guide explains the architecture and operating practices required to build context-aware AI agents that work reliably with real business data.
What Context Engineering Actually Means
A model produces an output from the information available in its current context window. Context engineering controls how that information is selected, structured, ordered, updated, and protected.
A useful context package may contain:
- system instructions and policies;
- the current user request;
- authenticated user and tenant information;
- structured workflow state;
- recent conversation turns;
- retrieved documents and citations;
- long-term preferences or historical facts;
- results from databases and APIs;
- available tools and their contracts;
- output schemas and validation requirements;
- examples relevant to the current task.
The goal is not to maximize how much information the model sees. The goal is to provide the smallest sufficient set of trusted information for the decision being made.
Prompt Engineering vs. Context Engineering
Prompt engineering focuses on how instructions are written.
Context engineering owns the complete information environment around the model.
| Prompt engineering | Context engineering |
|---|---|
| Wording and structure of instructions | Full runtime information architecture |
| Often tested with static examples | Changes with users, data, tools, and workflow state |
| Primarily model-facing | Connects models to application and data systems |
| Can improve one interaction | Must remain reliable across sessions and releases |
| Usually text-oriented | Includes structured state, retrieval, identity, tools, and memory |
A strong prompt cannot compensate for the wrong customer record. It cannot make a stale policy current. It cannot enforce tenant isolation or recover state after a failed workflow.
Those are system design problems.
A Reference Architecture
A production context layer should sit between business systems and the model.
User or business event
|
Identity, tenant, and authorization context
|
Workflow state and task definition
|
Context orchestrator
| | | |
RAG Memory Live tools Policies
| | | |
Context selection, ranking, redaction, and budgeting
|
Model or agent call
|
Validated output and tool decisions
|
Trace, evaluation, feedback, and state update
The context orchestrator is not necessarily one service. It is a set of responsibilities that may be implemented across the application backend, retrieval service, workflow engine, memory store, and agent framework.
Keeping those responsibilities explicit makes the system easier to secure and evaluate.
The Six Context Layers
1. Instructions and Policy
System instructions define the agent's role, task boundaries, communication style, and required behavior.
They should be stable, versioned, and separated from untrusted user or document content. Critical authorization and business rules must still be enforced outside the prompt.
2. Structured Runtime State
Runtime state describes what is happening now:
- workflow step;
- case ID;
- completed and pending actions;
- current status;
- deadlines;
- verified facts;
- approval state;
- retry count;
- output from previous deterministic steps.
This state belongs in a database or durable workflow system, not only inside conversation text.
3. Retrieved Knowledge
Retrieval brings relevant information from documents, databases, tickets, product catalogs, policies, and other knowledge sources into the current request.
This is the domain of RAG, search, knowledge graphs, and data APIs.
4. Short-Term Session Context
Short-term context includes recent messages and temporary task state. It helps the model resolve references and maintain continuity within one interaction.
Do not resend an unlimited transcript. Summarize or extract structured state when conversations become long.
5. Long-Term Memory
Long-term memory stores information that should survive across sessions: user preferences, previous outcomes, recurring constraints, and durable facts learned from interactions.
Memory needs its own lifecycle and governance. It is not simply an archive of every message.
6. Tools and Live Data
Tools let the agent query current systems or take actions. They are the best source for facts that change frequently, such as inventory, account balance, deployment status, or delivery tracking.
Our guide to MCP security in production explains how to expose these capabilities with identity, authorization, isolation, and audit controls.
RAG Is Not Memory
Retrieval-augmented generation and agent memory solve related but different problems.
RAG usually retrieves authoritative external knowledge:
- product documentation;
- policies;
- contracts;
- technical runbooks;
- historical tickets;
- research and knowledge-base articles.
Memory captures continuity from the agent's interactions:
- the user's preferred report format;
- a decision made in an earlier session;
- a recurring exception for one account;
- lessons from a completed task;
- a summary of an ongoing case.
A vector database does not automatically turn stored text into good memory. The system must decide what is worth remembering, which entity it belongs to, how long it remains valid, and when it should be updated or deleted.
Mixing authoritative policy with inferred personal memory in one untyped index creates provenance and trust problems.
Build a Retrieval Pipeline, Not Just a Vector Search
Naive RAG often follows a simple pattern: split documents, embed the chunks, retrieve the nearest vectors, and put them into the prompt.
That can work for a prototype. Production retrieval usually needs more stages.
Ingestion
Connect approved sources and preserve document identity, ownership, timestamps, access controls, and version history.
Parsing and Structure
Extract headings, tables, sections, links, and document relationships. Arbitrary fixed-size chunks can separate a rule from its exception or a table header from its values.
Metadata
Attach fields such as tenant, product, region, language, document type, effective date, status, and security classification.
Indexing
Use the retrieval method that matches the data. Semantic vectors are useful, but keyword search, relational queries, graph traversal, and exact identifiers may be better for specific questions.
Query Understanding
Rewrite or decompose the request, resolve entities, and apply filters derived from trusted identity and workflow context.
Retrieval and Reranking
Retrieve a candidate set and rerank it against the actual question. Hybrid semantic and lexical retrieval often handles names, codes, and domain terminology better than one method alone.
Context Assembly
Select the smallest evidence set that answers the question. Include source identity and dates, and remove duplicate or conflicting passages.
Citation and Verification
Require the final response or decision to reference the evidence used. When sources disagree, expose the conflict instead of silently choosing one.
Use Live Tools for Volatile Facts
Not every source belongs in a retrieval index.
Frequently changing operational data should usually come from a live API or database query:
- current subscription status;
- account balance;
- available inventory;
- active incidents;
- latest deployment;
- shipment location;
- today's price or exchange rate.
Indexing this information creates a freshness problem. The agent may retrieve an old snapshot and act as though it were current.
A practical rule is:
Stable knowledge -> retrieval index
Current operational fact -> live tool
Exact business state -> system of record
Cross-session preference -> governed memory
The source should match the lifetime and authority of the information.
Design Memory as a Data Product
Long-term memory needs a clear data model.
Useful memory categories include:
Semantic Memory
Durable facts about a user, account, or domain, such as preferred units or organization terminology.
Episodic Memory
Summaries of specific past interactions or completed tasks.
Procedural Memory
Learned instructions or workflow patterns. These require careful governance because allowing an agent to rewrite its own operating procedure can introduce hidden behavior changes.
Preference Memory
User-specific choices such as tone, output format, notification channel, or reporting cadence.
For every memory, record:
- subject and tenant;
- source interaction;
- memory type;
- creation and update time;
- confidence or verification status;
- expiration policy;
- access scope;
- deletion and correction path.
Users and operators should be able to inspect and correct important memories. An inferred preference should not silently become permanent truth.
Avoid the Long-Context Trap
Larger context windows are useful, but they do not eliminate context engineering.
Sending every available document and the full conversation can create several problems:
- higher inference cost;
- slower responses;
- important evidence buried in noise;
- contradictory facts;
- larger prompt-injection surface;
- weaker reproducibility;
- more sensitive data sent to the provider.
Context capacity is not the same as context quality.
Use the larger window for tasks that genuinely need broad comparison or synthesis. For repeated operational decisions, filtered and structured context is usually more reliable and economical.
Give Every Context Item Provenance
The agent should know where information came from and how much authority it has.
Useful provenance fields include:
- source system;
- document or record ID;
- author or owner;
- creation and effective dates;
- retrieval timestamp;
- tenant and access scope;
- authoritative, inferred, or user-supplied status;
- version;
- confidence where applicable.
Provenance lets the application prefer an active policy over an archived document, a live CRM record over conversation memory, and a verified fact over a model inference.
It also supports citations, debugging, compliance, and deletion requests.
Enforce Access Before Retrieval
Security filtering after retrieval is too late. Sensitive content may already have entered model context, logs, traces, or caches.
Authorization should constrain the search itself.
Apply user and tenant permissions to:
- source connectors;
- metadata filters;
- database queries;
- vector namespaces;
- graph traversals;
- cache keys;
- memory retrieval;
- tool calls;
- trace visibility.
Never ask the model to remove documents the user should not see. The model should not receive them.
For a customer-support agent, the authenticated employee's role, assigned accounts, region, and case should determine which records can enter context.
Treat Retrieved Content as Untrusted
Documents, email, websites, tickets, and tool responses may contain prompt injection or malicious instructions.
Context engineering must preserve the distinction between instructions and data.
Use controls such as:
- clear content boundaries;
- source allowlists;
- document sanitation;
- restricted tool availability by workflow step;
- outbound destination controls;
- deterministic authorization;
- human approval for high-impact actions;
- evaluation with adversarial documents.
A highly relevant malicious document is still malicious.
Budget Context Deliberately
Every request has a latency and token budget. Allocate it according to task value.
A context budget may reserve capacity for:
- system and policy instructions;
- current task state;
- recent conversation;
- retrieved evidence;
- tool results;
- output generation.
When the budget is exceeded, do not simply truncate the end. Use prioritization, deduplication, structured compression, and task-specific summaries.
Preserve exact data where precision matters, such as identifiers, monetary values, dates, and quoted policy language. Summaries are appropriate for narrative history but can lose critical exceptions.
Track the contribution of each context source to quality. A large memory layer that rarely changes outcomes is operational cost without proven value.
Example: Customer Support Agent
A support agent needs more than a knowledge-base search.
The context package may contain:
- authenticated agent and customer identity;
- current ticket and recent messages;
- account, subscription, and entitlement state from live APIs;
- relevant product documentation;
- known incidents;
- previous case summaries;
- approved response and escalation policies;
- available ticketing actions.
Each source has different authority and freshness.
The knowledge base explains how a feature works. The billing API confirms the current plan. Memory may record that the customer prefers email follow-up. The ticket system remains the source of truth for case status.
Combining everything into one vector store would erase those distinctions.
Example: Financial Operations Agent
A reconciliation or payment-operations agent may need:
- structured workflow state;
- ledger and transaction records;
- payment-provider status;
- customer correspondence;
- compliance policy;
- prior exception decisions;
- approval limits;
- current user authority.
The model may interpret correspondence and classify an exception. It must not replace the ledger or calculate authoritative balances from retrieved prose.
Deterministic services perform accounting and authorization. The context layer supplies relevant evidence for the bounded judgment.
Our fintech architecture guide explains how ledgers and durable workflows protect financial correctness.
Example: Engineering and Incident Agent
An engineering agent investigating an incident may retrieve:
- the active alert;
- recent deployments;
- service ownership;
- relevant runbooks;
- correlated logs, metrics, and traces;
- similar historical incidents;
- infrastructure state;
- approved remediation tools.
Static runbooks can come from RAG. Current metrics and deployment state should come from live observability APIs. Historical incident summaries can provide episodic memory, but the system should preserve links to original evidence.
The agent can propose a cause and next step. Production changes should remain behind policy and approval until the workflow has earned a higher level of autonomy.
Evaluate Context Separately From Model Output
End-to-end task success matters, but teams also need to know which context layer failed.
Retrieval Metrics
- recall of required evidence;
- precision of retrieved passages;
- ranking quality;
- metadata-filter correctness;
- freshness;
- citation coverage.
Memory Metrics
- correct entity association;
- useful recall rate;
- stale or contradictory memory rate;
- compression quality;
- user correction frequency;
- privacy and deletion compliance.
Context Assembly Metrics
- token usage by source;
- duplicate content;
- conflicting evidence detection;
- context-to-answer attribution;
- latency by stage.
Outcome Metrics
- answer correctness and groundedness;
- task completion;
- human acceptance or override;
- unsafe-action rate;
- cost per successful outcome;
- operational time saved.
Run evaluations when documents, chunking, embeddings, retrieval logic, memory extraction, models, prompts, or tool contracts change.
This belongs in the wider AgentOps lifecycle, with versioned artifacts and production feedback.
Observe the Context Path
Every agent trace should answer:
- which user and tenant initiated the task;
- which context sources were queried;
- what filters and query rewrites were applied;
- which documents or memories were selected;
- why they were ranked;
- what was redacted or excluded;
- which tools returned live data;
- how many tokens each source consumed;
- which evidence supported the final result;
- which context and model versions were used.
Do not log raw sensitive context by default. Store secure references and sanitized metadata where possible.
Observability should make failures diagnosable without creating a second uncontrolled copy of company data.
A Practical Delivery Plan
1. Define the Decision or Workflow
Start with the business outcome, users, risk, and definition of success.
2. Inventory Context Sources
List systems, owners, data classifications, freshness requirements, and existing permissions.
3. Assign Source Authority
Decide which system is authoritative for each fact. Separate knowledge, live operational data, workflow state, and memory.
4. Build a Thin End-to-End Context Path
Support one workflow with real identity, retrieval, citations, and a controlled tool set.
5. Create an Evaluation Dataset
Include normal cases, difficult exceptions, stale documents, conflicting sources, missing evidence, and unauthorized content.
6. Add Memory Only Where It Helps
Begin with explicit, inspectable memory types. Measure whether they improve outcomes before expanding capture.
7. Deploy With Human Review
Collect corrections and overrides. Use them to improve source quality, retrieval, and policies.
8. Automate Proven Segments
Increase autonomy by action type only when context quality and outcome metrics support it.
For complex environments, our forward deployed engineering approach keeps senior engineers close to the business workflow, data owners, and production systems throughout this process.
Build or Buy the Context Layer?
Cloud platforms now offer managed knowledge bases, memory services, agent runtimes, and retrieval systems. They can reduce infrastructure work.
Custom components remain useful when the organization needs:
- specialized retrieval and ranking;
- existing search infrastructure;
- strict data residency;
- complex authorization;
- domain-specific knowledge graphs;
- provider portability;
- unusual latency or scale requirements;
- full control over memory behavior.
The decision is rarely entirely managed or entirely custom. A team may use a managed vector service while owning ingestion, metadata, authorization, evaluation, and context assembly.
The durable asset is the organization's context model and evaluation data, not a particular framework.
How BoundLayer Can Help
We design and build context-aware AI systems, including:
- enterprise RAG and hybrid search;
- document ingestion and re-indexing pipelines;
- metadata and access-control architecture;
- short-term state and long-term memory;
- knowledge graphs and entity resolution;
- secure API and MCP tool integrations;
- durable agent workflows;
- evaluation datasets and automated quality checks;
- context observability and cost controls;
- AWS and cloud infrastructure for production deployment.
Our work connects the AI layer with the existing business systems that hold authoritative data. The result is not just a chatbot with more documents. It is software that can retrieve the right evidence, understand the current task, preserve controlled continuity, and take bounded actions.
For the broader product architecture, see our guide to building practical AI agent software.
Current Technical References
- Google Cloud: A developer's guide to production-ready AI agents;
- Google Cloud: Preparing data infrastructure for AI agents;
- AWS: Context intelligence for data and AI agents;
- AWS: AgentCore memory for context-aware agents;
- A Survey of Context Engineering for Large Language Models.
The Bottom Line
Production AI quality depends on more than the model and prompt.
Context engineering determines which instructions, facts, memories, tools, and workflow state reach the model, under which permissions, at what cost, and with what provenance.
The strongest architecture does not put everything into one vector store or one enormous prompt. It separates authoritative data from memory, stable knowledge from live state, instructions from untrusted content, and semantic judgment from deterministic business rules.
When context is treated as governed production infrastructure, AI agents become easier to secure, evaluate, debug, and improve.
That is the difference between a convincing demonstration and an AI system that can participate reliably in real business operations.
Related engineering articles
MCP Security in Production: How to Connect Enterprise AI Agents Without Losing Control
A practical guide to securing production MCP servers with OAuth, tool-level authorization, prompt-injection defenses, isolation, approvals, and auditability.
AI Agents and AI Software Development: How We Build Practical Systems for Real Business Workflows
How BoundLayer builds AI agents, RAG systems, workflow automation, data assistants, and custom AI software that connects to real business systems.
Jev and Typed Decision Models: When AI Software Should Decide Instead of Generate Text
A practical guide to Jev from TypeSafe AI: typed Choice, Score, and Noul decisions, production architecture, evaluation, limitations, and business use cases.
Building an AI agent that needs reliable business context?
We design and build secure RAG, memory, tool integration, workflow state, evaluation, and cloud infrastructure for production AI systems.