Amazon Web Services (AWS) has unveiled a new capability for building AI assistants that retain context across conversations, moving beyond the limitations of stateless interactions. This advancement leverages Amazon Bedrock AgentCore’s memory features in conjunction with the open-source agentic system OpenClaw, allowing AI assistants to remember past interactions, user preferences, and specific details, thereby offering a more personalized and efficient user experience.
The Chenab Times has learned that the new architecture addresses the common drawback of conventional AI assistants, where each query is treated as an isolated event without regard for previous exchanges. This can lead to users repeatedly providing the same information, diminishing the perceived intelligence and utility of the assistant. The integration with AgentCore memory transforms ephemeral chat sessions into persistent knowledge bases, enabling assistants to recall details like past gardening advice, fertilizer preferences, or plant health issues.
A Domain-Agnostic Architecture for Personalized AI
The system, demonstrated with a gardening assistant named Sprout, is designed to be adaptable to various domains. By modifying the persona and available skills, the same underlying architecture can power customer support bots, fitness coaches, or internal help desk tools. The entire system can be deployed using a single AWS CloudFormation template, operating on a consumption-based model that proponents suggest can cost as little as a few dollars per month for light personal use. Key to this is the AgentCore runtime, which provides a scalable platform for building and connecting agents with various frameworks and models.
Solution Overview and Technical Implementation
The system’s request flow begins with an inbound webhook from Telegram, which is routed through Amazon API Gateway and an AWS Lambda function. Scheduled jobs, such as reminder notifications, can also be initiated via Amazon EventBridge Scheduler and a cronjob Lambda function. Both entry points invoke the InvokeAgentRuntime API on the AgentCore runtime. This runtime coordinates with the OpenClaw gateway, AgentCore memory, and the Amazon Bedrock Converse API. Supporting AWS services include Amazon S3 for workspace storage, AWS Key Management Service for encryption, AWS Secrets Manager for bot tokens, and Amazon CloudWatch for logging and metrics.
The AgentCore runtime operates on a consumption-based pricing model, meaning users are billed only for the compute resources their agent actively uses, rather than for continuous uptime. This is a significant cost advantage over always-on instances for personal assistants used intermittently. The runtime enforces a minimal container contract, requiring agents to listen on port 8080 and expose GET /ping for health checks and POST /invocations as the primary agent entry point.
OpenClaw as the Agent Substrate and Multi-Model Routing
OpenClaw serves as the agent’s foundation, providing the agent loop, tool usage capabilities, and a skills system. A wrapper script, server.py, adapts OpenClaw to the AgentCore HTTP protocol. This wrapper handles tasks such as parsing payloads, retrieving memory, assembling context, forwarding requests to the OpenClaw gateway, and persisting results. The architecture also accommodates scenarios where AgentCore might need to thaw a frozen container, ensuring the gateway is ready before processing a turn.
To optimize for cost and performance, the system routes requests to different Amazon Bedrock models based on the task. Claude Haiku 4.5 is used for high-volume text conversations due to its speed and cost-effectiveness, while Claude Sonnet 4.5, with its stronger multimodal reasoning capabilities, is employed for tasks involving image understanding, such as diagnosing a plant from a photograph. Model IDs are configurable via environment variables, allowing for easy model swapping without rebuilding the agent image.
Memory: From Disposable Chats to Durable Knowledge
The core innovation lies in AgentCore memory, which enables context retention. This memory is structured into two layers: short-term memory, which stores each conversation turn as an event keyed by actor ID and session ID, and long-term memory, which is populated asynchronously through managed extraction strategies. These strategies identify and store different types of information, including explicit user preferences (e.g., organic fertilizer use), inferred facts (e.g., plant type and location), and session summaries.
Information is organized into per-user namespaces, ensuring that data from different users does not intermingle. For example, a namespace might be structured as sprout/{chat_id}/long_term. On each turn, the agent retrieves relevant long-term records, ranks them, and injects them into the system prompt. This retrieval process has a defined timeout, with the system designed to degrade gracefully and respond without memory if retrieval fails.
Metadata and Prompt Caching for Enhanced Efficiency
Structured metadata is crucial for effectively subgrouping memories within a namespace. Indexed keys such as ‘type’, ‘section’, and ‘plants’ allow for server-side filtering of memories, narrowing the scope of information before it is added to the prompt. This significantly improves retrieval accuracy and reduces the computational load.
To manage the cost implications of large system prompts, prompt caching on Amazon Bedrock is utilized. By structuring the prompt with stable elements like the persona and assembled memory at the beginning and the volatile user message at the end, Bedrock can cache the processed prefix across requests. This technique can dramatically reduce inference costs and latency.
Design Guidelines for Scalable AI Assistants
AWS provides several design guidelines for developers building on AgentCore and OpenClaw. These include wrapping agent frameworks rather than forking them, designing namespaces carefully for isolation, treating memory as an enhancement rather than a dependency, routing models appropriately based on task, ordering prompts for effective caching, planning for extraction latency, setting budgets from the outset, and keeping skills small and single-purpose to improve model selection and testability.
The full source code for this architecture is available on GitHub, enabling developers to deploy a pre-built solution or customize it for their specific needs. The cost for light personal use is estimated to be between $5–9 per month, factoring in infrastructure, text model, and vision model usage, with AWS Budgets alerts configured to monitor spending.
The Chenab Times News Desk

