DEV Community

Cover image for Power Up Your Agents: Execution Hooks and Smart Memory in Agent Kernel
AK for Agent Kernel

Posted on Originally published at kernel.yaala.ai

Power Up Your Agents: Execution Hooks and Smart Memory in Agent Kernel

By Yaala Labs

Ever wished you could intercept your AI agent's thoughts before they speak? Or give them a photographic memory that lasts exactly as long as you need? We've got you covered.

Agent Kernel now features Execution Hooks and Smart Memory Management - two game-changing capabilities that give you surgical control over how your agents think, remember, and respond. Whether you're building enterprise chatbots, RAG-powered assistants, or multi-agent systems, these features unlock new levels of sophistication without the complexity.

The Challenge: Control Meets Flexibility

Building production AI agents isn't just about picking the right LLM. You need:

  • Safety mechanisms that prevent harmful outputs before they escape
  • Context injection to make your agents smarter with domain knowledge
  • Memory that adapts to different use cases - ephemeral for some data, persistent for others
  • Audit trails that show exactly what your agent saw and said

Traditional approaches force you to hack this together with prompt engineering, custom middleware, or framework-specific workarounds. Agent Kernel takes a different path: first-class support for hooks and memory that works across any framework.

Execution Hooks: Your Agent's Control Panel

Think of hooks as strategic checkpoints in your agent's execution pipeline. You get two powerful interception points:

Pre-Execution Hooks: Shape the Input

These run before your agent sees the user's prompt. Perfect for:

๐Ÿ›ก๏ธ Guard Rails - Block inappropriate content before it reaches your agent:

class GuardRailHook(PreHook):
    BLOCKED_KEYWORDS = ["hack", "illegal", "exploit"]

    async def on_run(self, session, agent, requests):
        prompt = requests[0].prompt.lower()

        for keyword in self.BLOCKED_KEYWORDS:
            if keyword in prompt:
                return AgentReplyText(
                    text=f"I cannot assist with requests related to '{keyword}'."
                )

        return requests  # Safe - proceed to agent
Enter fullscreen mode Exit fullscreen mode

๐Ÿง  RAG Context Injection - Enrich prompts with knowledge from your databases:

class RAGHook(PreHook):
    async def on_run(self, session, agent, requests):
        prompt = requests[0].prompt

        # Search your knowledge base
        context = await search_knowledge_base(prompt)

        # Inject context into the prompt
        enriched_prompt = f"""Context: {context}

Question: {prompt}"""

        return [AgentRequestText(prompt=enriched_prompt)]
Enter fullscreen mode Exit fullscreen mode

๐Ÿ’ก Pro Tip: Custom Data in REST API Mode

When using Agent Kernel's REST API, you can pass custom data in your JSON request body beyond the standard agent, session_id, and prompt fields. Any additional keys are automatically converted to AgentRequestAny objects and passed to your pre-hooks:

# REST API request with custom fields
curl -X POST http://localhost:8000/api/v1/chat \
  -H "Content-Type: application/json" \
  -d '{
    "agent": "assistant",
    "session_id": "user-123",
    "prompt": "What is my account status?",
    "user_id": "12345",
    "additional_context": {
      "account_type": "premium",
      "region": "us-west"
    }
  }'
Enter fullscreen mode Exit fullscreen mode

In your pre-hook, access these custom fields:

class CustomDataHook(PreHook):
    async def on_run(self, session, agent, requests):
        # Find custom data in requests
        user_id = None
        additional_context = None

        for req in requests:
            if isinstance(req, AgentRequestAny):
                if req.name == "user_id":
                    user_id = req.content
                elif req.name == "additional_context":
                    additional_context = req.content

        # Use the data for personalization, auth, etc.
        if user_id and additional_context:
            # Store in cache for tools to access
            session.get_non_volatile_cache().set("user_id", user_id)
            session.get_non_volatile_cache().set("context", additional_context)

        return requests
Enter fullscreen mode Exit fullscreen mode

This is powerful for passing metadata like user IDs, session context, feature flags, or any custom data your hooks need without cluttering the actual prompt sent to the LLM. The custom fields are processed by your hooks but never sent to the agent unless you explicitly add them.

Post-Execution Hooks: Polish the Output

These run after your agent generates a response. Ideal for:

โš–๏ธ Adding Disclaimers - Compliance made automatic:

class DisclaimerHook(PostHook):
    async def on_run(self, session, requests, agent, agent_reply):
        disclaimer = "\n\n*This is AI-generated. Verify important decisions.*"
        agent_reply.response += disclaimer
        return agent_reply
Enter fullscreen mode Exit fullscreen mode

๐Ÿ”’ Output Moderation - Filter sensitive information from responses

๐Ÿ“Š Analytics - Log interactions for monitoring and improvement

Hook Chaining: The Power of Composition

Multiple hooks execute in sequence, each building on the last:

from agentkernel.openai import OpenAIModule
from agents import Agent

# Create agent
agent = Agent(name="assistant", instructions="You are a helpful assistant.")

# Register hooks in order using method chaining
OpenAIModule([agent]).pre_hook(agent, [
    RAGHook(),        # Add context first
    GuardRailHook(),  # Then validate everything
]).post_hook(agent, [
    ModerationHook(),   # Filter sensitive content
    DisclaimerHook(),   # Add legal disclaimer
])
Enter fullscreen mode Exit fullscreen mode

Flow: User Input โ†’ RAG โ†’ Guard Rails โ†’ Agent โ†’ Moderation โ†’ Disclaimer โ†’ Final Response

If any pre-hook returns an AgentReply (like guardrails blocking a request), execution stops immediately - the agent never sees the prompt.

Smart Memory: The Right Persistence at the Right Time

Not all data needs to live forever. Agent Kernel gives you two types of cache with identical APIs but different lifecycles:

Volatile Cache: Ephemeral Context

Data that lives only during a single request execution and vanishes afterward:

# In a pre-hook - inject context into cache
cache = session.get_volatile_cache()
cache.set("rag_context", retrieved_documents)

# In a tool - retrieve the context
cache = Session.current().get_volatile_cache()
docs = cache.get("rag_context")
return query_documents(docs, question)
Enter fullscreen mode Exit fullscreen mode

Perfect for:

  • ๐Ÿ“„ Document content from file uploads (don't clutter prompts)
  • ๐Ÿ” RAG search results (fresh every request)
  • ๐Ÿ”ข Temporary calculations and intermediate data
  • ๐Ÿงช Request-scoped feature flags

Non-Volatile Cache: Persistent Memory

Data that persists across multiple requests in the same session:

# First request - store user preferences
cache = session.get_non_volatile_cache()
cache.set("user_language", "Spanish")
cache.set("notification_enabled", True)

# Later requests - retrieve preferences
language = cache.get("user_language")  # Still "Spanish"
Enter fullscreen mode Exit fullscreen mode

Perfect for:

  • ๐Ÿ‘ค User preferences and settings
  • ๐Ÿ“ Extracted metadata from conversations
  • ๐ŸŽฏ Session-specific configurations
  • ๐Ÿท๏ธ Tags and classifications

Why This Matters: Clean Prompts, Lower Costs

Instead of stuffing everything into the prompt:

Before (bloated prompt):

prompt = f"""
User preferences: {json.dumps(prefs)}
Document content: {huge_document}
Previous context: {conversation_history}

Question: {user_question}
"""
Enter fullscreen mode Exit fullscreen mode

After (clean and efficient):

# Store in appropriate cache
volatile.set("document", huge_document)
non_volatile.set("preferences", prefs)

# Simple prompt
prompt = user_question  # Agent tools fetch from cache as needed
Enter fullscreen mode Exit fullscreen mode

Benefits:

  • โœ‚๏ธ Reduced token usage - Only send what the LLM needs to see
  • โšก Faster responses - Less to process per request
  • ๐Ÿ’ฐ Lower costs - Fewer tokens = smaller bills
  • ๐Ÿงน Cleaner prompts - Focus on the actual question

Real-World Example: RAG with Cache

Check out our key-value-cache example:

# Senior assistant with RAG hook
@function_tool
def query_knowledge_base(query: str) -> str:
    cache = Session.current().get_volatile_cache()
    context = cache.get("rag_context")  # Retrieved by pre-hook

    if context:
        return search_in_context(query, context)
    return "No information found."

# Register the hook with the Module
OpenAIModule([senior_assistant]).pre_hook(senior_assistant, [RAGHook()])
Enter fullscreen mode Exit fullscreen mode

The RAGHook searches your knowledge base and populates the cache. The tool accesses it transparently. The prompt stays clean.

Framework-Agnostic Magic

Here's the best part: this works with any agent framework Agent Kernel supports:

  • OpenAI Agents SDK
  • LangGraph
  • CrewAI
  • Google ADK

Same hook code. Same cache API. Different frameworks. One unified experience.

# Works with OpenAI
from agentkernel.openai import OpenAIModule
OpenAIModule([my_openai_agent]).pre_hook(my_openai_agent, [RAGHook()])

# Works with CrewAI
from agentkernel.crewai import CrewAIModule
CrewAIModule([my_crew_agent]).pre_hook(my_crew_agent, [RAGHook()])

# Same hooks and cache for both!
cache = session.get_volatile_cache()
Enter fullscreen mode Exit fullscreen mode

Production-Ready Architecture

Behind the scenes, Agent Kernel's memory system supports multiple backends:

Development:

export AK_SESSION__TYPE=in_memory
Enter fullscreen mode Exit fullscreen mode

Production (Redis):

export AK_SESSION__TYPE=redis
export AK_SESSION__REDIS__URL=redis://your-redis-instance
Enter fullscreen mode Exit fullscreen mode

Production (DynamoDB):

export AK_SESSION__TYPE=dynamodb
export AK_SESSION__DYNAMODB__TABLE_NAME=agent-sessions
Enter fullscreen mode Exit fullscreen mode

Same code, different backends. Swap them with environment variables.

Getting Started in Minutes

1. Install Agent Kernel

pip install agentkernel
Enter fullscreen mode Exit fullscreen mode

2. Create a Hook

from agentkernel import PreHook
from agentkernel import AgentRequestText

class SimpleRAGHook(PreHook):
    async def on_run(self, session, agent, requests):
        # Add context from your knowledge base
        context = get_relevant_context(requests[0].prompt)
        enriched = f"Context: {context}\n\nQuestion: {requests[0].prompt}"
        return [AgentRequestText(prompt=enriched)]

    def name(self):
        return "SimpleRAGHook"
Enter fullscreen mode Exit fullscreen mode

3. Register and Run

from agentkernel.openai import OpenAIModule
from agents import Agent

agent = Agent(
    name="assistant",
    instructions="You are a helpful assistant.",
)

# Register agent and hooks using method chaining
OpenAIModule([agent]).pre_hook(agent, [SimpleRAGHook()])

# Use CLI, REST API, or deploy to AWS
from agentkernel.api import RESTAPI
RESTAPI.run()
Enter fullscreen mode Exit fullscreen mode

4. Test It

curl -X POST http://localhost:8000/api/v1/chat \
  -H "Content-Type: application/json" \
  -d '{
    "agent": "assistant",
    "session_id": "test-123",
    "prompt": "What is Agent Kernel?"
  }'
Enter fullscreen mode Exit fullscreen mode

Your hook enriches the prompt before the agent sees it. Magic! โœจ

Learn More

Examples:

Documentation:

Why This Matters

Agent Kernel's hooks and memory system give you:

โœ… Separation of Concerns - Business logic separate from agent framework code

โœ… Reusability - Write hooks once, use across all agents and frameworks

โœ… Testability - Unit test hooks independently from agents

โœ… Composability - Chain hooks to build complex behaviors

โœ… Performance - Cache reduces token usage and costs

โœ… Flexibility - Swap memory backends without code changes

This is the kind of control you need for production AI agents that are safe, smart, and cost-effective.

What's Next?

We're not stopping here. Coming soon:

  • ๐Ÿ” RBAC Hooks - Role-based access control out of the box
  • ๐Ÿค– Human-in-the-Loop - Pause execution for human approval
  • ๐ŸŒŠ Streaming Hooks - Real-time interception for streaming responses
  • ๐Ÿ“ฆ Hook Marketplace - Share and discover community hooks

Ready to level up your agents? Get started with Agent Kernel today and join our Discord community to share what you build!


Built with โค๏ธ by Yaala Labs


Originally published at kernel.yaala.ai on December 18, 2025.

Top comments (0)