This article was originally published on the Actian blog. Read the original here.
You'll build a GPT-6 Astra coding agent that saves every test-confirmed bug fix to Actian VectorAI DB and searches those fixes before it edits code. Because the fixes live outside the response chain, a brand-new session can retrieve a repair that Astra confirmed in an earlier one.
You need GPT-6 Astra API access, Docker, and Python 3.12. The complete implementation, tests, and demo projects are in the companion GitHub repository, so you can clone it and follow along.
Before you start
Check these requirements first.
- You need an OpenAI API key with GPT-6 Astra access. Astra isn't available on the free API tier, and API billing is separate from a ChatGPT subscription.
- You need a small API budget. Standard processing costs $10 per million input tokens and $50 per million output tokens. The transport check and two live sessions in this build cost about $0.15, and your run may cost more or less depending on token usage and request count.
- You need Docker with Compose to run VectorAI DB locally.
- You need Python 3.12 in a Linux shell. The original commands ran from the repository root in WSL2.
How the memory loop works
In the Responses API, previous_response_id and durable Conversation objects preserve context for a thread the application already knows it wants to continue. A fresh response chain can't see a fix that Astra confirmed during separate work. This build stores each confirmed fix in VectorAI DB so any later session can search for it.
The agent follows this order on every failure.
- Astra sends the error details to
search_codebase_memory. The tool creates a local embedding and searches VectorAI DB for similar confirmed fixes. - The agent loop returns the matches as a
function_call_outputlinked to the original request by itscall_id. Each match includes the error type, affected files, solution, and test outcome. - Astra compares any retrieved fix with the current code before it changes anything. If the search finds nothing relevant, Astra investigates using the project files and test output.
- Astra runs the acceptance tests after it implements a fix. A passing result lets it call
store_fix_memory, which saves the failure, solution, affected files, and verified outcome. - The handler rejects unconfirmed records, which keeps failed attempts out of later searches.
Memory retrieval always comes before editing, and memory storage always comes after passing tests.
Set up the project
Project layout
astra-codebase-memory/
├── docker-compose.yml
├── requirements.txt
├── settings.py
├── vectoraidb_memory_tools.py
├── astra_agent.py
├── coding_tools.py
├── cost_guard.py
├── main.py
├── smoke_test.py
├── session_demo.py
├── demo_projects/
└── tests/
vectoraidb_memory_tools.py holds the VectorAI DB backend, the local embedder, and the memory handlers. The Responses API loop lives in astra_agent.py. main.py wires these pieces together for a single coding session, and session_demo.py runs the controlled cross-session experiment.
Start VectorAI DB
docker pull actian/vectorai:latest
Pull the current VectorAI DB image, then start the tested configuration with the project's docker-compose.yml.
docker compose -p astra-codebase-memory up -d
docker compose -p astra-codebase-memory ps
The Compose file maps the REST endpoint to http://localhost:16573, which is the address the Python client uses.
Install the Python dependencies
python3.12 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python -m pip install --no-deps --no-build-isolation -e .
Run these commands from the repository root to create a virtual environment and install the pinned versions from requirements.txt.
Add your API key
cp .env.example .env
Copy the example environment file, then open .env and add your OpenAI API key.
OPENAI_API_KEY=your_api_key
The remaining values in .env.example configure the VectorAI DB connection, Standard API processing, and the project's spending limit.
Create the collection and load local embeddings
created = await self._call(
"collections_create",
self.collection_name,
{"vectors": {"size": EMBEDDING_DIMENSION, "distance": "Cosine"}},
timeout=30.0,
)
vectoraidb_memory_tools.py creates the astra_codebase_memory collection with 384-dimensional vectors and cosine distance. Each stored vector carries a payload with the fix description, error type, file path, confirmed outcome, timestamp, and session ID.
self._model = factory(self.model_name, device="cpu")
encoded = self._load_model().encode(
clean_text,
normalize_embeddings=True,
show_progress_bar=False,
convert_to_numpy=True,
)
The same file loads sentence-transformers/all-MiniLM-L6-v2 on the CPU and normalizes the generated vectors. The model downloads the first time it loads. Because it generates embeddings locally, it doesn't use OpenAI embedding credits.
Run the smoke test
.venv/bin/python -u smoke_test.py store
.venv/bin/python -u smoke_test.py verify --memory-id "PASTE_MEMORY_ID_FROM_STORE"
The first command creates or validates the collection, stores a temporary fix, and searches for it. Copy the returned memory_id into the second command, which confirms the record persisted and then deletes it.
VectorAI DB can now store and retrieve embedded fix records. Next, you'll give Astra a controlled way to use that database.
Build the memory tools
Both tools live in vectoraidb_memory_tools.py and share one MemoryService. The service receives the VectorAI DB backend, the local embedder, and the current session ID when the agent starts.
Search for confirmed fixes
async def search_codebase_memory(
self,
query: str,
*,
error_type: str | None = None,
top_k: int = 3,
) -> str:
clean_query = _required_text(query, "query")
_validate_search_arguments(top_k, error_type)
embedding = await self.embedder.embed(clean_query)
results = await self.backend.search(
embedding,
top_k=top_k,
error_type=error_type,
)
return format_search_results(results)
Astra calls search_codebase_memory before it changes any code. query holds the error message or failed-test details. The optional error_type narrows the search to one kind of failure, and top_k caps the number of matches. The method embeds the query locally and waits for VectorAI DB to finish the search before Astra continues.
Format matches for Astra
def format_search_results(
results: Sequence[MemorySearchResult],
) -> str:
"""Return compact readable evidence for a function-call output."""
if not results:
return "No relevant confirmed fixes found."
blocks = []
for index, result in enumerate(results, start=1):
record = result.record
blocks.append(
"\n".join(
(
f"Match {index} (score={result.score:.4f})",
f"Fix: {record.fix_description}",
f"Error type: {record.error_type}",
f"Files: {', '.join(record.file_paths)}",
f"Outcome: {record.outcome}",
f"Confirmed: {record.timestamp}",
f"Session: {record.session_id}",
)
)
)
return "\n\n".join(blocks)
format_search_results turns matching records into text that the Responses API can return as a function_call_output. Each match gives Astra the earlier fix plus the details it needs to judge relevance. VectorAI DB supplies the similarity score, which the formatter rounds to four decimal places. The file paths and test outcome add context about the stored fix.
Store a fix only after the tests pass
def confirm_outcome(self, outcome: str) -> None:
self._confirmed_outcomes.add(
_required_text(outcome, "outcome")
)
async def store_fix_memory(
self,
*,
fix_description: str,
error_type: str,
file_paths: Sequence[str],
outcome: str,
timestamp: datetime | None = None,
) -> str:
clean_fix = _required_text(
fix_description,
"fix_description",
)
clean_error = _required_text(error_type, "error_type")
clean_outcome = _required_text(outcome, "outcome")
if clean_outcome not in self._confirmed_outcomes:
raise UnconfirmedOutcomeError(
"outcome was not confirmed by the current run"
)
if isinstance(file_paths, (str, bytes)):
raise MemoryValidationError(
"file_paths must be a sequence of paths"
)
paths = tuple(file_paths)
confirmed_at = timestamp or datetime.now(UTC)
if (
confirmed_at.tzinfo is None
or confirmed_at.utcoffset() is None
):
raise MemoryValidationError(
"timestamp must include a timezone"
)
memory_id = str(
uuid5(
NAMESPACE_URL,
"\n".join(
(
self.session_id,
clean_fix,
clean_error,
clean_outcome,
*paths,
)
),
)
)
record = MemoryRecord(
memory_id=memory_id,
fix_description=clean_fix,
error_type=clean_error,
outcome=clean_outcome,
file_paths=paths,
timestamp=confirmed_at.isoformat(),
session_id=self.session_id,
embedding=await self.embedder.embed(clean_fix),
)
await self.backend.upsert(record)
return memory_id
After Astra makes a repair, the agent loop runs the project's tests. The loop calls confirm_outcome with the latest passing test result right before it stores the fix. store_fix_memory checks that value, builds an ID from the session and repair details, embeds the fix description locally, and writes the record to VectorAI DB.
The generated UUID becomes the record's ID in VectorAI DB. If the same session retries an identical storage request, it produces the same UUID and updates the existing record, so the collection doesn't fill with duplicates.
Register the tools with the Responses API
MEMORY_TOOL_DEFINITIONS = (
{
"type": "function",
"name": "search_codebase_memory",
"description": (
"Search confirmed past debugging results "
"using observed failure symptoms."
),
"async": True,
"strict": True,
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string"},
"error_type": {
"type": ["string", "null"]
},
"top_k": {
"type": "integer",
"minimum": 1,
"maximum": 10,
},
},
"required": [
"query",
"error_type",
"top_k",
],
"additionalProperties": False,
},
},
{
"type": "function",
"name": "store_fix_memory",
"description": (
"Store a debugging result after its outcome "
"is confirmed by the current run."
),
"strict": True,
"parameters": {
"type": "object",
"properties": {
"fix_description": {
"type": "string"
},
"error_type": {"type": "string"},
"file_paths": {
"type": "array",
"items": {"type": "string"},
},
"outcome": {"type": "string"},
},
"required": [
"fix_description",
"error_type",
"file_paths",
"outcome",
],
"additionalProperties": False,
},
},
)
These definitions describe each function and the arguments Astra must supply. Setting strict to True keeps every call inside its declared schema. Astra must provide each required argument, and "additionalProperties": False rejects any field outside the definition.
The search tool uses Astra's async tool-calling option, which lets the model continue with independent work while the application runs the lookup. It's marked asynchronous because it waits on the local embedding and the VectorAI DB search.
Build the agent loop
The loop in astra_agent.py controls when Astra can search memory, inspect the project, edit a file, and save a confirmed fix.
Set the system instructions
SYSTEM_INSTRUCTIONS = """You are fixing a bug in a confined demonstration workspace.
Call search_codebase_memory using only observed failure symptoms before stating a diagnosis or
editing code. Wait for its function output. An empty result still completes the required search.
Use only the supplied coding tools. Store a fix only after run_tests returns exit code 0, and use
the exact confirmed_outcome string returned by that test call. Do not assume memory will help."""
These instructions set the required order at the start of every session. The loop then enforces the same order with its own checks.
Run the loop
async def run(
self,
prompt: str,
*,
session_id: str | None = None,
) -> AgentRunResult:
if not isinstance(prompt, str) or not prompt.strip():
raise ValueError("prompt must be a non-empty string")
run_session_id = session_id or str(uuid4())
previous_response_id: str | None = None
next_input: object = [
{"role": "user", "content": prompt.strip()}
]
memory_result_delivered = False
pending_memory_delivery = False
last_confirmed_outcome: str | None = None
start_event_index = len(self.event_logger.events)
run_model_requests = 0
run_local_tools = 0
while True:
if run_model_requests >= self.limits.max_turns:
raise AgentLimitError(
"model turn limit reached"
)
payload: dict[str, Any] = {
"model": "gpt-6-astra",
"instructions": SYSTEM_INSTRUCTIONS,
"input": next_input,
"tools": self.tool_definitions,
"reasoning": {"effort": "low"},
"text": {"verbosity": "low"},
"max_output_tokens": (
self.limits.max_output_tokens
),
}
if previous_response_id is not None:
payload["previous_response_id"] = (
previous_response_id
)
if pending_memory_delivery:
memory_result_delivered = True
pending_memory_delivery = False
response = await self._request(
payload,
run_session_id,
)
run_model_requests += 1
response_id, output = self._validate_response(
response
)
function_calls = [
item
for item in output
if item.get("type") == "function_call"
]
if not function_calls:
if not memory_result_delivered:
raise MemoryGateError(
"agent produced a diagnosis before "
"receiving memory output"
)
final_text = self._extract_text(output)
return AgentRunResult(
session_id=run_session_id,
final_text=final_text,
final_response_id=response_id,
model_request_count=run_model_requests,
local_tool_call_count=run_local_tools,
events=tuple(
self.event_logger.events[
start_event_index:
]
),
)
outputs: list[dict[str, str]] = []
memory_ready_at_response_start = (
memory_result_delivered
)
search_completed = False
for item in function_calls:
if (
run_local_tools
>= self.limits.max_tool_calls
):
raise AgentLimitError(
"local tool-call limit reached"
)
(
tool_output,
confirmed_outcome,
was_search,
) = await self._dispatch(
item,
session_id=run_session_id,
response_id=response_id,
memory_ready=(
memory_ready_at_response_start
),
last_confirmed_outcome=(
last_confirmed_outcome
),
)
run_local_tools += 1
if item.get("name") == "run_tests":
last_confirmed_outcome = (
confirmed_outcome
)
if item.get("name") == "apply_edit":
last_confirmed_outcome = None
search_completed = (
search_completed or was_search
)
outputs.append(
{
"type": "function_call_output",
"call_id": str(item["call_id"]),
"output": tool_output,
}
)
if search_completed:
pending_memory_delivery = True
previous_response_id = response_id
next_input = outputs
Every call to run starts with previous_response_id set to None, so the first API request opens a new response chain. After Astra responds, the loop keeps the returned response ID and attaches it to the next request in the same session.
Each tool result carries the call_id from Astra's original function call, which lets the next response match the returned value to the correct request. The memory gate records when the search output has reached Astra. If Astra tries to finish before that happens, the loop raises MemoryGateError. last_confirmed_outcome keeps the latest passing test result available for store_fix_memory, and any later apply_edit call clears it.
Gate storage inside _dispatch
elif name == "store_fix_memory":
self._require_keys(
arguments,
{
"fix_description",
"error_type",
"file_paths",
"outcome",
},
)
if not memory_ready:
raise MemoryGateError(
"memory storage attempted before "
"memory result delivery"
)
if (
last_confirmed_outcome is None
or arguments["outcome"]
!= last_confirmed_outcome
):
raise ConfirmedFixGateError(
"store_fix_memory requires the latest "
"passing-test confirmed_outcome"
)
self.memory_service.confirm_outcome(
last_confirmed_outcome
)
result = {
"memory_id": (
await self.memory_service.store_fix_memory(
**arguments
)
)
}
was_search = False
confirmed_outcome = last_confirmed_outcome
_dispatch runs the final storage checks. It compares the outcome Astra supplies with the latest passing test result before it confirms and saves the record. A storage attempt before the memory output arrives raises MemoryGateError, and a mismatched outcome raises ConfirmedFixGateError.
Send requests to the Responses API
response = await asyncio.to_thread(
requests.post,
RESPONSES_URL,
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
},
json=payload,
timeout=self.timeout_seconds,
)
response.raise_for_status()
result = response.json()
The project sends each payload with requests.post. That call sits inside RealHTTPResponsesTransport, which also checks that live mode, the API key, and the project budget are configured before any request leaves.
Run one coding session
backend = ActianVectorAIBackend(
settings.vectorai_url,
collection_name=settings.vectorai_collection,
grpc_url=settings.vectorai_grpc_url,
)
await backend.ensure_collection()
memory_service = MemoryService(
backend,
SentenceTransformerEmbedder(
settings.embedding_model
),
session_id=run_id,
)
transport = RecordedLiveTransport(
RealHTTPResponsesTransport(
guard,
enabled=True,
api_key=settings.api_key,
processing_tier=settings.processing_tier,
),
ledger=ledger,
guard=guard,
session_id=run_id,
transcript_path=(
evidence_dir
/ "private"
/ f"{run_id}-transcript.jsonl"
),
)
agent = AstraAgent(
transport,
memory_service,
SafeCodingTools(workspace),
limits=AgentLimits(
max_turns=LIVE_MAX_TURNS,
max_output_tokens=LIVE_MAX_OUTPUT_TOKENS,
),
)
result = await agent.run(
prompt,
session_id=run_id,
)
main.py loads the settings, checks the spending limit, and then connects the VectorAI DB backend, memory service, coding tools, and Responses API transport. The complete file in the repository also validates the workspace, API key, processing tier, run ID, and live-cost approval before the session starts.
Scenario Alpha contains a deliberate bug for the agent to investigate. Copy it into a working folder so Astra can edit files without touching the original version.
mkdir -p runs/tutorial
cp -R demo_projects/templates/scenario_alpha runs/tutorial/scenario_alpha
Run the agent against the copied project.
.venv/bin/python main.py \
--workspace runs/tutorial/scenario_alpha \
--run-id tutorial-alpha \
--approve-live-cost
The --approve-live-cost flag confirms that the session may send paid Responses API requests. When the run ends, main.py prints Astra's response, the API request and tool-call counts, and the session cost.
Test memory across separate response chains
for scenario, suffix in (
("scenario_alpha", "alpha"),
("scenario_beta", "beta"),
):
session_id = f"{run_id}-session-{suffix}"
workspace = reset_workspace(run_id, scenario)
coding_tools = SafeCodingTools(workspace)
transport = RecordedLiveTransport(
transport_factory(),
ledger=ledger,
guard=guard,
session_id=session_id,
transcript_path=(
evidence_dir
/ "private"
/ f"{run_id}-{suffix}-transcript.jsonl"
),
)
service = TrackedMemoryService(
backend,
embedder,
session_id=session_id,
tracker=tracker,
)
agent = AstraAgent(
transport,
service,
coding_tools,
limits=AgentLimits(
max_turns=LIVE_MAX_TURNS,
max_output_tokens=LIVE_MAX_OUTPUT_TOKENS,
),
event_logger=EventLogger(
evidence_dir
/ "private"
/ f"{run_id}-{suffix}-events.jsonl"
),
)
result = await agent.run(
DEMO_PROMPT,
session_id=session_id,
)
new_chains = all(
bool(transport.requests)
and "previous_response_id"
not in transport.requests[0]
for transport in transports
)
session_demo.py runs two projects against the same VectorAI DB collection. It creates a new memory service and agent for each project, and the full function also handles validation and cleanup. new_chains checks that no session's first request included previous_response_id.
The first session had no stored fixes to draw from. Astra traced the failed origin check to whitespace around the comma-separated values, updated app/service.py, and saved the result after all four tests passed.
Fixed `app/service.py` to trim whitespace around comma-separated allowed origins. The leading space caused the second origin to be rejected.
All 4 tests pass. Stored the confirmed fix result.
The second session switched to a different project and started a fresh response chain. Its memory search found the record from the first run.
Match 1 (score=0.1650)
Fix: Trim surrounding whitespace from each comma-separated APP_ALLOWED_ORIGINS value in configured_values so the second configured origin matches exactly and receives Access-Control-Allow-Origin.
Error type: AssertionError
Files: app/service.py
Outcome: pytest passed with exit code 0
Confirmed: 2026-09-15T12:11:51.934136+00:00
Session: live-session1-20260915-1207-session-alpha
The transcript places this retrieval before any relevant file reads, diagnosis, or code change. Astra later fixed a separate normalization bug in app/service.py, converting values such as Audit-Log to audit_log. Because the live run also changed an assertion in tests/test_service.py, the application fix was verified once more in a fresh workspace with the original, untouched tests.
.... [100%]
4 passed in 0.17s
All four original acceptance tests passed in the fresh workspace, which confirms the Session 2 application fix. The transcript also verifies that a confirmed fix from Session 1 reached Astra in a separate response chain before diagnosis and editing began.
Results across five sessions
The same Scenario Beta task ran across five live Astra sessions. Each session used a fresh workspace and a separate response chain, so its first request didn't include previous_response_id. Session 1 began with an empty VectorAI DB collection and stored its confirmed fix after the tests passed. Sessions 2 through 5 started with that record as their only available memory, and records created by those later sessions were removed before the next run.
| Session | Session 1 memory retrieved | Responses API requests | Local tool calls | Untouched tests | Cost |
|---|---|---|---|---|---|
| 1 | No | 7 | 9 | 4 passed | $0.0779135 |
| 2 | Yes | 7 | 9 | 4 passed | $0.0758070 |
| 3 | Yes | 7 | 9 | 4 passed | $0.0702995 |
| 4 | Yes | 7 | 9 | 4 passed | $0.0718085 |
| 5 | Yes | 7 | 9 | 4 passed | $0.0704115 |
Each session made seven Responses API requests and nine local tool calls. Astra edited app/service.py and tests/test_service.py, so each application fix was checked in a fresh workspace with the original tests. All four tests passed every time.
Sessions 2 through 5 retrieved the tested Session 1 fix despite starting new response chains. This pattern helps most when a failure resembles one the agent has solved before. New errors and broader design decisions still require investigation of the current project.
Wrapping up
The five-session run showed that a confirmed fix can move from one Astra response chain to another through VectorAI DB. Sessions 2 through 5 retrieved the records saved during Session 1, and each session passed the four original acceptance tests.
VectorAI DB Community Edition gives you a local database to try the same approach with your own coding tasks. Pull the image to get started.
docker pull actian/vectorai:latest
From there, connect the search and storage tools from this tutorial to Astra so it can save successful fixes and find them again in later sessions.
What's next
- Clone the companion repository and run the two-session workflow yourself.
- Read What are Self-Evolving AI Agents? on the Actian blog.
- Download VectorAI DB Community Edition and try the memory tools on your own codebase.
FAQ
Does GPT-6 Astra remember previous sessions?
It doesn't remember them automatically. previous_response_id or a Conversation object can preserve earlier context, but an independent response chain needs an external memory tool to retrieve fixes from other sessions.
How do I add persistent memory to a GPT-6 Astra agent?
Store confirmed fixes outside the response chain and expose tools for searching and adding records. In this tutorial, Astra searches before editing and stores a fix only after the tests pass.
Can I use an external vector database with GPT-6 Astra?
Yes, you can. Define Responses API function tools that search and update the database, then return the results to Astra as tool output.
What's the difference between OpenAI's file_search tool and an external vector database?
OpenAI's file_search is a hosted tool for searching uploaded files. An external vector database gives your application direct control over storage, embeddings, metadata, filtering, updates, and deletion.
Top comments (0)