Build a self-improving GPT-6 Astra coding agent with codebase memory
This article was originally published on the Actian blog. Read the original here.
GPT-6 Astra keeps context inside a response chain. In the Responses API, previous_response_id links each response to the one before it, and durable Conversation objects carry a thread across sessions, devices, or jobs. Both mechanisms assume the application already knows which earlier interaction belongs to the current one.
A coding agent that opens a new response chain for each task has no link back to its earlier work. A fix Astra confirmed in one session sits in a chain that the next session can't see, even when the next failure looks the same.
This guide adds a memory layer that lives outside the response chain. Confirmed fixes go into Actian VectorAI DB with local embeddings and structured metadata, and Astra searches that store before it diagnoses or edits code. By the end, you'll have implemented each of these pieces.
Create a VectorAI DB collection backed by local 384-dimensional embeddings.
Define two Responses API function tools,
search_codebase_memoryandstore_fix_memory.Build an agent loop that gates retrieval ahead of editing and storage behind passing tests.
Run a two-session experiment that shows a confirmed fix crossing response chains.
The complete implementation, tests, and demo projects are in the companion GitHub repository.
Architecture overview
Why the response chain needs an external store
History becomes useful when an agent needs a detail that didn't make it into the working summary. BlackwellBoy gives a practical example.
"If a weird test failure happened three hours ago and didn't make the summary notes, Astra can dig back through the logs to retrieve it."
That example finds an earlier event inside the logs of one long-running task. This build targets a second need, where a later session finds a relevant fix from separate work. Each confirmed fix goes into VectorAI DB so a new response chain can search for it.
How the components fit together
Every failure moves through the same sequence.
When a test fails, Astra sends the error details to
search_codebase_memory. The tool creates a local embedding and searches VectorAI DB for similar confirmed fixes. Each result contains the error type, affected files, solution, and test outcome.The agent loop returns the search result to Astra as a
function_call_outputlinked to the original request through itscall_id.Astra compares any retrieved fix with the current code before it makes changes. If the search finds nothing relevant, Astra continues the investigation using the project files and test output.
After implementing a fix, Astra runs the acceptance tests. A passing result allows it to call
store_fix_memory, which saves the failure, solution, affected files, and verified outcome.The handler rejects unconfirmed records, which keeps failed attempts out of later searches.
Three design decisions shape the rest of the build.
Retrieval comes before editing. The loop refuses to accept a final answer until the memory search output has reached Astra.
Storage sits behind passing tests. The handler accepts only the outcome string that the current run's latest passing test produced.
Embeddings run locally. The embedder runs on the CPU, so searches and stores don't spend OpenAI embedding credits. Structured metadata makes each fix searchable by similarity, error type, affected files, and outcome.
Since VectorAI DB stores the records outside the response chain, later sessions can retrieve fixes saved during earlier work. The stack needs Astra API access, a local VectorAI DB instance, and a local embedding model.
Prerequisites and cost
GPT-6 Astra is unavailable on the free API tier, and API billing is separate from a ChatGPT subscription. Standard processing currently costs $10 per million input tokens and $50 per million output tokens. In this experiment, the transport check and two live sessions cost about $0.15. A new run may cost more or less depending on its token usage and how many Responses API requests the agent needs.
You'll also need Docker with Compose and Python 3.12. The commands below ran from the repository root in WSL2.
Set up the stack
Project layout
The project separates the memory layer, the Astra agent loop, the local coding tools, and the demo workflow into their own files.
astra-codebase-memory/
├── docker-compose.yml
├── requirements.txt
├── settings.py
├── vectoraidb_memory_tools.py
├── astra_agent.py
├── coding_tools.py
├── cost_guard.py
├── main.py
├── smoke_test.py
├── session_demo.py
├── demo_projects/
└── tests/
vectoraidb_memory_tools.py contains the VectorAI DB backend, local embedder, and memory handlers. The Responses API loop lives in astra_agent.py. The main.py entry point connects these components for a single coding session, while session_demo.py runs the controlled cross-session experiment.
Start VectorAI DB
Pull the current VectorAI DB image.
docker pull actian/vectorai:latest
The project starts the tested configuration through docker-compose.yml.
docker compose -p astra-codebase-memory up -d
docker compose -p astra-codebase-memory ps
The Compose file maps the REST endpoint to http://localhost:16573, which is the address the Python client uses.
Install dependencies and configure the environment
The pinned package versions are stored in requirements.txt. Create the environment and install them from the repository root.
python3.12 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python -m pip install --no-deps --no-build-isolation -e .
Copy the example environment file.
cp .env.example .env
Open .env and add your OpenAI API key.
OPENAI_API_KEY=your_api_key
The remaining values in .env.example configure the VectorAI DB connection, Standard API processing, and the spending limit the project enforces.
Create the collection
# Create the astra_codebase_memory collection in VectorAI DB.
# EMBEDDING_DIMENSION is 384 to match the local embedding model's output,
# and cosine distance pairs with the normalized vectors the embedder produces.
created = await self._call(
"collections_create",
self.collection_name,
{"vectors": {"size": EMBEDDING_DIMENSION, "distance": "Cosine"}},
timeout=30.0,
)
Each stored vector carries a payload containing the fix description, error type, file path, confirmed outcome, timestamp, and session ID. Keeping that metadata next to the vector means a search result arrives with everything Astra needs to judge it.
Load the local embedder
# Load sentence-transformers/all-MiniLM-L6-v2 on the CPU.
# normalize_embeddings=True keeps every vector unit-length for cosine search.
self._model = factory(self.model_name, device="cpu")
encoded = self._load_model().encode(
clean_text,
normalize_embeddings=True,
show_progress_bar=False,
convert_to_numpy=True,
)
The model downloads the first time it loads. Since it generates these embeddings locally, it doesn't use OpenAI embedding credits.
Verify the stack with the smoke test
.venv/bin/python -u smoke_test.py store
.venv/bin/python -u smoke_test.py verify --memory-id "PASTE_MEMORY_ID_FROM_STORE"
The first command creates or validates the collection, stores a temporary fix, and searches for it. Copy the returned memory_id into the second command, which confirms that the record persisted and deletes it afterward.
At this point, VectorAI DB can store and retrieve embedded fix records. Astra still needs a controlled way to use that database, which is the job of the two memory tools.
Build the memory tools
Both memory tools live in vectoraidb_memory_tools.py and use the same MemoryService. The service receives the VectorAI DB backend, local embedder, and current session ID when the agent starts.
The search tool
# Search confirmed fixes that resemble the current failure.
# query holds the observed error text, error_type optionally narrows the
# search to one failure class, and top_k caps the number of matches.
async def search_codebase_memory(
self,
query: str,
*,
error_type: str | None = None,
top_k: int = 3,
) -> str:
clean_query = _required_text(query, "query")
_validate_search_arguments(top_k, error_type)
embedding = await self.embedder.embed(clean_query)
results = await self.backend.search(
embedding,
top_k=top_k,
error_type=error_type,
)
return format_search_results(results)
When Astra encounters a failure, it calls search_codebase_memory before changing any code. The method converts the query into a local embedding and waits for VectorAI DB to complete the search before Astra continues.
Why this approach
Searching by embedding lets Astra surface a stored fix even when the new error message is worded differently from the original one. The optional error_type filter can restrict results to a single class of failure when the symptoms are clear. Running the embedder locally means every search and every store uses the same model with no per-call API charge. Because the payload stores file paths and test outcomes beside each vector, Astra can weigh a match against the current code before it acts on it.
Formatting results for Astra
# Turn VectorAI DB matches into plain text for a function_call_output.
# Each block carries the score, fix, error type, files, outcome, and session
# so Astra can judge relevance without another lookup.
def format_search_results(
results: Sequence[MemorySearchResult],
) -> str:
"""Return compact readable evidence for a function-call output."""
if not results:
return "No relevant confirmed fixes found."
blocks = []
for index, result in enumerate(results, start=1):
record = result.record
blocks.append(
"\n".join(
(
f"Match {index} (score={result.score:.4f})",
f"Fix: {record.fix_description}",
f"Error type: {record.error_type}",
f"Files: {', '.join(record.file_paths)}",
f"Outcome: {record.outcome}",
f"Confirmed: {record.timestamp}",
f"Session: {record.session_id}",
)
)
)
return "\n\n".join(blocks)
Each match gives Astra the earlier fix and the details it needs to judge relevance to the current problem. The similarity score comes from VectorAI DB and is rounded to four decimal places in the formatted result.
The store tool
# confirm_outcome registers a passing test result for the current run.
# store_fix_memory refuses any outcome that wasn't registered, and its
# deterministic uuid5 ID turns a repeated store into an update.
def confirm_outcome(self, outcome: str) -> None:
self._confirmed_outcomes.add(
_required_text(outcome, "outcome")
)
async def store_fix_memory(
self,
*,
fix_description: str,
error_type: str,
file_paths: Sequence[str],
outcome: str,
timestamp: datetime | None = None,
) -> str:
clean_fix = _required_text(
fix_description,
"fix_description",
)
clean_error = _required_text(error_type, "error_type")
clean_outcome = _required_text(outcome, "outcome")
if clean_outcome not in self._confirmed_outcomes:
raise UnconfirmedOutcomeError(
"outcome was not confirmed by the current run"
)
if isinstance(file_paths, (str, bytes)):
raise MemoryValidationError(
"file_paths must be a sequence of paths"
)
paths = tuple(file_paths)
confirmed_at = timestamp or datetime.now(UTC)
if (
confirmed_at.tzinfo is None
or confirmed_at.utcoffset() is None
):
raise MemoryValidationError(
"timestamp must include a timezone"
)
memory_id = str(
uuid5(
NAMESPACE_URL,
"\n".join(
(
self.session_id,
clean_fix,
clean_error,
clean_outcome,
*paths,
)
),
)
)
record = MemoryRecord(
memory_id=memory_id,
fix_description=clean_fix,
error_type=clean_error,
outcome=clean_outcome,
file_paths=paths,
timestamp=confirmed_at.isoformat(),
session_id=self.session_id,
embedding=await self.embedder.embed(clean_fix),
)
await self.backend.upsert(record)
return memory_id
After Astra makes a repair, the agent loop runs the project's tests and calls confirm_outcome with the latest passing result immediately before storing the fix. store_fix_memory checks that value, creates an ID from the session and repair details, embeds the fix description locally, and writes the completed record to VectorAI DB.
The generated UUID becomes the record's ID in VectorAI DB. If the same session retries an identical storage request, it generates the same UUID and updates the existing record, which keeps the collection free of duplicates.
Declaring the tools
# Responses API function definitions for both memory tools.
# strict mode holds every call to its schema. The search tool is async so
# Astra can keep working while the embedding and lookup complete.
MEMORY_TOOL_DEFINITIONS = (
{
"type": "function",
"name": "search_codebase_memory",
"description": (
"Search confirmed past debugging results "
"using observed failure symptoms."
),
"async": True,
"strict": True,
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string"},
"error_type": {
"type": ["string", "null"]
},
"top_k": {
"type": "integer",
"minimum": 1,
"maximum": 10,
},
},
"required": [
"query",
"error_type",
"top_k",
],
"additionalProperties": False,
},
},
{
"type": "function",
"name": "store_fix_memory",
"description": (
"Store a debugging result after its outcome "
"is confirmed by the current run."
),
"strict": True,
"parameters": {
"type": "object",
"properties": {
"fix_description": {
"type": "string"
},
"error_type": {"type": "string"},
"file_paths": {
"type": "array",
"items": {"type": "string"},
},
"outcome": {"type": "string"},
},
"required": [
"fix_description",
"error_type",
"file_paths",
"outcome",
],
"additionalProperties": False,
},
},
)
Astra can call these methods once they're declared as Responses API function tools. The search definition uses Astra's async tool-calling option, which allows the model to continue with independent work while the application runs the memory search.
Setting strict to True keeps each function call within its declared schema. Astra must provide every required argument, and "additionalProperties": False rejects fields outside the definition. With the memory functions registered, the agent loop can pass Astra's calls to MemoryService and return each result to the correct response.
Build the agent loop
The loop in astra_agent.py controls when Astra can search the memory collection, inspect the project, edit a file, and save a confirmed fix.
System instructions
# The required order, stated to the model at the start of every session.
SYSTEM_INSTRUCTIONS = """You are fixing a bug in a confined demonstration workspace.
Call search_codebase_memory using only observed failure symptoms before stating a diagnosis or
editing code. Wait for its function output. An empty result still completes the required search.
Use only the supplied coding tools. Store a fix only after run_tests returns exit code 0, and use
the exact confirmed_outcome string returned by that test call. Do not assume memory will help."""
The application reinforces those instructions with checks inside the loop, so the order holds even if the model drifts from them.
The run loop
# Each call to run opens a new response chain because previous_response_id
# starts as None. The loop enforces two gates. Astra can't finish before the
# memory output reaches it, and storage needs the latest passing test result.
async def run(
self,
prompt: str,
*,
session_id: str | None = None,
) -> AgentRunResult:
if not isinstance(prompt, str) or not prompt.strip():
raise ValueError("prompt must be a non-empty string")
run_session_id = session_id or str(uuid4())
previous_response_id: str | None = None
next_input: object = [
{"role": "user", "content": prompt.strip()}
]
memory_result_delivered = False
pending_memory_delivery = False
last_confirmed_outcome: str | None = None
start_event_index = len(self.event_logger.events)
run_model_requests = 0
run_local_tools = 0
while True:
if run_model_requests >= self.limits.max_turns:
raise AgentLimitError(
"model turn limit reached"
)
payload: dict[str, Any] = {
"model": "gpt-6-astra",
"instructions": SYSTEM_INSTRUCTIONS,
"input": next_input,
"tools": self.tool_definitions,
"reasoning": {"effort": "low"},
"text": {"verbosity": "low"},
"max_output_tokens": (
self.limits.max_output_tokens
),
}
if previous_response_id is not None:
payload["previous_response_id"] = (
previous_response_id
)
if pending_memory_delivery:
memory_result_delivered = True
pending_memory_delivery = False
response = await self._request(
payload,
run_session_id,
)
run_model_requests += 1
response_id, output = self._validate_response(
response
)
function_calls = [
item
for item in output
if item.get("type") == "function_call"
]
if not function_calls:
if not memory_result_delivered:
raise MemoryGateError(
"agent produced a diagnosis before "
"receiving memory output"
)
final_text = self._extract_text(output)
return AgentRunResult(
session_id=run_session_id,
final_text=final_text,
final_response_id=response_id,
model_request_count=run_model_requests,
local_tool_call_count=run_local_tools,
events=tuple(
self.event_logger.events[
start_event_index:
]
),
)
outputs: list[dict[str, str]] = []
memory_ready_at_response_start = (
memory_result_delivered
)
search_completed = False
for item in function_calls:
if (
run_local_tools
>= self.limits.max_tool_calls
):
raise AgentLimitError(
"local tool-call limit reached"
)
(
tool_output,
confirmed_outcome,
was_search,
) = await self._dispatch(
item,
session_id=run_session_id,
response_id=response_id,
memory_ready=(
memory_ready_at_response_start
),
last_confirmed_outcome=(
last_confirmed_outcome
),
)
run_local_tools += 1
if item.get("name") == "run_tests":
last_confirmed_outcome = (
confirmed_outcome
)
if item.get("name") == "apply_edit":
last_confirmed_outcome = None
search_completed = (
search_completed or was_search
)
outputs.append(
{
"type": "function_call_output",
"call_id": str(item["call_id"]),
"output": tool_output,
}
)
if search_completed:
pending_memory_delivery = True
previous_response_id = response_id
next_input = outputs
Once Astra responds, the loop keeps the returned response ID and attaches it to the next request in the same session. Each tool result carries the call_id from Astra's original function call, which allows the next response to associate the returned value with the correct request.
The memory gate records when the search output has reached Astra. A final answer before that point raises MemoryGateError. last_confirmed_outcome keeps the latest passing test result available for store_fix_memory, and an apply_edit call resets it to None so an edit after the tests always needs a fresh passing run.
Storage checks in _dispatch
# Final gate before a record reaches VectorAI DB. Storage needs the memory
# output delivered and an outcome that matches the latest passing test.
elif name == "store_fix_memory":
self._require_keys(
arguments,
{
"fix_description",
"error_type",
"file_paths",
"outcome",
},
)
if not memory_ready:
raise MemoryGateError(
"memory storage attempted before "
"memory result delivery"
)
if (
last_confirmed_outcome is None
or arguments["outcome"]
!= last_confirmed_outcome
):
raise ConfirmedFixGateError(
"store_fix_memory requires the latest "
"passing-test confirmed_outcome"
)
self.memory_service.confirm_outcome(
last_confirmed_outcome
)
result = {
"memory_id": (
await self.memory_service.store_fix_memory(
**arguments
)
)
}
was_search = False
confirmed_outcome = last_confirmed_outcome
The _dispatch method compares the outcome supplied by Astra with the latest passing test result before confirming and saving the record. A mismatch raises ConfirmedFixGateError, so Astra can't store a fix the tests haven't confirmed.
The HTTP transport
# Post each payload to the Responses API from a worker thread.
response = await asyncio.to_thread(
requests.post,
RESPONSES_URL,
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
},
json=payload,
timeout=self.timeout_seconds,
)
response.raise_for_status()
result = response.json()
The HTTP call sits inside RealHTTPResponsesTransport, which also checks that live mode, the API key, and the project budget have been configured before a request leaves.
Run the agent
A single coding session
# Wire the VectorAI DB backend, memory service, coding tools, and
# Responses API transport, then run one session.
backend = ActianVectorAIBackend(
settings.vectorai_url,
collection_name=settings.vectorai_collection,
grpc_url=settings.vectorai_grpc_url,
)
await backend.ensure_collection()
memory_service = MemoryService(
backend,
SentenceTransformerEmbedder(
settings.embedding_model
),
session_id=run_id,
)
transport = RecordedLiveTransport(
RealHTTPResponsesTransport(
guard,
enabled=True,
api_key=settings.api_key,
processing_tier=settings.processing_tier,
),
ledger=ledger,
guard=guard,
session_id=run_id,
transcript_path=(
evidence_dir
/ "private"
/ f"{run_id}-transcript.jsonl"
),
)
agent = AstraAgent(
transport,
memory_service,
SafeCodingTools(workspace),
limits=AgentLimits(
max_turns=LIVE_MAX_TURNS,
max_output_tokens=LIVE_MAX_OUTPUT_TOKENS,
),
)
result = await agent.run(
prompt,
session_id=run_id,
)
main.py loads the settings and checks the spending limit before it connects these components. The complete file in the repository also validates the workspace, API key, processing tier, run ID, and live-cost approval before starting the session.
Scenario Alpha contains the deliberate bug that the agent will investigate. Copy it into a working folder so Astra can edit the files without changing the original version.
mkdir -p runs/tutorial
cp -R demo_projects/templates/scenario_alpha runs/tutorial/scenario_alpha
Run the agent against the copied project.
.venv/bin/python main.py \
--workspace runs/tutorial/scenario_alpha \
--run-id tutorial-alpha \
--approve-live-cost
The --approve-live-cost flag confirms that the session may send paid Responses API requests. When the run ends, main.py prints Astra's response, the API request and tool-call counts, and the session cost.
Memory across separate response chains
# Run two scenarios against one VectorAI DB collection. Each gets a new
# memory service, agent, and response chain, and new_chains confirms that
# no first request carried previous_response_id.
for scenario, suffix in (
("scenario_alpha", "alpha"),
("scenario_beta", "beta"),
):
session_id = f"{run_id}-session-{suffix}"
workspace = reset_workspace(run_id, scenario)
coding_tools = SafeCodingTools(workspace)
transport = RecordedLiveTransport(
transport_factory(),
ledger=ledger,
guard=guard,
session_id=session_id,
transcript_path=(
evidence_dir
/ "private"
/ f"{run_id}-{suffix}-transcript.jsonl"
),
)
service = TrackedMemoryService(
backend,
embedder,
session_id=session_id,
tracker=tracker,
)
agent = AstraAgent(
transport,
service,
coding_tools,
limits=AgentLimits(
max_turns=LIVE_MAX_TURNS,
max_output_tokens=LIVE_MAX_OUTPUT_TOKENS,
),
event_logger=EventLogger(
evidence_dir
/ "private"
/ f"{run_id}-{suffix}-events.jsonl"
),
)
result = await agent.run(
DEMO_PROMPT,
session_id=session_id,
)
new_chains = all(
bool(transport.requests)
and "previous_response_id"
not in transport.requests[0]
for transport in transports
)
session_demo.py runs two projects against the same collection, and the full function also handles validation and cleanup.
The first session had no stored fixes to draw from. Astra traced the failed origin check to whitespace around the comma-separated values, updated app/service.py, and saved the result after all four tests passed.
Fixed `app/service.py` to trim whitespace around comma-separated allowed origins. The leading space caused the second origin to be rejected.
All 4 tests pass. Stored the confirmed fix result.
For the second session, the run switched to a different project and started a fresh response chain. Its memory search found the record from the first run.
Match 1 (score=0.1650)
Fix: Trim surrounding whitespace from each comma-separated APP_ALLOWED_ORIGINS value in configured_values so the second configured origin matches exactly and receives Access-Control-Allow-Origin.
Error type: AssertionError
Files: app/service.py
Outcome: pytest passed with exit code 0
Confirmed: 2026-09-15T12:11:51.934136+00:00
Session: live-session1-20260915-1207-session-alpha
The transcript places the retrieval before any relevant file reads, diagnosis, or code change. Astra later fixed a separate normalization bug in app/service.py, converting values such as Audit-Log to audit_log. Because the live run also changed an assertion in tests/test_service.py, the application fix was verified once more in a fresh workspace containing the original, untouched tests.
.... [100%]
4 passed in 0.17s
Passing all four original acceptance tests in the fresh workspace confirms the Session 2 application fix. The transcript also verifies that a confirmed fix from Session 1 reached Astra in a separate response chain before diagnosis and editing began.
What the five-session run showed
Every tool call was recorded to see whether memory changed how much work later sessions did. The same Scenario Beta task ran across five live Astra sessions. Each session used a fresh workspace and a separate response chain, so its first request didn't include previous_response_id.
Session 1 began with an empty VectorAI DB collection and stored its confirmed fix after the tests passed. Sessions 2 through 5 started with that record as their only available memory. Records created by those later sessions were removed before the next run.
| Session | Session 1 memory retrieved | Responses API requests | Local tool calls | Untouched tests | Cost |
|---|---|---|---|---|---|
| 1 | No | 7 | 9 | 4 passed | $0.0779135 |
| 2 | Yes | 7 | 9 | 4 passed | $0.0758070 |
| 3 | Yes | 7 | 9 | 4 passed | $0.0702995 |
| 4 | Yes | 7 | 9 | 4 passed | $0.0718085 |
| 5 | Yes | 7 | 9 | 4 passed | $0.0704115 |
Each session made seven Responses API requests and nine local tool calls. Astra edited app/service.py and tests/test_service.py, so each application fix was checked in a fresh workspace containing the original tests. All four tests passed every time.
Sessions 2 through 5 retrieved the tested Session 1 fix despite starting new response chains. This is most useful when a failure resembles one the agent has solved before. New errors and broader design decisions still require investigation of the current project.
Conclusion
The five-session run showed that a confirmed fix could move from one Astra response chain to another through VectorAI DB. Sessions 2 through 5 retrieved the records saved during Session 1, and each session passed the four original acceptance tests.
VectorAI DB Community Edition gives you a local database to try the same approach with your own coding tasks. Pull the image to start.
docker pull actian/vectorai:latest
From there, you can connect the search and storage tools from this tutorial to Astra so it can save successful fixes and find them again in later sessions.
Where to take this
Clone the companion repository and reproduce the two-session workflow against your own VectorAI DB instance.
Read What are Self-Evolving AI Agents? on the Actian blog.
Start with VectorAI DB Community Edition and connect the memory tools to your own coding tasks.
Frequently asked questions
Does GPT-6 Astra remember previous sessions?
It doesn't remember them automatically. previous_response_id or a Conversation object can preserve earlier context, but an independent response chain needs an external memory tool to retrieve fixes from other sessions.
How do I add persistent memory to a GPT-6 Astra agent?
Store confirmed fixes outside the response chain and expose tools for searching and adding records. In this tutorial, Astra searches before editing and stores a fix only after the tests pass.
Can I use an external vector database with GPT-6 Astra?
Yes, you can. Define Responses API function tools that search and update the database, then return the results to Astra as tool output.
What's the difference between OpenAI's file_search tool and an external vector database?
OpenAI's file_search is a hosted tool for searching uploaded files. An external vector database gives your application direct control over storage, embeddings, metadata, filtering, updates, and deletion.



