I want to walk a reconstructed click, not a metric from a customer I do not have. The spinner completes, the panel shows a summary, and a refresh after the worker recycles comes back blank. The call was already spent, and the only paragraph lived in a temp file on the worker. Does a green toast mean you own the result, or does it only mean the process was still breathing?
I think we keep mistaking a cheap seat for a system of record, and that mistake gets expensive the moment a host recycles. A free model allotment is useful fuel for one vertical slice while you learn the seams. A free server is a fair place to run that worker, but it is not memory you are allowed to trust. If the host can forget the bytes, the handler must not tell the user that the answer is saved.
The tempting story is that a working button on a free host means the feature is shipped, and I do not buy that story. Vibe-coded glue can paint the panel, call a model, and even sound confident while the only durable artifact is a screenshot. A screenshot is not a receipt, and a host you can lose overnight is not a place to keep customer text. If the prototype cannot survive a kill and a select, it is still a prototype, however fast the first demo felt.
The break is usually a boring handoff, not a dramatic failure buried inside the model itself. The UI posts a ticket id, the API checks the session, and the worker calls the model with text it already stored. The model returns a completion, the worker drops it into a temp file, and only then does the handler try to answer the browser. Which of those steps is the commit, and which are just hallway gossip between processes that can die?
I would rather return a hard error than let that temp file pretend to be the commit. The analogy I keep using is a coat check in a crowded hallway after a show. The model hands you a coat, and the free host is the hallway that gets swept when the lights come up. You do not tell the guest the coat is safe until the stub is written in a book you actually keep.
I am placing the disclosure here, because this is the first point where the product actually enters the method. Disclosure: This article was prepared as part of MonkeyCode's product outreach. I am using its free model access and its free server option as the cheap worker seat, not as the database I trust.
The outreach brief also cites a free allotment on the order of ten million tokens, which I have not re-measured for this draft. Confirm that figure on the current project page before you plan a sprint around it, because quotas move and old screenshots lie. I will not treat an unverified allotment as a capacity plan, and I will not omit the claim the brief asked me to surface.
My position is simple enough to argue with, which is how I like a production rule. The free host may execute the call, but your own store must commit the receipt before any success status leaves the API. If the commit fails, the client should see a retryable error, not a toast that points at memory the next restart will erase. Would you ship a payment capture that lived only in the web process, and why is a model completion any less of a write?
What follows is a proposed slice, not a benchmark I ran against a live free host this week. I am labeling the SQL and the handler as unexecuted examples, so you can adapt them without inheriting fake timings. The receipt row is the stub in the coat-check book, and the response body is copied from that row after the commit. If you cannot point at the row, you do not get to say the feature worked.
The table is deliberately dull, because dull is what you want when a free host disappears under you. Status is a small state machine, not a boolean the UI can paint green whenever the socket closes. The unique constraint is what stops a browser refresh from looking like a brand-new user intent.
create table completion_receipts (
id uuid primary key,
actor_id uuid not null,
ticket_id uuid not null,
attempt_key text not null,
status text not null check (status in ('reserved', 'stored', 'failed')),
output_text text,
error_code text,
created_at timestamptz not null default now(),
unique (actor_id, ticket_id, attempt_key)
);
A reserved row means we intended to spend a call, and a stored row means the text now belongs to us. A failed row means we owe the user an honest error instead of a second silent spend. The unique key is the part people skip, and then they wonder why a retry doubles the bill against a finite allotment.
Read the handler from the reservation downward, because that order is the entire opinion in code. The actor id and the connection are injected from the session and the pool, not fields the browser may invent. We commit the reserved row before the model speaks, so a crash leaves a visible intention instead of a ghost success. We refuse an empty completion, then we store the text in the same database the API already uses for tickets.
Only after that commit do we return 200, and the body should be read back from the stored row. A local variable is just another hallway, and that hallway gets swept the moment the process exits. The client must reuse the same attempt key when it retries, or the unique constraint cannot protect the allotment. A fresh key is a fresh spend, which is fine when the user asked again and wrong when the browser merely timed out.
I generate that key on the click, keep it beside the button state, and send the same key until the receipt settles. Settled means stored or failed, not a spinner that gave up while the worker was still talking. Reserving before the call has a cost, and I still prefer that cost to a silent double spend. A crashed attempt blocks that key until the reconciler marks it failed, so the button must surface 409 instead of spinning forever.
That friction is the point, because a finite allotment is not a retry budget you can spill on the floor. If you hate the 409, fix the reconciler window, do not delete the unique key and hope the host is kind. The handler below is the proposed shape, and the provider call stays a seam you can swap.
import uuid
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
app = FastAPI()
class SummarizeIn(BaseModel):
ticket_id: str
attempt_key: str
def load_ticket(actor_id: str, ticket_id: str) -> str:
"""Proposed: return persisted text this actor may read, or raise."""
raise NotImplementedError
def call_free_model(prompt: str) -> str:
"""Proposed seam. Do not write the prompt onto the worker disk."""
raise NotImplementedError
def mark_failed(conn, receipt_id: str, error_code: str) -> None:
conn.execute(
"""
update completion_receipts
set status = 'failed', error_code = %s
where id = %s and status = 'reserved'
""",
(error_code, receipt_id),
)
conn.commit()
@app.post("/tickets/summarize")
def summarize(body: SummarizeIn, actor_id: str, conn):
existing = conn.execute(
"""
select id, status, output_text
from completion_receipts
where actor_id = %s and ticket_id = %s and attempt_key = %s
""",
(actor_id, body.ticket_id, body.attempt_key),
).fetchone()
if existing and existing["status"] == "stored":
return {"receipt_id": existing["id"], "summary": existing["output_text"]}
if existing and existing["status"] == "reserved":
raise HTTPException(status_code=409, detail="attempt_in_progress")
if existing and existing["status"] == "failed":
raise HTTPException(status_code=409, detail="attempt_failed")
receipt_id = str(uuid.uuid4())
inserted = conn.execute(
"""
insert into completion_receipts
(id, actor_id, ticket_id, attempt_key, status)
values (%s, %s, %s, %s, 'reserved')
on conflict (actor_id, ticket_id, attempt_key)
do nothing
returning id
""",
(receipt_id, actor_id, body.ticket_id, body.attempt_key),
).fetchone()
if not inserted:
raise HTTPException(status_code=409, detail="attempt_raced")
conn.commit()
try:
source = load_ticket(actor_id, body.ticket_id)
except HTTPException:
mark_failed(conn, receipt_id, "ticket_rejected")
raise
try:
text = call_free_model(source)
except Exception:
mark_failed(conn, receipt_id, "provider_error")
raise HTTPException(status_code=502, detail="provider_error")
if not text or not text.strip():
mark_failed(conn, receipt_id, "empty_completion")
raise HTTPException(status_code=502, detail="empty_completion")
conn.execute(
"""
update completion_receipts
set status = 'stored', output_text = %s
where id = %s and status = 'reserved'
""",
(text, receipt_id),
)
conn.commit()
stored = conn.execute(
"""
select id, output_text
from completion_receipts
where id = %s and status = 'stored'
""",
(receipt_id,),
).fetchone()
if not stored:
raise HTTPException(status_code=500, detail="receipt_missing_after_write")
return {"receipt_id": stored["id"], "summary": stored["output_text"]}
Wire the actor id from your real session dependency before you copy this, or the receipt will trust a caller-supplied owner. The same warning applies to the connection: it must point at a database you administer, not at a directory on the free worker. I left both as arguments so the control flow stays readable, and that readability is not a deployment guide.
The drill I want before anyone calls this production is a kill between the model return and the stored update. You should see a non-200 from the API and a reserved or failed row in the database you control. You should not find the only copy of the paragraph in a temp path on the worker. If curl prints 200 and the select shows no stored row, the slice is still a demo.
# Proposed drill. Point APP_DATABASE_URL at your store, not the worker disk.
curl -sS -D - -o /tmp/sum.json \
-H 'content-type: application/json' \
-d '{"ticket_id":"t_104","attempt_key":"a_9"}' \
http://127.0.0.1:8000/tickets/summarize
psql "$APP_DATABASE_URL" -c \
"select status, error_code, output_text is null as missing
from completion_receipts
where ticket_id = 't_104' and attempt_key = 'a_9';"
A reserved row that sits too long is a worker that died, and leaving it reserved forever blocks the same attempt key. I mark those rows failed with a worker_lost code after a short window I have not tuned in production. Two minutes in the example is a placeholder, not a measured timeout for any particular free server. If your provider is slower than that window, you will fail live calls, so measure your own latency before you copy the interval.
update completion_receipts
set status = 'failed', error_code = 'worker_lost'
where status = 'reserved'
and created_at < now() - interval '2 minutes';
Before I would trust the slice, I want four proofs and I want them in this order. The session must reject a stranger, the reservation must exist before the provider call, and the stored row must match the body we returned. A killed worker must leave a non-success status, and a repeated attempt key must not create a second provider call. If any proof fails, I keep the feature behind the flag, even when the free seat makes the happy path feel finished.
This approach is the wrong tool when your database also lives on that same free host. You have only moved the temp file into a local Postgres directory, and a recycle still deletes the receipt. Skip it if you need a contractual quota, a latency SLO, or a promise that the free seat will still exist next quarter. I am not claiming those properties, and a free option that can change is a prototype seat rather than an architecture.
Streaming tokens into the UI before the receipt commits is a different contract, and this slice will feel too slow for that product. Do not use it as an authorization design either, because a stored paragraph does not prove the actor was allowed to read the ticket. Load the ticket through the same permission check you already trust, and keep the model seam from becoming a side door. If load_ticket is a stub that returns whatever the prompt contains, you have not built the feature, you have built a leak.
If you want a cheap seat for the worker while the receipt stays in a database you administer, look at the current free options. MonkeyCode is the seat I have in mind for that exercise, and only the limits listed on its page today should enter your plan. I already disclosed the outreach relationship above, so I will not repeat a slogan where a constraint belongs.
The handoff I trust least is still that gap between the model return and the commit. Which handoff is least stable in your stack, the session check, the provider call, or the write that should outlive the host? Send me the failure state you actually saw, including the response code, and tell me whether a receipt row existed afterward. If the code was 200 and the row was missing, we are describing the same bug, just on a different ticket.
Top comments (0)