If you put an LLM agent on a public link, every visitor can spend your API quota. My review agent uses Groq for the model and Hindsight for memory, so I wanted a public link that shows the results and cannot be used to run reviews. This article covers how I did that, what I hardened, and what I did not cover.
Two modes, one codebase
The Review Desk is a FastAPI app. On localhost, everything works: you can review a demo pull request, accept or reject comments, and paste a GitHub PR link. On the public link, I turn writes off. That split is controlled by environment variables:
-
READ_ONLY=1blocks/review,/feedbackand/github-pr. -
PUBLIC_MODE=1hides the/docspage. -
BANK_ID=review-agent-publicpoints the public deployment at its own Hindsight memory bank, so it does not share memory with the local one. -
ALLOWED_ORIGINSsets which origins may call the API. There is no wildcard CORS.
The switch lives in main.py. This is the real code:
READ_ONLY = os.getenv("READ_ONLY") == "1"
PUBLIC_MODE = os.getenv("PUBLIC_MODE") == "1"
# ...
app = FastAPI(
title="Code Review Agent with Memory",
lifespan=lifespan,
docs_url=None if PUBLIC_MODE else "/docs",
redoc_url=None if PUBLIC_MODE else "/redoc",
openapi_url=None if PUBLIC_MODE else "/openapi.json",
)
# ...
def block_if_read_only():
if READ_ONLY:
raise HTTPException(403, "Live reviews are turned off on this deployment.")
@app.post("/review")
def review(body: ReviewIn, request: Request):
block_if_read_only()
The public link still serves the pages that do not cost anything to show: the Conventions, Results, Replay and How it works pages, and the measured results.
The deployment
The public copy runs on Render as a free web service. The build command is pip install -r requirements.txt. The start command is python -m uvicorn main:app --host 0.0.0.0 --port $PORT. The environment variables are HINDSIGHT_API_URL, HINDSIGHT_API_KEY, GROQ_API_KEY, READ_ONLY=1, PUBLIC_MODE=1, BANK_ID=review-agent-public and PYTHON_VERSION=3.11.9. The .env file is never committed. Pushing to the repository redeploys the site automatically.
A free Render service sleeps when idle, so the first load can take 30 to 90 seconds. I would rather say that in the article than have someone think the link is broken.
The decision I would defend hardest is the separate memory bank. While building the app I accepted and rejected plenty of comments just to test the interface, and none of those clicks should shape what a visitor sees on the public copy. Giving the public deployment its own BANK_ID in render.yaml means my test decisions can never leak into it. The trade-off I accepted is the slow first load. A free host that sleeps is a fair price for a link that costs nothing to leave online, and a warning line in the article is cheaper than paying to keep it awake.
Before and after
Here is the difference between the two modes in practice.
| Localhost | Public link | |
|---|---|---|
| Run a review | Yes | Blocked by READ_ONLY
|
| Accept or reject a comment | Yes | Blocked by READ_ONLY
|
| Paste a GitHub PR link | Yes | Blocked by READ_ONLY
|
/docs page |
Visible | Hidden by PUBLIC_MODE
|
| Hindsight bank | Local bank | review-agent-public |
| Results and Replay pages | Yes | Yes |
The design goal is that a stranger with the link can look at the measured results and cannot cause a model call or a memory write.
Hardening inside the app
Even on localhost, the app should not trust its input. What I added:
- A limit on diff size (200,000 characters).
- Rejection reasons limited to three allowed values.
- Categories sanitized to lowercase letters, digits and underscores.
- Comment text flattened and capped at 500 characters.
- Per-IP rate limits.
- Duplicate votes on a comment return a 409.
- Writes to
decisions.jsonare atomic, and a corrupt file is moved aside to a.badfile. - The in-memory list of comments is capped at 1,000.
-
/github-praccepts only a link shaped likehttps://github.com/owner/repo/pull/N, and only ever callsapi.github.com.
There is no database, so there is no SQL to inject into.
What this does not cover
This is not a security audit, and I want to be specific about the gaps.
- A diff can contain instructions aimed at the model, and prompt injection cannot be fully prevented. It only matters where live review is on, which means localhost.
- There is no login.
- The rate limiter is coarse behind Render's proxy, and the GET routes have no limit.
- There is no Content Security Policy and no integrity hashes on the CDN scripts.
- The
linefield the model returns is not cleaned in/review. -
decisions.jsonis stored as plain text with no encryption at rest. The public link uses HTTPS through Render, which protects data in transit only.
There are three checks I still want to run: confirm that .env never entered the git history, confirm that the write routes return 403 on the live link and /docs returns 404, and run pip-audit on requirements.txt. I will update this article with the results.
What I would add next
The order I would work through the gaps is the order of how much damage each one can do. First, cap and clean the line field the model returns, because it is a small change. Second, read the real client address behind the proxy, so the per-IP rate limit means something on Render, and put a limit on the GET routes. Third, add a Content Security Policy and integrity hashes for the


Top comments (0)