DEV Community

Nardeep Singh
Nardeep Singh

Posted on

Steadywag: I Vibe-Coded a Dog Care App Where Missing Paperwork Is Data

Sanity Challenge Path Two Submission

This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange

What I Built

Steadywag is a care companion for a dog with a chronic illness, and the strange part is what it is about: the things his paperwork does not say.

My dog Theo is an 8-year-old Shih Tzu mix with copper storage liver disease, recurrent pancreatitis and high triglycerides. His care lives in about 180 pages of PDFs, emails and things said out loud in exam rooms. One medication is still on his written list, but a specialist told us in person to stop it, and no document says so. A sitter, an emergency vet or a chatbot reading that list would get it wrong.

So I built an app where absence is data. A Sanity record can say a medication is "listed, not given," that a fact is "conflicting" across two reports, or that a lab result a specialist promised never arrived. An agent reads all of it, and the same record drives a day plan, a sitter brief, a daily check-in log, an appointments calendar and a vet-visit prep page. It installs as an app on a phone or a computer.

It is for families who care for a sick pet, and first of all for mine.

Demo

Try these on the first screen:

  • What does he need today? It opens with the medication that is listed but never given.
  • I'm watching Theo. It writes a sitter brief you can print.
  • Can he have blueberries? It answers with a verdict, then shows its checks against his own diet plan.
  • Can he have mango? His plan does not cover it, so it looks up USDA numbers and labels them "not from his vet."

The app limits questions each day to keep the demo running. Every other page works without the assistant.

Code

GitHub logo narlynars07 / steadywag

A care companion for a dog with a chronic illness. An agent on Sanity Context that tells you what the records don't say. DEV Sanity Challenge, Path One.

Steadywag

A care companion for a dog with a chronic illness. It reads his records and tells you what the paperwork doesn't say.

Built on one real case: Theo, an 8-year-old Shih Tzu mix with copper storage hepatopathy, recurrent pancreatitis and high triglycerides. His care lived in about 180 pages of reports, emails and things said out loud in exam rooms. Steadywag turns that into one structured, sourced record in Sanity, then puts an agent, a daily plan, check-ins and appointments on top of it. Built for the DEV Sanity Challenge, Path One: an agent that queries real content.

The idea in one example. A medication was on the written list, but the dog never got it: a specialist told one caregiver out loud to skip it, and nobody updated the paper. The structure knows (medication.status = "listed-not-given" plus a recordGap). The PDFs don't. The agent, the…

My Build Process

The setup

I used Claude Code in a terminal as my AI IDE, plus Claude in Chrome to drive a real browser so it could look at its own work. I direct the product. I did not write the code. That meant my job was deciding what to build, supplying the facts only a family knows, and approving every step that touches real data or goes public.

The working rules I set early, and kept:

  • Plan, then wait. For anything that was a real decision (stack, storage, what is public), Claude Code had to lay out the choice and wait for my answer. I never let it present a decision as settled.
  • Nothing loads without approval. The dataset was built by script, dry-run, then loaded only after I said so. The family's own observations sat behind a flag that stayed off until I approved them.
  • Check before claiming. Before it told me something was fixed, it had to show me: a screenshot at a real screen width, a query result, a test.

The final repo has 40 commits over about two days, around 8,900 lines of app code and 1,300 lines of Sanity schema.

Prompts that worked

Short, specific, and usually a reaction to something I could see. These are real messages from my session, typos left in:

  • "how can we make the agent more useful… separate knowledge base? web search? guardrails?" This one reshaped the agent. It came back with verdict-first food answers ("Fits his plan's rules," "Probably not," "No," "Can't tell"), a small scope gate, and a web search limited to a short list of veterinary sources.
  • "can you rate this answer? he didnt find the label information" (with a screenshot). Claude rated its own answer honestly: it had implied it could search for a product's label, and it cannot. The fix was a rule: ask for the ingredient list instead of pretending.
  • "is this saved only on the browser because it's a demo, should we explain it in the video?" This surfaced a design decision I had never made out loud. Check-ins and appointments stay in the visitor's browser on purpose, because they are private health details. Now every page says so at the top.
  • "on the today section the card with the info about the medication that we are not giving him takes most of the screen can that card be minimizes" A one-line complaint with a screenshot. The card went from about 540px to 155px, with the warning still always visible.
  • "can we do a security check make sure we didnt leave any persinal information, or api keys or tokens open" It scanned the repo, all of git history, the live dataset and the deployed code for 34 real emails, phone numbers and ID numbers from my private files. It found none. It also found two real issues I had not asked about: raw visitor IP addresses in the rate limiter, and missing browser security headers. Both are fixed.

Prompts and approaches that didn't work

  • "Can you search for this treat's label?" The agent talked as if it could. A model that sounds helpful will fill a gap with a guess, and a guess about a dog's food is the wrong kind of helpful.
  • Hard-coding data in the page. Early on, the diets from before his diagnosis lived in the page instead of Sanity, so the agent could not see them. It answered a plain question wrongly. That was my build's fault, not the model's.
  • npm audit fix --force. GitHub's Dependabot raised 41 alerts. The suggested fix downgraded two core packages to old major versions. Claude Code read the audit, traced which alerts could reach the live site (none; all were in Sanity's command-line tooling and the linter), and tested overrides on a separate branch first. One override left the install in a broken state, so it rolled that one back.
  • A background test run. A long audit I started in the background only advanced while the model was actively working. I stopped it and ran the rest in the foreground.

Where the model got stuck, and how we got out

1. Blank answers. "What does he need today?" and the sitter brief came back empty. The agent spent all 8 of its tool steps on lookups and never wrote an answer. The fix was a structural one, not a prompt: raise the cap to 12 and make the final step text-only, so an answer always gets written.

2. The desktop chat window cut off answers. This took three stacked fixes. First the page scrolled instead of the chat panel. Then the card stopped being a flex container once a thread opened. Then its grid row was unconstrained. Each fix revealed the next. The only way out was to measure: a long real answer, in a frame at exact desktop widths, until nothing was clipped. The same bug came back later in a different form (Theo's photo clipped at the top of a short window). That time the root cause was centering content in a box shorter than the content.

3. The filter trap, twice. Sanity Context's chart endpoint only sees the document types in its filter, and it reads the schema from the hosted Studio. When I added new types, the agent could not see them. It said, honestly, "I couldn't retrieve his family routine." That honest failure is how we found the problem, both times.

4. The Knowledge Base invented a number. It added a "400-600 borderline" row to a liver-copper table that no source gives, and the agent repeated it. The fix was a standing instruction on the Knowledge Base: do not add reference ranges that no cited source states. The same Knowledge Base also read a date written "3/10" as March 3 instead of March 10. I now spell out months.

5. Environment variables that looked empty. Vercel marks production variables as "sensitive" by default, and they pull back empty. Claude Code read that as "not set," and spent a while on a problem that did not exist. The fix was to set them as non-sensitive.

6. My own security header broke my own test. After the security pass added a header that stops other sites from embedding the app, the layout test Claude Code had been relying on (loading the app in a frame at a phone width) stopped working. That is the header doing its job. I switched to resizing the panel directly.

7. The first audit. I ran 12 questions through the agent. 4 failed: blueberries, the pre-diagnosis diet, the copper reference, and an Atopica date. Every fix was a change to the data or an instruction, never a prompt trick. The audit now lives in the repo as evals/: 24 questions, each with the reason it exists, and a runner. They pass against the live site, and the README says plainly that a pass means "avoided the known traps," not "good."

What was mine, and what was Claude Code's

Claude Code wrote the code, the schema, the scripts and most of the first drafts of the records. I supplied what no document holds: the verbal instruction about the medication, the family's real daily routine, what we noticed before each flare. I also did every step in the Sanity dashboard that the AI could not do for me. Creating the Context endpoints and the Knowledge Base is a dashboard job, so Claude Code wrote me a click-by-click setup sheet (docs/context-setup.md) and I followed it.

The App SDK and Workflows

I did not use the App SDK or Sanity Workflows, and I would rather say so than stretch. The closest thing is a change log stored as data (recordUpdate). When his vet changes the written list, I mark the medication stopped in the Studio, resolve its gap, and add a log entry with who said so. Within about a minute the home page, Today, the sitter brief and the agent all update, because they read the same record. That is a review step encoded in documents, not a Sanity Workflow.

What I would do differently

  • Write the audit questions on day one, not after the first failures.
  • Model the sources first. Once the vet's words and the family's routine were separate documents, the agent stopped blending them.
  • Run the security scan from the start, not at the end.

Sanity Project Details

  • Project ID: yahsq70q
  • Dataset: production (public read)
  • Dataset query URL: https://yahsq70q.api.sanity.io/v2026-10-01/data/query/production?query=count(*)
  • Studio: https://steadywag.sanity.studio

Schema: 22 document types and one shared object, 516 documents in all (288 lab results, 26 visits, 26 weights, 18 medications). The design choices I would point a reviewer to:

  • sourceNote on every fact from a record: the kind of document, its date, and a confidence of confirmed, single-source or conflicting.
  • recordGap: documents for what the chart does not know. Absence cannot be searched for, so it has to be modeled. There are 9 open gaps.
  • medication.status = "listed-not-given" and medication.writtenInstruction, kept apart from careRoutine, the family's own daily times. The vet's words and the family's routine are different sources, and the agent labels each one.
  • historyChapter and historyPattern, which hold references to the visits, labs and medications they rest on, so every claim points at its evidence.
  • recordUpdate, the change log.

The agent reads the dataset through two Sanity Context endpoints (a GROQ-mode chart endpoint and a Knowledge Base endpoint built from 26 documents of cited guidance and diet rules), read-only, with Claude Sonnet 5.5.

Steadywag tracks and prepares. It never diagnoses, doses or replaces his vet, and it is not medical advice. Theo's records are de-identified: first name and birth year only.

Top comments (0)