DEV Community

Shaikh Ubaid Ahmed
Shaikh Ubaid Ahmed

Posted on AI-assisted

AccessPath-ai: I built my sibling a subway planner that won't strand him when a lift breaks

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🀝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.

What I Built

My sibling relies on step-free transit. A route planner can call a trip accessible and still send them through a transfer that becomes unusable when one lift fails. The usual route card does not say which equipment the journey depends on, when it was last checked, or where to go if it stops working.

I built AccessPath for that situation. It plans step-free NYC Subway journeys, checks the equipment and outage evidence behind each route, and prepares another valid route when the first one breaks. The rider's requirements stay fixed even when the network changes.

The demo starts at Grand Central-42 St and ends at Astoria Blvd. Plan A changes from the 7 to the N at Queensboro Plaza. I inject a clearly labelled failure of the transfer lift, which invalidates that plan and runs the graph search again. Plan B keeps the step-free requirement and transfers at Times Sq-42 St.

Both routes stay on screen so the change is easy to inspect. The result also includes current warnings, model confidence, source timestamps, and an evidence ledger. A simulated outage is always labelled as a simulation, never as a live MTA fact.

Demo

For the stable judging path:

  1. Leave Grand Central-42 St and Astoria Blvd selected.
  2. Keep Step-free and Short transfers checked.
  3. Choose Outage demo.
  4. Press Build my access plan.
  5. Compare the interrupted Queensboro Plaza route with the Times Square reroute.
  6. Open Evidence and review the sources, timestamps, and confidence labels.
  7. Play the spoken journey briefing.

Code

GitHub logo shaikhubaidahmed / accesspath-ai

Evidence-first accessible journey planner for transit riders.

AccessPath

AccessPath is an evidence-first journey planner for riders who depend on step-free transit. It checks the infrastructure a route needs, shows the source behind each access claim, and calculates a backup when a lift failure invalidates the original trip.

The working demonstration plans a trip from Grand Central-42 St to Astoria Blvd. Its original route changes at Queensboro Plaza. A clearly labelled simulated transfer-lift outage blocks that interchange, so the planner finds a second route through Times Sq-42 St without relaxing the rider's step-free requirement.

Why this project exists

Most route planners treat accessibility as a filter attached to a station. A wheelchair user needs a stronger answer. The station entrance, platform, direction of travel, transfer path, and the equipment serving them all have to work at the same time.

My sibling relies on step-free transit, but ordinary route planners do not reliably account for inaccessible transfers or elevator…

The repository contains the app, tests, deployment manifests, and a reproducible public-data pipeline. It also includes the TabPFN benchmark, the Tinker training and evaluation pipeline, and partner activation evidence. A checked-in replay fixture keeps the judging path available when an external service is down.

git clone https://github.com/shaikhubaidahmed/accesspath-ai.git
cd accesspath-ai
npm install
npm run dev
Enter fullscreen mode Exit fullscreen mode

How I Built It

Where models and code separate

I decided early that the model should interpret what a rider says, while official data and deterministic code should decide whether a route is usable.

When configured, Gemma turns a short rider note into typed preferences such as stepFree, avoidLongWalks, and needsAccessibleToilet. It also writes a short briefing after the route has been calculated. Gemma cannot mark a station accessible, clear an outage, or loosen a hard preference to save time.

Official MTA records and the graph search handle those decisions. AccessPath checks every start, arrival, and transfer point. It adds a larger cost to transfers when the rider asks for shorter station changes, and it blocks an interchange when an outage affects a required platform path.

The Mastra workflow

Mastra runs the planning process as five typed steps:

rider note
  -> interpret access needs
  -> resolve MTA evidence
  -> compute constrained route
  -> explain verified result
  -> ground extra official evidence
Enter fullscreen mode Exit fullscreen mode

Each step returns complete or fallback, along with its duration and a plain-language detail. If local Gemma is down, the workflow keeps the rider's explicit checkbox choices and uses local rules. If the live MTA endpoint times out, it switches to the recorded response and reports the fallback in the workflow result.

Building the transit graph

The data script combines the MTA Subway Stations dataset with Regular Subway GTFS. Adjacent stops become travel edges, while connections inside station complexes become walking edges. Equipment and outage feeds provide the current infrastructure state. Monthly availability records provide the reliability history.

The processed graph currently contains 496 stops and the travel and interchange edges used for routing. npm run data:sync rebuilds it from the official endpoints.

There are three evidence modes:

  • Live asks the current MTA outage feed.
  • Replay uses a timestamped response for a stable demonstration.
  • Outage demo adds one local, labelled Queensboro Plaza transfer-lift failure.

Each claim in the evidence ledger has a source type, URL, observation time, freshness note, and confidence level.

Estimating lift risk with TabPFN

The outage feed tells AccessPath what is broken now. TabPFN estimates which lift dependencies deserve more caution based on their history.

The pipeline uses monthly MTA availability and unscheduled-outage records. It creates one-month lags and three-month rolling features, then reserves the latest 20 percent of months for testing. Keeping the split chronological prevents future months from leaking into training.

I compare TabPFN with a histogram gradient-boosting baseline using balanced accuracy, F1, ROC AUC, and runtime. A real TabPFN run writes station-complex forecasts that the TypeScript route engine loads directly.

Model Balanced accuracy F1 ROC AUC
Baseline 0.7878 0.8574 0.8908
TabPFN 0.7888 0.8583 0.8946

The script writes a tabpfn artifact only when it has run with a real token. Fallback scores use a different label.

Fine-tuning the preference extractor with Tinker

The Tinker pipeline fine-tunes Qwen3.5-4B for one task: turning a rider's plain-language note into five typed constraints. The dataset covers wheelchair use, fatigue, crowd sensitivity, toilet access, and explicit walking limits. Four examples are kept out of training for evaluation.

The script evaluates the base model, performs 12 supervised LoRA updates, and evaluates the same holdout again. It records exact JSON match, per-field accuracy, loss, and elapsed time.

  • Base exact match: 0.0%
  • Fine-tuned exact match: 50.0%
  • Base field accuracy: 0.0%
  • Fine-tuned field accuracy: 90.0%
  • Training loss: 1.2858 (step 00) β†’ 0.0002 (step 11)

The complete 12-step loss sequence was:

1.2858 β†’ 0.1938 β†’ 0.0861 β†’ 0.0591 β†’ 0.0380 β†’ 0.0320
       β†’ 0.0219 β†’ 0.0125 β†’ 0.0042 β†’ 0.0020 β†’ 0.00022 β†’ 0.00019
Enter fullscreen mode Exit fullscreen mode

On the four-example holdout, the base model produced no parseable JSON. The fine-tuned model matched two examples exactly and got 18 of 20 individual fields right.

The authenticated rerun preserved the baseline and fine-tuned sampler checkpoints under Tinker training model 828027e1-aa15-5721-abd1-eee47c49e1fd:train:0. Their complete tinker:// paths, model identity, dependency versions, and holdout outputs are in training/reports/tinker-evaluation.json.

Monitoring a route with Temporal

A route planned now may be wrong by the time the rider leaves. The Temporal workflow checks live evidence on a timer, retries failed planning activities with exponential backoff, compares the new route with the previous one, and records any change. If a worker restarts, Temporal reconstructs the monitor from its event history.

I would use this monitor for the hours before a concert or appointment. It records route changes without requiring a browser tab to stay open.

What the other services do

  • ElevenLabs speaks the evidence-bound route briefing. Browser speech synthesis remains available when the API is not configured.
  • Backboard runs the same held-out extraction notes through three open models and records field accuracy, latency, resolved model, and cost.
  • MongoDB Atlas stores expiring journey-plan documents.
  • Tiger Data stores source claims as timestamped evidence events for later time-aware retrieval.
  • SerpApi adds a current result from an official domain when destination evidence is needed.
  • Sentry wraps the planning workflow in an agent span and captures API failures.
  • The Render Blueprint deploys the public web application from the Docker image.
  • The DigitalOcean manifests deploy either the application or the private Ollama and Gemma inference service.

Technical checks

The current repository passes:

4 test files
11 tests
Oxc TypeScript lint
TypeScript client check
TypeScript server check
Vite production build
npm audit: 0 vulnerabilities
production SPA and API smoke test
Enter fullscreen mode Exit fullscreen mode

The tests cover the hard access constraint and the Queensboro Plaza reroute. They also check TabPFN forecast loading with historical fallbacks, evidence labels, all five Mastra steps, malformed API input, health metadata, and Temporal's degraded response.

Focus on open-source infrastructure

A rider may write about wheelchair use, pain, fatigue, or toilet access. AccessPath can process that note with Gemma running through Ollama beside the app, including on a laptop, instead of requiring a closed model provider.

Mastra keeps the orchestration readable and replaceable. The deterministic fallback still enforces hard requirements when no model service is available. If another open-weight model handles the extraction task better, the Backboard benchmark and Tinker evaluation let me measure the difference without rebuilding the product around one vendor.

The same design helps with debugging. I can inspect the profile passed to the route engine, the graph rule that rejected a transfer, the data behind a warning, and the text Gemma received after the route decision. Model output never becomes an accessibility fact on its own.

AccessPath keeps rider notes local where possible. Deterministic code enforces safety constraints, and every factual claim carries a source.

My Agent Session

Pre-submission review of AccessPath, an evidence-first step-free transit planner
You

Review AccessPath before I submit it to the Hacktoberfest Weekend Challenge (Build for a Friend). Check that SUBMISSION.md matches what the repository and the live deployment actually do, and keep secrets out of anything public.

Agent

I'll start with the write-up, then verify each claim against the code, the checks, and the live app.

Agent

Before running anything, a secret scan of every tracked file.

Agent

Now the technical checks the submission lists.

Agent

All four pass. Next, the public deployment and the exact judging path from the write-up.

Agent

I read the remaining public routes, since anything the server exposes is part of the submission too.

Agent

Review result: the write-up holds up.

Verified
- The outage demo behaves as described. The reroute still passes through Queensboro Plaza, but stays on the N there. Only the 7β†’N transfer lift is marked failed, so the step-free requirement is kept rather than relaxed.
- The simulated outage is labelled as a simulation in the outage record, the evidence ledger, and the trace.
- The five-step workflow trace shows each step's status as complete or fallback, as the write-up says.
- Tests, lint, both typechecks, and npm audit pass.
- /api/debug/recent exposes no rider text.

Things to be aware of
- On the public Render deploy, Gemma runs in fallback mode, so the deterministic rules interpret the rider's preferences. The write-up already says "When configured", which is accurate, but a judge who opens the trace will see "fallback" on step 1. A sentence saying the hosted demo uses the rule-based fallback and Gemma runs via Ollama locally or on DigitalOcean would remove the surprise.
- /api/voice forwards any text to ElevenLabs with no rate limit. It is harmless on the current deploy because no ElevenLabs key is set. If one is added, put a limit on it first.
- The local SUBMISSION.md copy edit is not committed or pushed, so the public repo still has the older wording.

You

Prepare a clean session for the submission. Build it from this review only, not from earlier chats.

Agent

Prepared. The session contains only this review. It leaves out:
- credentials, tokens, and account identifiers, including the Tinker training-run ID
- local machine paths and personal contact details
- personal details about the person the project was built for. The transcript describes the users only as riders who depend on step-free transit.
- request IDs from the live API calls

Long tool output is trimmed to the lines that support each claim.

I used an AI coding agent during the challenge window. The work included data-source checks, route-engine debugging, an accessibility review, and a production smoke test. The agent also caught a compiled-path bug after the development build had passed. The embedded session is a curated, secret-free record of the final project review: it re-runs the checks, exercises the outage demo on the live deployment, and scans the repository for secrets.

Prize Categories

I have evidence for these categories:

  • Best Use of Mastra
  • Best Use of Render
  • Best Use of TabPFN
  • Best Use of Tinker

Top comments (0)