Originally published on tamiz.pro.
The best code is written while walking. Not at a desk, not in an IDE, not even on a laptop balanced on a park bench — but in the rhythm of footsteps, the cadence of breath, the ambient noise of a neighborhood waking up. The Touch Grass Manifesto isn't wellness advice. It's a technical constraint: build AI applications that cannot be used sitting down.
This article dissects the architecture of screen-free, mobility-first AI systems. We'll cover local-first model deployment, voice-only interaction patterns, sensor-fused storytelling engines, and the infrastructure that makes "walk-driven development" not just possible but preferable to chair-bound coding.
Table of Contents
- 1. The Constraint as Architecture
- 2. Offline-First Model Stack
- 3. Voice-Only Interaction Layer
- 4. Walk-Driven Storytelling Engine
- 5. Power, Thermal, and Hardware Realities
- 6. Developer Experience: Coding While Walking
- 7. Failure Modes and Graceful Degradation
- 8. Frequently Asked Questions
1. The Constraint as Architecture
Most mobile apps treat mobility as a feature. The Touch Grass Manifesto treats immobility as a bug. The core architectural principle: if the user can operate the app while stationary, the design has failed.
This constraint cascades through every layer:
| Layer | Traditional Mobile | Touch Grass Architecture |
|---|---|---|
| Compute | Cloud-first, fallback to local | Local-first, cloud only for sync |
| Input | Touch, keyboard, voice | Voice + sensors only |
| Output | Screen, haptics, audio | Audio + haptics only |
| Context | GPS optional | Motion + location + env required |
| Session | Tap-to-start | Starts on walk detection |
| Error | Modal dialog | Spoken recovery + continue |
The constraint eliminates entire categories of complexity: no responsive layouts, no focus management, no virtual keyboards, no screen readers to test against (the app is a screen reader), no visual regression tests. What remains is harder: reliability without visual confirmation.
1.1 The "No Pixel" Invariant
// Core invariant: zero screen dependencies in the runtime
// This compiles away if any UI framework is linked
#[cfg(target_os = "ios")]
compile_error!("Touch Grass apps cannot link UIKit. Use AudioToolbox + CoreMotion only.");
#[cfg(target_os = "android")]
compile_error!("Touch Grass apps cannot link Android View system. Use AudioTrack + SensorManager only.");
This isn't performative. It forces the team to solve hard problems: how does a user know the model loaded? How do they confirm a destructive action? How do they debug without logs on screen?
2. Offline-First Model Stack
Cloud inference breaks the walk. Latency variance, connectivity drops, battery drain from radio — all unacceptable. The model must run on-device, cold-start in <2s, and sustain 15+ tokens/sec on a phone NPU.
2.1 Model Selection Criteria
| Requirement | Target | Rationale |
|---|---|---|
| Size | ≤ 4B params quantized | Fits in 3-4GB RAM, leaves headroom for OS |
| Quantization | 4-bit (Q4_K_M) or 3-bit | Best quality/speed tradeoff on mobile NPUs |
| Context | 8K+ tokens | Walk sessions accumulate context |
| Latency | <200ms first token | Perceived instant on voice |
| Throughput | >15 tok/s sustained | Natural conversation pace |
| Languages | Multilingual | Walks happen everywhere |
Current picks (2024):
- Llama 3.2 3B Instruct (Q4_K_M) — 2.1GB, 18 tok/s on iPhone 15 Pro NPU
- Gemma 2 2B IT (Q4_K_M) — 1.4GB, 22 tok/s, Apache 2.0
- Phi-3.5-mini (Q4_K_M) — 2.2GB, 20 tok/s, strong reasoning
- Qwen 2.5 3B (Q4_K_M) — 2.0GB, best multilingual
2.2 Local Inference Runtime
llama.cpp via llama.rn (React Native) or native Swift/Kotlin bindings remains the most battle-tested path. MLC-LLM and ExecuTorch are promising but less mature for audio-streaming workloads.
// Swift: Minimal llama.cpp wrapper for voice streaming
// TouchGrassInference.swift
import Foundation
import llama
final class TouchGrassInference {
private var ctx: OpaquePointer?
private var batch: llama_batch
private let nCtx = 8192
private let nThreads = ProcessInfo.processInfo.activeProcessorCount - 1
init(modelPath: String) throws {
var params = llama_context_default_params()
params.n_ctx = Int32(nCtx)
params.n_threads = Int32(nThreads)
params.n_threads_batch = Int32(nThreads)
params.offload_kqv = true // Metal/GPU offload
params.flash_attn = true
ctx = llama_init_from_file(modelPath, params)
guard ctx != nil else { throw InferenceError.modelLoadFailed }
batch = llama_batch_init(512, 0, 1)
}
// Streaming callback for token-by-token TTS handoff
func generate(
prompt: String,
onToken: @escaping (String) -> Bool // return false to stop
) throws {
let tokens = tokenize(prompt)
batch.n_tokens = Int32(tokens.count)
for (i, tok) in tokens.enumerated() {
batch.token[i] = tok
batch.pos[i] = Int32(i)
batch.n_seq_id[i][0] = 0
batch.logits[i] = (i == tokens.count - 1) ? 1 : 0
}
if llama_decode(ctx!, batch) != 0 { throw InferenceError.decodeFailed }
var nPast = tokens.count
while true {
let logits = llama_get_logits_ith(ctx!, Int32(batch.n_tokens - 1))
let nextTok = sampleTopP(logits, p: 0.9, temp: 0.7)
if llama_token_is_eog(ctx!, nextTok) { break }
let piece = String(cString: llama_token_to_piece(ctx!, nextTok))
if !onToken(piece) { break } // TTS or user interrupted
batch.n_tokens = 1
batch.token[0] = nextTok
batch.pos[0] = Int32(nPast)
batch.logits[0] = 1
nPast += 1
if llama_decode(ctx!, batch) != 0 { throw InferenceError.decodeFailed }
if nPast >= nCtx { throw InferenceError.contextFull }
}
}
deinit { if let ctx { llama_free(ctx) } }
}
Key optimization: The onToken callback hands each token directly to the TTS engine before the full response completes. This cuts perceived latency from "model done" to "first phoneme" — critical for voice-only UX.
2.3 Model Swapping Without Screen
Users need to switch models (coding assistant → storytelling → navigation) without looking. Solution: voice-triggered model hot-swap with haptic confirmation.
// ModelManager.swift
enum ModelProfile: String, CaseIterable {
case coder = "llama-3.2-3b-coder-q4"
case storyteller = "gemma-2-2b-story-q4"
case navigator = "phi-3.5-mini-nav-q4"
var hapticPattern: CHHapticPattern { /* distinct per model */ }
var earcon: String { /* unique 200ms audio ID */ }
}
final class ModelManager {
private var current: TouchGrassInference?
private let audio: AudioEngine
private let haptics: HapticEngine
func switchTo(_ profile: ModelProfile) async throws {
// Play earcon *before* unload so user hears intent
try await audio.playEarcon(profile.earcon)
try await haptics.play(profile.hapticPattern)
current = try TouchGrassInference(modelPath: profile.path)
// Confirmation: speak model name via TTS
try await audio.speak("Switched to \(profile.rawValue)")
}
}
3. Voice-Only Interaction Layer
No screen means no visual turn-taking cues. The voice stack must handle: barge-in, endpointing, noise rejection, speaker diarization, and implicit confirmation — all locally.
3.1 Audio Pipeline Architecture
┌─────────────┐ ┌──────────────┐ ┌─────────────┐ ┌──────────────┐
│ Mic Array │──▶│ VAD + AEC │──▶│ ASR Stream │──▶│ Intent/Slot │
│ (beamform) │ │ (RNNoise) │ │ (Whisper) │ │ (LLM) │
└─────────────┘ └──────────────┘ └─────────────┘ └──────────────┘
▲ │ │
│ ┌──────────────────┘ ▼
│ ▼ ┌─────────────────┐
│ ┌──────────────┐ │ TTS (Piper/ │
└─────────────│ Barge-in │◀─────────────────│ StyleTTS2) │
│ Detection │ token stream └─────────────────┘
└──────────────┘ │
│ ▼
▼ ┌─────────────────┐
┌──────────────┐ │ Audio Mixer │
│ Turn Manager│ │ (spatialized) │
└──────────────┘ └─────────────────┘
3.2 Streaming ASR with Whisper.cpp
Whisper.cpp supports streaming via whisper_full_with_state but it's frame-based. For true streaming, use faster-whisper (CTranslate2) or whisper.cpp with tiny.en + sliding window.
# faster-whisper streaming wrapper
# asr_stream.py
from faster_whisper import WhisperModel
import numpy as np
import queue
import threading
class StreamingASR:
def __init__(self, model_size="tiny.en", device="cpu", compute_type="int8"):
self.model = WhisperModel(model_size, device=device, compute_type=compute_type)
self.audio_queue = queue.Queue()
self.result_queue = queue.Queue()
self.running = False
self.buffer = np.array([], dtype=np.float32)
self.sample_rate = 16000
self.chunk_duration = 0.5 # 500ms chunks
def start(self):
self.running = True
threading.Thread(target=self._process_loop, daemon=True).start()
def push_audio(self, chunk: np.ndarray):
self.audio_queue.put(chunk)
def _process_loop(self):
while self.running:
try:
chunk = self.audio_queue.get(timeout=0.1)
self.buffer = np.concatenate([self.buffer, chunk])
# Process when we have enough context
if len(self.buffer) >= self.sample_rate * 2: # 2s window
segments, _ = self.model.transcribe(
self.buffer[-self.sample_rate*4:], # last 4s
language="en",
vad_filter=True,
vad_parameters=dict(min_silence_duration_ms=500)
)
text = " ".join(s.text for s in segments).strip()
if text:
self.result_queue.put(text)
# Keep last 1s for context overlap
self.buffer = self.buffer[-self.sample_rate:]
except queue.Empty:
continue
3.3 Barge-In: The Critical UX Primitive
Users will interrupt. The system must stop TTS instantly and pivot to listening.
// BargeInManager.swift
final class BargeInManager {
private let vad: VoiceActivityDetector // Silero VAD, 1.5ms/frame
private let tts: StreamingTTS
private let asr: StreamingASR
private var isSpeaking = false
func onTTSStart() { isSpeaking = true }
func onTTSChunk(_ audio: Data) {
// Feed TTS output back to VAD for echo cancellation reference
vad.referenceSignal(audio)
}
func onMicAudio(_ audio: Data) -> BargeInDecision {
let voiceProb = vad.process(audio)
if isSpeaking && voiceProb > 0.85 {
// User speaking over TTS → barge in
tts.stopImmediately()
isSpeaking = false
asr.reset() // Clear any partial hypothesis
return .bargeIn
}
if !isSpeaking && voiceProb > 0.6 {
return .startListening
}
return .continue
}
}
Latency budget: VAD must decide in <20ms. Silero VAD (ONNX, 1.5MB) runs at ~0.5ms/frame on mobile NPU. The TTS stop must be synchronous — no fade-out, hard cutoff.
3.4 Implicit Confirmation Patterns
No "tap to confirm." Use progressive commitment:
// ConfirmationStrategy.swift
enum ConfirmationLevel {
case none // Read-only, low risk
case implicit // "I'll save that note" → 3s undo window
case explicit // "Delete all notes? Say 'yes delete' to confirm"
case multiModal // "Say 'yes' AND tap phone twice" (for destructive)
}
func confirmationLevel(for action: Action) -> ConfirmationLevel {
switch action {
case .createNote: return .implicit
case .sendMessage: return .explicit
case .deleteAll: return .multiModal
case .runCode: return .explicit // Code execution = explicit
}
}
Implicit confirmation: speak the action, start a 3-second timer. If user says "cancel" or "undo" within window, revert. No beep, no vibration — just the absence of the follow-up "Done."
4. Walk-Driven Storytelling Engine
The killer app: narrative that unfolds at walking pace. Not audiobooks — generative, location-aware, motion-responsive stories where the walk is the interface.
4.1 Sensor Fusion for Narrative Context
// WalkContext.swift
import CoreMotion
import CoreLocation
struct WalkContext {
let pace: Pace // .stroll / .walk / .brisk / .run
let terrain: Terrain // .pavement / .trail / .stairs / .unknown
let environment: Env // .quiet / .street / .park / .transit
let location: CLLocation?
let timeOfDay: TimePhase
let weather: Weather?
let heartRate: Double? // if HealthKit authorized
var narrativeTempo: NarrativeTempo {
switch (pace, terrain) {
case (.stroll, .park): return .reflective
case (.brisk, .pavement): return .urgent
case (.walk, .trail): return .exploratory
case (.run, _): return .action
default: return .neutral
}
}
}
final class WalkContextEngine {
private let motion = CMMotionManager()
private let location = CLLocationManager()
private let altimeter = CMAltimeter()
private var contextContinuation: AsyncStream<WalkContext>.Continuation?
var contextStream: AsyncStream<WalkContext> {
AsyncStream { continuation in
self.contextContinuation = continuation
self.startUpdates()
}
}
private func startUpdates() {
motion.startDeviceMotionUpdates(to: .main) { [weak self] data, _ in
guard let data, let self else { return }
let pace = self.classifyPace(data)
let terrain = self.classifyTerrain(data)
// ... fuse with location, weather, time
let ctx = WalkContext(...)
self.contextContinuation?.yield(ctx)
}
}
private func classifyPace(_ data: CMDeviceMotion) -> Pace {
// Vertical oscillation + step frequency from userAcceleration
let vertical = data.userAcceleration.z
let freq = self.stepFrequency(from: data)
switch (freq, abs(vertical)) {
case (0..<1.5, _): return .stroll
case (1.5..<2.0, 0..<0.15): return .walk
case (1.5..<2.0, _): return .brisk
case (2.0..., _): return .run
default: return .unknown
}
}
}
4.2 Narrative State Machine
The story isn't a linear script. It's a state graph where nodes are narrative beats, edges are transitions gated by walk context.
// story_graph.json - loaded at startup, no screen needed
{
{
"nodes": {
"start": {
"text": "The trailhead is empty. Your boots hit dirt. What's the first thing you notice?",
"choices": [
{ "label": "The smell of pine", "next": "pine", "requires": {} },
{ "label": "Birdsong overhead", "next": "birds", "requires": {} },
{ "label": "The weight of your pack", "next": "pack", "requires": {} }
]
},
"pine": {
"text": "Pine needles cushion each step. The air tastes like resin and rain. A deer watches from thirty meters.",
"choices": [
{ "label": "Freeze. Watch it leave.", "next": "deer_gone", "requires": {} },
{ "label": "Whistle low. See if it approaches.", "next": "deer_curious", "requires": { "has_whistle": true } }
]
},
"birds": {
"text": "A Steller's jay scolds you from a cedar. Two chickadees mob it. The canopy is alive with argument.",
"choices": [
{ "label": "Identify the jay's alarm call.", "next": "bird_id", "requires": { "skill": "birding" } },
{ "label": "Walk on. They're not your problem.", "next": "ridge", "requires": {} }
]
},
"pack": {
"text": "Twenty kilos. Water, shelter, the satellite messenger your partner insisted on. Every gram earned its place.",
"choices": [
{ "label": "Adjust the hip belt. Keep moving.", "next": "ridge", "requires": {} },
{ "label": "Dump the camp chair. Save six hundred grams.", "next": "lighter", "requires": { "has_chair": true } }
]
},
"ridge": {
"text": "The trail breaks onto a ridge. Valley spreads below—river silver, forest green, the highway a gray scar.",
"choices": [
{ "label": "Sit. Eat the apple you packed.", "next": "apple", "requires": { "has_apple": true } },
{ "label": "Pull out the messenger. Check in.", "next": "checkin", "requires": { "has_messenger": true } },
{ "label": "Just breathe. No screens.", "next": "breathe", "requires": {} }
]
},
"breathe": {
"text": "Wind carries wet earth and snowmelt. Your shoulders drop. This is why you came.",
"is_ending": true,
"ending_type": "pure"
},
"checkin": {
"text": "Three taps. 'All good. On ridge. Back by dark.' The satellite swallows it. Peace of mind, delivered.",
"is_ending": true,
"ending_type": "connected"
}
},
"entry": "start",
"meta": {
"version": "1.0",
"estimated_walk_minutes": 45,
"difficulty": "moderate"
}
}
The graph lives in a static JSON file. No database, no CMS, no admin panel. Writers edit it in VS Code. Version control is the content history.
Runtime evaluation is a pure function:
# story_engine.py - zero dependencies, ~80 lines
import json
from dataclasses import dataclass
from typing import Optional
@dataclass
class WalkContext:
step_count: int
elevation_gain_m: float
heart_rate_bpm: Optional[int]
inventory: set[str]
skills: set[str]
time_of_day: str # "dawn" | "day" | "dusk" | "night"
weather: str # "clear" | "cloudy" | "rain" | "storm"
@dataclass
class Choice:
label: str
next_node: str
requires: dict
@dataclass
class Node:
text: str
choices: list[Choice]
is_ending: bool = False
ending_type: str = ""
def load_graph(path: str) -> dict[str, Node]:
with open(path) as f:
raw = json.load(f)
nodes = {}
for nid, nd in raw["nodes"].items():
nodes[nid] = Node(
text=nd["text"],
choices=[Choice(**c) for c in nd.get("choices", [])],
is_ending=nd.get("is_ending", False),
ending_type=nd.get("ending_type", "")
)
return nodes
def available_choices(node: Node, ctx: WalkContext) -> list[Choice]:
"""Filter choices by context. Pure, testable, no I/O."""
def check(req: dict) -> bool:
for k, v in req.items():
if k == "has_whistle" and v and "whistle" not in ctx.inventory:
return False
if k == "has_chair" and v and "camp_chair" not in ctx.inventory:
return False
if k == "has_apple" and v and "apple" not in ctx.inventory:
return False
if k == "has_messenger" and v and "sat_messenger" not in ctx.inventory:
return False
if k == "skill" and v not in ctx.skills:
return False
return True
return [c for c in node.choices if check(c.requires)]
def render_node(node: Node, ctx: WalkContext) -> str:
lines = [node.text]
for i, ch in enumerate(available_choices(node, ctx), 1):
lines.append(f" {i}. {ch.label}")
return "\n".join(lines)
No LLM calls at runtime. The story is authored, not generated. LLMs are a design-time tool—writers use them to brainstorm branches, check for dead ends, translate. The shipped artifact is deterministic.
4.3 The Audio Pipeline: TTS That Doesn't Suck
Text-to-speech on mobile in 2024 is a solved problem if you accept tradeoffs.
| Approach | Latency | Quality | Offline | Battery |
|---|---|---|---|---|
| Cloud (ElevenLabs, OpenAI) | 800ms–2s | ★★★★★ | ❌ | Low (network only) |
| On-device (piper, whisper.cpp) | 50–200ms | ★★★☆☆ | ✅ | Medium |
| Hybrid (cache + cloud fallback) | 50ms cached | ★★★★☆ | Partial | Low |
Our choice: hybrid with aggressive pre-generation.
At build time, every node's text is rendered to .opus files via Piper (local, fast, decent voices). The app bundles ~50 MB of audio—trivial for modern phones. Zero runtime TTS latency. Zero network dependency.
# build_audio.sh - runs in CI, commits artifacts
#!/usr/bin/env bash
set -euo pipefail
GRAPH="story_graph.json"
OUT_DIR="assets/audio"
VOICE="en_US-lessac-medium" # Piper voice model
mkdir -p "$OUT_DIR"
# Extract all unique node texts
jq -r '.nodes[].text' "$GRAPH" | sort -u | while IFS= read -r text; do
# Stable filename: sha256 of text
fname=$(echo -n "$text" | sha256sum | cut -c1-16)
out="$OUT_DIR/$fname.opus"
[[ -f "$out" ]] && continue # idempotent
echo "$text" | piper --model "$VOICE" --output_raw | \
opusenc --bitrate 24 --raw --raw-channels 1 --raw-rate 22050 - "$out"
done
# Generate manifest for the app
jq -r '.nodes | to_entries[] | "\(.key) \(.value.text)"' "$GRAPH" | \
while read -r nid text; do
fname=$(echo -n "$text" | sha256sum | cut -c1-16)
echo "{\"node\":\"$nid\",\"audio\":\"$fname.opus\"}"
done > "$OUT_DIR/manifest.jsonl"
Runtime playback is a single AVAudioPlayer / MediaPlayer call. The manifest maps node ID → audio file. No streaming, no buffering, no "waiting for voice to load."
Edge case: dynamic text (step count, heart rate). Solved by template fragments.
// story_graph.json snippet
{
"nodes": {
"checkin": {
"text": "Three taps. 'All good. {{step_count}} steps. {{elevation}} meters up. Back by dark.'",
"audio_template": "checkin_base.opus",
"dynamic_slots": ["step_count", "elevation"]
}
}
}
Pre-render the static base ("Three taps. 'All good. ... steps. ... meters up. Back by dark.'"). At runtime, splice in tiny TTS fragments for numbers only—generated on-device via Piper's espeak-ng backend in <50ms. The splice is seamless because prosody matches.
4.4 Sensor Fusion Without The PhD
You don't need a Kalman filter. You need heuristics that survive real users.
# context_builder.py - runs on phone, updates every 5s
from dataclasses import dataclass
from enum import Enum
import time
class Weather(Enum):
CLEAR = "clear"
CLOUDY = "cloudy"
RAIN = "rain"
STORM = "storm"
@dataclass
class WalkContext:
step_count: int
elevation_gain_m: float
heart_rate_bpm: int | None
inventory: set[str]
skills: set[str]
time_of_day: str
weather: Weather
class ContextBuilder:
# ponytail: thresholds tuned for iOS/Android sensor noise profiles.
# Upgrade path: per-device calibration on first run.
PACE_WINDOW_SEC = 30
HR_MIN_VALID = 40
HR_MAX_VALID = 220
ELEVATION_SMOOTH_ALPHA = 0.3
def __init__(self):
self._step_timestamps: list[float] = []
self._elevation_ewma: float | None = None
self._last_pressure_hpa: float | None = None
def update_steps(self, new_steps: int, timestamp: float) -> int:
self._step_timestamps.append(timestamp)
# Drop old
cutoff = timestamp - self.PACE_WINDOW_SEC
self._step_timestamps = [t for t in self._step_timestamps if t > cutoff]
return new_steps
def current_pace_spm(self) -> float:
if len(self._step_timestamps) < 2:
return 0.0
dt = self._step_timestamps[-1] - self._step_timestamps[0]
return len(self._step_timestamps) / dt * 60 # steps/min
def update_elevation(self, pressure_hpa: float) -> float:
"""Barometric altitude. No GPS vertical—too noisy."""
if self._last_pressure_hpa is None:
self._last_pressure_hpa = pressure_hpa
self._elevation_ewma = 0.0
return 0.0
# Hypsometric formula, simplified
delta_h = 8.3 * (self._last_pressure_hpa - pressure_hpa) # meters per hPa approx
self._elevation_ewma = (
self.ELEVATION_SMOOTH_ALPHA * delta_h +
(1 - self.ELEVATION_SMOOTH_ALPHA) * self._elevation_ewma
)
self._last_pressure_hpa = pressure_hpa
return max(0.0, self._elevation_ewma)
def update_heart_rate(self, bpm: int | None) -> int | None:
if bpm is None:
return None
if not (self.HR_MIN_VALID <= bpm <= self.HR_MAX_VALID):
return None # discard artifact
return bpm
def infer_weather(self, pressure_hpa: float, humidity: float, temp_c: float) -> Weather:
# ponytail: naive heuristic. Ceiling: no microclimate awareness.
# Upgrade: on-device TinyML model (TensorFlow Lite, <100KB).
if pressure_hpa < 990 and humidity > 80:
return Weather.STORM
if pressure_hpa < 1000 and humidity > 70:
return Weather.RAIN
if humidity > 60:
return Weather.CLOUDY
return Weather.CLEAR
def time_of_day(self) -> str:
h = time.localtime().tm_hour
if 5 <= h < 8: return "dawn"
if 8 <= h < 18: return "day"
if 18 <= h < 21: return "dusk"
return "night"
No ML models shipped. The heuristic is 30 lines, auditable, explainable. If it misclassifies "cloudy" as "clear" once, the user hears a slightly mismatched line. Nobody dies. The app keeps walking.
4.5 The "No Screen" Contract
The app never wakes the screen. Not for notifications, not for choices, not for errors.
// iOS: AudioSession + Background Modes = "audio" + "location updates"
// Android: Foreground Service (mediaPlayback) + PARTIAL_WAKE_LOCK
class AudioSessionManager {
func configure() throws {
let session = AVAudioSession.sharedInstance()
try session.setCategory(.playback, mode: .spokenAudio, options: [.duckOthers, .mixWithOthers])
try session.setActive(true)
// Lock screen controls appear automatically. No UI code needed.
}
}
User interaction happens via:
- Headphone buttons — single press = choice 1, double = choice 2, triple = choice 3, long press = repeat current node
-
Voice — "Next", "Repeat", "Choice two" (on-device speech recognition,
SFSpeechRecognizer/SpeechRecognizer, no cloud) - Watch complication — tap to advance, crown to scroll choices (watchOS only, optional)
The phone stays in the pack. The watch stays on the wrist. The headphones stay in the ears.
4.6 Testing: Property-Based, Not Example-Based
# test_story_engine.py
import hypothesis.strategies as st
from hypothesis import given, settings
from story_engine import load_graph, available_choices, WalkContext
GRAPH = load_graph("story_graph.json")
@given(
step_count=st.integers(0, 50000),
elevation=st.floats(0, 3000),
hr=st.one_of(st.none(), st.integers(30, 250)),
inventory=st.sets(st.sampled_from(["whistle", "camp_chair", "apple", "sat_messenger"])),
skills=st.sets(st.sampled_from(["birding", "navigation", "first_aid"])),
time_of_day=st.sampled_from(["dawn", "day", "dusk", "night"]),
weather=st.sampled_from(["clear", "cloudy", "rain", "storm"]),
)
@settings(max_examples=500, deadline=None)
def test_no_crashes_on_valid_context(step_count, elevation, hr, inventory, skills, time_of_day, weather):
ctx = WalkContext(step_count, elevation, hr, inventory, skills, time_of_day, weather)
for node in GRAPH.values():
# Should never raise, even with nonsense context
_ = available_choices(node, ctx)
@given(nid=st.sampled_from(list(GRAPH.keys())))
def test_every_node_reachable_from_start(nid):
# BFS from entry
from collections import deque
visited = set()
q = deque([GRAPH["start"]])
while q:
node = q.popleft()
if id(node) in visited: continue
visited.add(id(node))
for ch in node.choices:
if ch.next_node == nid:
return # found
q.append(GRAPH[ch.next_node])
assert False, f"Node {nid} unreachable from start"
def test_no_dead_ends_except_endings():
for nid, node in GRAPH.items():
if node.is_ending: continue
# At least one choice with empty requires (always available)
assert any(len(c.requires) == 0 for c in node.choices), f"{nid}: all choices gated"
500 random contexts × 12 nodes = 6,000 executions per CI run. Catches:
- Missing
requireskeys - Typos in
next_noderefs - Unreachable nodes
- Choices that can never fire (over-constrained)
4.7 Shipping: One Binary, Zero Config
trail/
├── Cargo.toml # or pyproject.toml / package.json / go.mod
├── src/
│ ├── main.rs # 200 lines: init sensors → loop → render → play
│ ├── story_engine.rs # the pure logic above
│ ├── context.rs # sensor readers (platform-specific, ~150 lines each)
│ └── audio.rs # playback + manifest lookup
├── assets/
│ ├── story_graph.json
│ └── audio/ # 50 MB .opus + manifest.jsonl (git-lfs)
└── .github/workflows/
├── build.yml # compiles, runs tests, generates audio
└── release.yml # codesign → notarize → upload to TestFlight / Play Console
No feature flags. No A/B tests. No analytics. The app knows nothing about its users. It doesn't phone home. It doesn't check for updates on launch. It's a tool, not a service.
5. What We Didn't Build (And Why)
| Feature | Requested? | Verdict |
|---|---|---|
| User accounts / cloud sync | "Would be nice" | ❌ YAGNI. The walk is local. |
| Social sharing | "Viral growth" | ❌ Antithetical to the premise. |
| AI-generated side quests | "Infinite content" | ❌ Authored > generated. Quality > quantity. |
| Map view | "Safety" | ❌ Paper map in pack. Phone GPS kills battery. |
| Achievements / streaks | "Retention" | ❌ Gamification ruins the point. |
| Accessibility: screen reader | "Compliance" | ✅ Built in—entire app is audio-first. |
| Offline maps | "Safety" | ❌ Separate app (Organic Maps). Unix philosophy. |
6. The Real Metric
Not DAU. Not retention. Not session length.
Walks completed without phone leaving pack.
-- The only query that matters
SELECT COUNT(*)
FROM walks
WHERE phone_screen_on_seconds < 30
AND duration_minutes BETWEEN 20 AND 180
AND ended_at_ridge = true;
Every line of code above serves that metric. If it doesn't, delete it.
7. Build Your Own
The stack is boring on purpose:
- Language: Whatever compiles to a single binary (Rust, Go, Swift, Kotlin/Native)
- Audio: Piper TTS (build-time) + Opus (runtime)
- Sensors: Platform APIs directly—no wrappers, no abstractions
- Story: JSON + pure functions
- Tests: Hypothesis / proptest / quickcheck
- Distribution: App stores + F-Droid / GitHub Releases
Start this weekend.
- Write 10 nodes in
story_graph.json. - Run
build_audio.sh. - Write the 200-line main loop.
- Walk.
The trail doesn't care about your tech stack. It only cares that you're on it.
Touch grass. The code will wait.
Top comments (0)