DEV Community

Tamiz Uddin
Tamiz Uddin

Posted on Originally published at tamiz.pro

The Touch Grass Manifesto: Building AI Apps That Force Developers Outside

Originally published on tamiz.pro.

The best code is written while walking. Not at a desk, not in an IDE, not even on a laptop balanced on a park bench — but in the rhythm of footsteps, the cadence of breath, the ambient noise of a neighborhood waking up. The Touch Grass Manifesto isn't wellness advice. It's a technical constraint: build AI applications that cannot be used sitting down.

This article dissects the architecture of screen-free, mobility-first AI systems. We'll cover local-first model deployment, voice-only interaction patterns, sensor-fused storytelling engines, and the infrastructure that makes "walk-driven development" not just possible but preferable to chair-bound coding.

Table of Contents

1. The Constraint as Architecture

Most mobile apps treat mobility as a feature. The Touch Grass Manifesto treats immobility as a bug. The core architectural principle: if the user can operate the app while stationary, the design has failed.

This constraint cascades through every layer:

Layer Traditional Mobile Touch Grass Architecture
Compute Cloud-first, fallback to local Local-first, cloud only for sync
Input Touch, keyboard, voice Voice + sensors only
Output Screen, haptics, audio Audio + haptics only
Context GPS optional Motion + location + env required
Session Tap-to-start Starts on walk detection
Error Modal dialog Spoken recovery + continue

The constraint eliminates entire categories of complexity: no responsive layouts, no focus management, no virtual keyboards, no screen readers to test against (the app is a screen reader), no visual regression tests. What remains is harder: reliability without visual confirmation.

1.1 The "No Pixel" Invariant

// Core invariant: zero screen dependencies in the runtime
// This compiles away if any UI framework is linked
#[cfg(target_os = "ios")]
compile_error!("Touch Grass apps cannot link UIKit. Use AudioToolbox + CoreMotion only.");

#[cfg(target_os = "android")]
compile_error!("Touch Grass apps cannot link Android View system. Use AudioTrack + SensorManager only.");
Enter fullscreen mode Exit fullscreen mode

This isn't performative. It forces the team to solve hard problems: how does a user know the model loaded? How do they confirm a destructive action? How do they debug without logs on screen?

2. Offline-First Model Stack

Cloud inference breaks the walk. Latency variance, connectivity drops, battery drain from radio — all unacceptable. The model must run on-device, cold-start in <2s, and sustain 15+ tokens/sec on a phone NPU.

2.1 Model Selection Criteria

Requirement Target Rationale
Size ≤ 4B params quantized Fits in 3-4GB RAM, leaves headroom for OS
Quantization 4-bit (Q4_K_M) or 3-bit Best quality/speed tradeoff on mobile NPUs
Context 8K+ tokens Walk sessions accumulate context
Latency <200ms first token Perceived instant on voice
Throughput >15 tok/s sustained Natural conversation pace
Languages Multilingual Walks happen everywhere

Current picks (2024):

  • Llama 3.2 3B Instruct (Q4_K_M) — 2.1GB, 18 tok/s on iPhone 15 Pro NPU
  • Gemma 2 2B IT (Q4_K_M) — 1.4GB, 22 tok/s, Apache 2.0
  • Phi-3.5-mini (Q4_K_M) — 2.2GB, 20 tok/s, strong reasoning
  • Qwen 2.5 3B (Q4_K_M) — 2.0GB, best multilingual

2.2 Local Inference Runtime

llama.cpp via llama.rn (React Native) or native Swift/Kotlin bindings remains the most battle-tested path. MLC-LLM and ExecuTorch are promising but less mature for audio-streaming workloads.

// Swift: Minimal llama.cpp wrapper for voice streaming
// TouchGrassInference.swift
import Foundation
import llama

final class TouchGrassInference {
    private var ctx: OpaquePointer?
    private var batch: llama_batch
    private let nCtx = 8192
    private let nThreads = ProcessInfo.processInfo.activeProcessorCount - 1

    init(modelPath: String) throws {
        var params = llama_context_default_params()
        params.n_ctx = Int32(nCtx)
        params.n_threads = Int32(nThreads)
        params.n_threads_batch = Int32(nThreads)
        params.offload_kqv = true  // Metal/GPU offload
        params.flash_attn = true

        ctx = llama_init_from_file(modelPath, params)
        guard ctx != nil else { throw InferenceError.modelLoadFailed }

        batch = llama_batch_init(512, 0, 1)
    }

    // Streaming callback for token-by-token TTS handoff
    func generate(
        prompt: String,
        onToken: @escaping (String) -> Bool  // return false to stop
    ) throws {
        let tokens = tokenize(prompt)
        batch.n_tokens = Int32(tokens.count)
        for (i, tok) in tokens.enumerated() {
            batch.token[i] = tok
            batch.pos[i] = Int32(i)
            batch.n_seq_id[i][0] = 0
            batch.logits[i] = (i == tokens.count - 1) ? 1 : 0
        }

        if llama_decode(ctx!, batch) != 0 { throw InferenceError.decodeFailed }

        var nPast = tokens.count
        while true {
            let logits = llama_get_logits_ith(ctx!, Int32(batch.n_tokens - 1))
            let nextTok = sampleTopP(logits, p: 0.9, temp: 0.7)

            if llama_token_is_eog(ctx!, nextTok) { break }

            let piece = String(cString: llama_token_to_piece(ctx!, nextTok))
            if !onToken(piece) { break }  // TTS or user interrupted

            batch.n_tokens = 1
            batch.token[0] = nextTok
            batch.pos[0] = Int32(nPast)
            batch.logits[0] = 1
            nPast += 1

            if llama_decode(ctx!, batch) != 0 { throw InferenceError.decodeFailed }

            if nPast >= nCtx { throw InferenceError.contextFull }
        }
    }

    deinit { if let ctx { llama_free(ctx) } }
}
Enter fullscreen mode Exit fullscreen mode

Key optimization: The onToken callback hands each token directly to the TTS engine before the full response completes. This cuts perceived latency from "model done" to "first phoneme" — critical for voice-only UX.

2.3 Model Swapping Without Screen

Users need to switch models (coding assistant → storytelling → navigation) without looking. Solution: voice-triggered model hot-swap with haptic confirmation.

// ModelManager.swift
enum ModelProfile: String, CaseIterable {
    case coder = "llama-3.2-3b-coder-q4"
    case storyteller = "gemma-2-2b-story-q4"
    case navigator = "phi-3.5-mini-nav-q4"

    var hapticPattern: CHHapticPattern { /* distinct per model */ }
    var earcon: String { /* unique 200ms audio ID */ }
}

final class ModelManager {
    private var current: TouchGrassInference?
    private let audio: AudioEngine
    private let haptics: HapticEngine

    func switchTo(_ profile: ModelProfile) async throws {
        // Play earcon *before* unload so user hears intent
        try await audio.playEarcon(profile.earcon)
        try await haptics.play(profile.hapticPattern)

        current = try TouchGrassInference(modelPath: profile.path)

        // Confirmation: speak model name via TTS
        try await audio.speak("Switched to \(profile.rawValue)")
    }
}
Enter fullscreen mode Exit fullscreen mode

3. Voice-Only Interaction Layer

No screen means no visual turn-taking cues. The voice stack must handle: barge-in, endpointing, noise rejection, speaker diarization, and implicit confirmation — all locally.

3.1 Audio Pipeline Architecture

┌─────────────┐   ┌──────────────┐   ┌─────────────┐   ┌──────────────┐
│  Mic Array  │──▶│  VAD + AEC   │──▶│  ASR Stream │──▶│  Intent/Slot │
│  (beamform) │   │  (RNNoise)   │   │  (Whisper)  │   │  (LLM)       │
└─────────────┘   └──────────────┘   └─────────────┘   └──────────────┘
       ▲                                       │                   │
       │                    ┌──────────────────┘                   ▼
       │                    ▼                          ┌─────────────────┐
       │             ┌──────────────┐                  │  TTS (Piper/    │
       └─────────────│  Barge-in    │◀─────────────────│  StyleTTS2)     │
                     │  Detection   │   token stream   └─────────────────┘
                     └──────────────┘                        │
                            │                                ▼
                            ▼                       ┌─────────────────┐
                     ┌──────────────┐                │  Audio Mixer    │
                     │  Turn Manager│                │  (spatialized)  │
                     └──────────────┘                └─────────────────┘
Enter fullscreen mode Exit fullscreen mode

3.2 Streaming ASR with Whisper.cpp

Whisper.cpp supports streaming via whisper_full_with_state but it's frame-based. For true streaming, use faster-whisper (CTranslate2) or whisper.cpp with tiny.en + sliding window.

# faster-whisper streaming wrapper
# asr_stream.py
from faster_whisper import WhisperModel
import numpy as np
import queue
import threading

class StreamingASR:
    def __init__(self, model_size="tiny.en", device="cpu", compute_type="int8"):
        self.model = WhisperModel(model_size, device=device, compute_type=compute_type)
        self.audio_queue = queue.Queue()
        self.result_queue = queue.Queue()
        self.running = False
        self.buffer = np.array([], dtype=np.float32)
        self.sample_rate = 16000
        self.chunk_duration = 0.5  # 500ms chunks

    def start(self):
        self.running = True
        threading.Thread(target=self._process_loop, daemon=True).start()

    def push_audio(self, chunk: np.ndarray):
        self.audio_queue.put(chunk)

    def _process_loop(self):
        while self.running:
            try:
                chunk = self.audio_queue.get(timeout=0.1)
                self.buffer = np.concatenate([self.buffer, chunk])

                # Process when we have enough context
                if len(self.buffer) >= self.sample_rate * 2:  # 2s window
                    segments, _ = self.model.transcribe(
                        self.buffer[-self.sample_rate*4:],  # last 4s
                        language="en",
                        vad_filter=True,
                        vad_parameters=dict(min_silence_duration_ms=500)
                    )
                    text = " ".join(s.text for s in segments).strip()
                    if text:
                        self.result_queue.put(text)
                    # Keep last 1s for context overlap
                    self.buffer = self.buffer[-self.sample_rate:]
            except queue.Empty:
                continue
Enter fullscreen mode Exit fullscreen mode

3.3 Barge-In: The Critical UX Primitive

Users will interrupt. The system must stop TTS instantly and pivot to listening.

// BargeInManager.swift
final class BargeInManager {
    private let vad: VoiceActivityDetector  // Silero VAD, 1.5ms/frame
    private let tts: StreamingTTS
    private let asr: StreamingASR
    private var isSpeaking = false

    func onTTSStart() { isSpeaking = true }

    func onTTSChunk(_ audio: Data) {
        // Feed TTS output back to VAD for echo cancellation reference
        vad.referenceSignal(audio)
    }

    func onMicAudio(_ audio: Data) -> BargeInDecision {
        let voiceProb = vad.process(audio)

        if isSpeaking && voiceProb > 0.85 {
            // User speaking over TTS → barge in
            tts.stopImmediately()
            isSpeaking = false
            asr.reset()  // Clear any partial hypothesis
            return .bargeIn
        }

        if !isSpeaking && voiceProb > 0.6 {
            return .startListening
        }

        return .continue
    }
}
Enter fullscreen mode Exit fullscreen mode

Latency budget: VAD must decide in <20ms. Silero VAD (ONNX, 1.5MB) runs at ~0.5ms/frame on mobile NPU. The TTS stop must be synchronous — no fade-out, hard cutoff.

3.4 Implicit Confirmation Patterns

No "tap to confirm." Use progressive commitment:

// ConfirmationStrategy.swift
enum ConfirmationLevel {
    case none      // Read-only, low risk
    case implicit  // "I'll save that note" → 3s undo window
    case explicit  // "Delete all notes? Say 'yes delete' to confirm"
    case multiModal // "Say 'yes' AND tap phone twice" (for destructive)
}

func confirmationLevel(for action: Action) -> ConfirmationLevel {
    switch action {
    case .createNote: return .implicit
    case .sendMessage: return .explicit
    case .deleteAll: return .multiModal
    case .runCode: return .explicit  // Code execution = explicit
    }
}
Enter fullscreen mode Exit fullscreen mode

Implicit confirmation: speak the action, start a 3-second timer. If user says "cancel" or "undo" within window, revert. No beep, no vibration — just the absence of the follow-up "Done."

4. Walk-Driven Storytelling Engine

The killer app: narrative that unfolds at walking pace. Not audiobooks — generative, location-aware, motion-responsive stories where the walk is the interface.

4.1 Sensor Fusion for Narrative Context

// WalkContext.swift
import CoreMotion
import CoreLocation

struct WalkContext {
    let pace: Pace          // .stroll / .walk / .brisk / .run
    let terrain: Terrain    // .pavement / .trail / .stairs / .unknown
    let environment: Env    // .quiet / .street / .park / .transit
    let location: CLLocation?
    let timeOfDay: TimePhase
    let weather: Weather?
    let heartRate: Double?  // if HealthKit authorized

    var narrativeTempo: NarrativeTempo {
        switch (pace, terrain) {
        case (.stroll, .park): return .reflective
        case (.brisk, .pavement): return .urgent
        case (.walk, .trail): return .exploratory
        case (.run, _): return .action
        default: return .neutral
        }
    }
}

final class WalkContextEngine {
    private let motion = CMMotionManager()
    private let location = CLLocationManager()
    private let altimeter = CMAltimeter()
    private var contextContinuation: AsyncStream<WalkContext>.Continuation?

    var contextStream: AsyncStream<WalkContext> {
        AsyncStream { continuation in
            self.contextContinuation = continuation
            self.startUpdates()
        }
    }

    private func startUpdates() {
        motion.startDeviceMotionUpdates(to: .main) { [weak self] data, _ in
            guard let data, let self else { return }
            let pace = self.classifyPace(data)
            let terrain = self.classifyTerrain(data)
            // ... fuse with location, weather, time
            let ctx = WalkContext(...)
            self.contextContinuation?.yield(ctx)
        }
    }

    private func classifyPace(_ data: CMDeviceMotion) -> Pace {
        // Vertical oscillation + step frequency from userAcceleration
        let vertical = data.userAcceleration.z
        let freq = self.stepFrequency(from: data)

        switch (freq, abs(vertical)) {
        case (0..<1.5, _): return .stroll
        case (1.5..<2.0, 0..<0.15): return .walk
        case (1.5..<2.0, _): return .brisk
        case (2.0..., _): return .run
        default: return .unknown
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

4.2 Narrative State Machine

The story isn't a linear script. It's a state graph where nodes are narrative beats, edges are transitions gated by walk context.

// story_graph.json - loaded at startup, no screen needed
{

{
  "nodes": {
    "start": {
      "text": "The trailhead is empty. Your boots hit dirt. What's the first thing you notice?",
      "choices": [
        { "label": "The smell of pine", "next": "pine", "requires": {} },
        { "label": "Birdsong overhead", "next": "birds", "requires": {} },
        { "label": "The weight of your pack", "next": "pack", "requires": {} }
      ]
    },
    "pine": {
      "text": "Pine needles cushion each step. The air tastes like resin and rain. A deer watches from thirty meters.",
      "choices": [
        { "label": "Freeze. Watch it leave.", "next": "deer_gone", "requires": {} },
        { "label": "Whistle low. See if it approaches.", "next": "deer_curious", "requires": { "has_whistle": true } }
      ]
    },
    "birds": {
      "text": "A Steller's jay scolds you from a cedar. Two chickadees mob it. The canopy is alive with argument.",
      "choices": [
        { "label": "Identify the jay's alarm call.", "next": "bird_id", "requires": { "skill": "birding" } },
        { "label": "Walk on. They're not your problem.", "next": "ridge", "requires": {} }
      ]
    },
    "pack": {
      "text": "Twenty kilos. Water, shelter, the satellite messenger your partner insisted on. Every gram earned its place.",
      "choices": [
        { "label": "Adjust the hip belt. Keep moving.", "next": "ridge", "requires": {} },
        { "label": "Dump the camp chair. Save six hundred grams.", "next": "lighter", "requires": { "has_chair": true } }
      ]
    },
    "ridge": {
      "text": "The trail breaks onto a ridge. Valley spreads below—river silver, forest green, the highway a gray scar.",
      "choices": [
        { "label": "Sit. Eat the apple you packed.", "next": "apple", "requires": { "has_apple": true } },
        { "label": "Pull out the messenger. Check in.", "next": "checkin", "requires": { "has_messenger": true } },
        { "label": "Just breathe. No screens.", "next": "breathe", "requires": {} }
      ]
    },
    "breathe": {
      "text": "Wind carries wet earth and snowmelt. Your shoulders drop. This is why you came.",
      "is_ending": true,
      "ending_type": "pure"
    },
    "checkin": {
      "text": "Three taps. 'All good. On ridge. Back by dark.' The satellite swallows it. Peace of mind, delivered.",
      "is_ending": true,
      "ending_type": "connected"
    }
  },
  "entry": "start",
  "meta": {
    "version": "1.0",
    "estimated_walk_minutes": 45,
    "difficulty": "moderate"
  }
}
Enter fullscreen mode Exit fullscreen mode

The graph lives in a static JSON file. No database, no CMS, no admin panel. Writers edit it in VS Code. Version control is the content history.

Runtime evaluation is a pure function:

# story_engine.py - zero dependencies, ~80 lines
import json
from dataclasses import dataclass
from typing import Optional

@dataclass
class WalkContext:
    step_count: int
    elevation_gain_m: float
    heart_rate_bpm: Optional[int]
    inventory: set[str]
    skills: set[str]
    time_of_day: str  # "dawn" | "day" | "dusk" | "night"
    weather: str      # "clear" | "cloudy" | "rain" | "storm"

@dataclass
class Choice:
    label: str
    next_node: str
    requires: dict

@dataclass
class Node:
    text: str
    choices: list[Choice]
    is_ending: bool = False
    ending_type: str = ""

def load_graph(path: str) -> dict[str, Node]:
    with open(path) as f:
        raw = json.load(f)
    nodes = {}
    for nid, nd in raw["nodes"].items():
        nodes[nid] = Node(
            text=nd["text"],
            choices=[Choice(**c) for c in nd.get("choices", [])],
            is_ending=nd.get("is_ending", False),
            ending_type=nd.get("ending_type", "")
        )
    return nodes

def available_choices(node: Node, ctx: WalkContext) -> list[Choice]:
    """Filter choices by context. Pure, testable, no I/O."""
    def check(req: dict) -> bool:
        for k, v in req.items():
            if k == "has_whistle" and v and "whistle" not in ctx.inventory:
                return False
            if k == "has_chair" and v and "camp_chair" not in ctx.inventory:
                return False
            if k == "has_apple" and v and "apple" not in ctx.inventory:
                return False
            if k == "has_messenger" and v and "sat_messenger" not in ctx.inventory:
                return False
            if k == "skill" and v not in ctx.skills:
                return False
        return True
    return [c for c in node.choices if check(c.requires)]

def render_node(node: Node, ctx: WalkContext) -> str:
    lines = [node.text]
    for i, ch in enumerate(available_choices(node, ctx), 1):
        lines.append(f"  {i}. {ch.label}")
    return "\n".join(lines)
Enter fullscreen mode Exit fullscreen mode

No LLM calls at runtime. The story is authored, not generated. LLMs are a design-time tool—writers use them to brainstorm branches, check for dead ends, translate. The shipped artifact is deterministic.

4.3 The Audio Pipeline: TTS That Doesn't Suck

Text-to-speech on mobile in 2024 is a solved problem if you accept tradeoffs.

Approach Latency Quality Offline Battery
Cloud (ElevenLabs, OpenAI) 800ms–2s ★★★★★ ❌ Low (network only)
On-device (piper, whisper.cpp) 50–200ms ★★★☆☆ ✅ Medium
Hybrid (cache + cloud fallback) 50ms cached ★★★★☆ Partial Low

Our choice: hybrid with aggressive pre-generation.

At build time, every node's text is rendered to .opus files via Piper (local, fast, decent voices). The app bundles ~50 MB of audio—trivial for modern phones. Zero runtime TTS latency. Zero network dependency.

# build_audio.sh - runs in CI, commits artifacts
#!/usr/bin/env bash
set -euo pipefail

GRAPH="story_graph.json"
OUT_DIR="assets/audio"
VOICE="en_US-lessac-medium"  # Piper voice model

mkdir -p "$OUT_DIR"

# Extract all unique node texts
jq -r '.nodes[].text' "$GRAPH" | sort -u | while IFS= read -r text; do
  # Stable filename: sha256 of text
  fname=$(echo -n "$text" | sha256sum | cut -c1-16)
  out="$OUT_DIR/$fname.opus"
  [[ -f "$out" ]] && continue  # idempotent
  echo "$text" | piper --model "$VOICE" --output_raw | \
    opusenc --bitrate 24 --raw --raw-channels 1 --raw-rate 22050 - "$out"
done

# Generate manifest for the app
jq -r '.nodes | to_entries[] | "\(.key) \(.value.text)"' "$GRAPH" | \
while read -r nid text; do
  fname=$(echo -n "$text" | sha256sum | cut -c1-16)
  echo "{\"node\":\"$nid\",\"audio\":\"$fname.opus\"}"
done > "$OUT_DIR/manifest.jsonl"
Enter fullscreen mode Exit fullscreen mode

Runtime playback is a single AVAudioPlayer / MediaPlayer call. The manifest maps node ID → audio file. No streaming, no buffering, no "waiting for voice to load."

Edge case: dynamic text (step count, heart rate). Solved by template fragments.

// story_graph.json snippet
{
  "nodes": {
    "checkin": {
      "text": "Three taps. 'All good. {{step_count}} steps. {{elevation}} meters up. Back by dark.'",
      "audio_template": "checkin_base.opus",
      "dynamic_slots": ["step_count", "elevation"]
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Pre-render the static base ("Three taps. 'All good. ... steps. ... meters up. Back by dark.'"). At runtime, splice in tiny TTS fragments for numbers only—generated on-device via Piper's espeak-ng backend in <50ms. The splice is seamless because prosody matches.

4.4 Sensor Fusion Without The PhD

You don't need a Kalman filter. You need heuristics that survive real users.

# context_builder.py - runs on phone, updates every 5s
from dataclasses import dataclass
from enum import Enum
import time

class Weather(Enum):
    CLEAR = "clear"
    CLOUDY = "cloudy"
    RAIN = "rain"
    STORM = "storm"

@dataclass
class WalkContext:
    step_count: int
    elevation_gain_m: float
    heart_rate_bpm: int | None
    inventory: set[str]
    skills: set[str]
    time_of_day: str
    weather: Weather

class ContextBuilder:
    # ponytail: thresholds tuned for iOS/Android sensor noise profiles.
    # Upgrade path: per-device calibration on first run.
    PACE_WINDOW_SEC = 30
    HR_MIN_VALID = 40
    HR_MAX_VALID = 220
    ELEVATION_SMOOTH_ALPHA = 0.3

    def __init__(self):
        self._step_timestamps: list[float] = []
        self._elevation_ewma: float | None = None
        self._last_pressure_hpa: float | None = None

    def update_steps(self, new_steps: int, timestamp: float) -> int:
        self._step_timestamps.append(timestamp)
        # Drop old
        cutoff = timestamp - self.PACE_WINDOW_SEC
        self._step_timestamps = [t for t in self._step_timestamps if t > cutoff]
        return new_steps

    def current_pace_spm(self) -> float:
        if len(self._step_timestamps) < 2:
            return 0.0
        dt = self._step_timestamps[-1] - self._step_timestamps[0]
        return len(self._step_timestamps) / dt * 60  # steps/min

    def update_elevation(self, pressure_hpa: float) -> float:
        """Barometric altitude. No GPS vertical—too noisy."""
        if self._last_pressure_hpa is None:
            self._last_pressure_hpa = pressure_hpa
            self._elevation_ewma = 0.0
            return 0.0
        # Hypsometric formula, simplified
        delta_h = 8.3 * (self._last_pressure_hpa - pressure_hpa)  # meters per hPa approx
        self._elevation_ewma = (
            self.ELEVATION_SMOOTH_ALPHA * delta_h +
            (1 - self.ELEVATION_SMOOTH_ALPHA) * self._elevation_ewma
        )
        self._last_pressure_hpa = pressure_hpa
        return max(0.0, self._elevation_ewma)

    def update_heart_rate(self, bpm: int | None) -> int | None:
        if bpm is None:
            return None
        if not (self.HR_MIN_VALID <= bpm <= self.HR_MAX_VALID):
            return None  # discard artifact
        return bpm

    def infer_weather(self, pressure_hpa: float, humidity: float, temp_c: float) -> Weather:
        # ponytail: naive heuristic. Ceiling: no microclimate awareness.
        # Upgrade: on-device TinyML model (TensorFlow Lite, <100KB).
        if pressure_hpa < 990 and humidity > 80:
            return Weather.STORM
        if pressure_hpa < 1000 and humidity > 70:
            return Weather.RAIN
        if humidity > 60:
            return Weather.CLOUDY
        return Weather.CLEAR

    def time_of_day(self) -> str:
        h = time.localtime().tm_hour
        if 5 <= h < 8: return "dawn"
        if 8 <= h < 18: return "day"
        if 18 <= h < 21: return "dusk"
        return "night"
Enter fullscreen mode Exit fullscreen mode

No ML models shipped. The heuristic is 30 lines, auditable, explainable. If it misclassifies "cloudy" as "clear" once, the user hears a slightly mismatched line. Nobody dies. The app keeps walking.

4.5 The "No Screen" Contract

The app never wakes the screen. Not for notifications, not for choices, not for errors.

// iOS: AudioSession + Background Modes = "audio" + "location updates"
// Android: Foreground Service (mediaPlayback) + PARTIAL_WAKE_LOCK

class AudioSessionManager {
    func configure() throws {
        let session = AVAudioSession.sharedInstance()
        try session.setCategory(.playback, mode: .spokenAudio, options: [.duckOthers, .mixWithOthers])
        try session.setActive(true)
        // Lock screen controls appear automatically. No UI code needed.
    }
}
Enter fullscreen mode Exit fullscreen mode

User interaction happens via:

  1. Headphone buttons — single press = choice 1, double = choice 2, triple = choice 3, long press = repeat current node
  2. Voice — "Next", "Repeat", "Choice two" (on-device speech recognition, SFSpeechRecognizer / SpeechRecognizer, no cloud)
  3. Watch complication — tap to advance, crown to scroll choices (watchOS only, optional)

The phone stays in the pack. The watch stays on the wrist. The headphones stay in the ears.

4.6 Testing: Property-Based, Not Example-Based

# test_story_engine.py
import hypothesis.strategies as st
from hypothesis import given, settings
from story_engine import load_graph, available_choices, WalkContext

GRAPH = load_graph("story_graph.json")

@given(
    step_count=st.integers(0, 50000),
    elevation=st.floats(0, 3000),
    hr=st.one_of(st.none(), st.integers(30, 250)),
    inventory=st.sets(st.sampled_from(["whistle", "camp_chair", "apple", "sat_messenger"])),
    skills=st.sets(st.sampled_from(["birding", "navigation", "first_aid"])),
    time_of_day=st.sampled_from(["dawn", "day", "dusk", "night"]),
    weather=st.sampled_from(["clear", "cloudy", "rain", "storm"]),
)
@settings(max_examples=500, deadline=None)
def test_no_crashes_on_valid_context(step_count, elevation, hr, inventory, skills, time_of_day, weather):
    ctx = WalkContext(step_count, elevation, hr, inventory, skills, time_of_day, weather)
    for node in GRAPH.values():
        # Should never raise, even with nonsense context
        _ = available_choices(node, ctx)

@given(nid=st.sampled_from(list(GRAPH.keys())))
def test_every_node_reachable_from_start(nid):
    # BFS from entry
    from collections import deque
    visited = set()
    q = deque([GRAPH["start"]])
    while q:
        node = q.popleft()
        if id(node) in visited: continue
        visited.add(id(node))
        for ch in node.choices:
            if ch.next_node == nid:
                return  # found
            q.append(GRAPH[ch.next_node])
    assert False, f"Node {nid} unreachable from start"

def test_no_dead_ends_except_endings():
    for nid, node in GRAPH.items():
        if node.is_ending: continue
        # At least one choice with empty requires (always available)
        assert any(len(c.requires) == 0 for c in node.choices), f"{nid}: all choices gated"
Enter fullscreen mode Exit fullscreen mode

500 random contexts × 12 nodes = 6,000 executions per CI run. Catches:

  • Missing requires keys
  • Typos in next_node refs
  • Unreachable nodes
  • Choices that can never fire (over-constrained)

4.7 Shipping: One Binary, Zero Config

trail/
├── Cargo.toml           # or pyproject.toml / package.json / go.mod
├── src/
│   ├── main.rs          # 200 lines: init sensors → loop → render → play
│   ├── story_engine.rs  # the pure logic above
│   ├── context.rs       # sensor readers (platform-specific, ~150 lines each)
│   └── audio.rs         # playback + manifest lookup
├── assets/
│   ├── story_graph.json
│   └── audio/           # 50 MB .opus + manifest.jsonl (git-lfs)
└── .github/workflows/
    ├── build.yml        # compiles, runs tests, generates audio
    └── release.yml      # codesign → notarize → upload to TestFlight / Play Console
Enter fullscreen mode Exit fullscreen mode

No feature flags. No A/B tests. No analytics. The app knows nothing about its users. It doesn't phone home. It doesn't check for updates on launch. It's a tool, not a service.


5. What We Didn't Build (And Why)

Feature Requested? Verdict
User accounts / cloud sync "Would be nice" ❌ YAGNI. The walk is local.
Social sharing "Viral growth" ❌ Antithetical to the premise.
AI-generated side quests "Infinite content" ❌ Authored > generated. Quality > quantity.
Map view "Safety" ❌ Paper map in pack. Phone GPS kills battery.
Achievements / streaks "Retention" ❌ Gamification ruins the point.
Accessibility: screen reader "Compliance" ✅ Built in—entire app is audio-first.
Offline maps "Safety" ❌ Separate app (Organic Maps). Unix philosophy.

6. The Real Metric

Not DAU. Not retention. Not session length.

Walks completed without phone leaving pack.

-- The only query that matters
SELECT COUNT(*)
FROM walks
WHERE phone_screen_on_seconds < 30
  AND duration_minutes BETWEEN 20 AND 180
  AND ended_at_ridge = true;
Enter fullscreen mode Exit fullscreen mode

Every line of code above serves that metric. If it doesn't, delete it.


7. Build Your Own

The stack is boring on purpose:

  • Language: Whatever compiles to a single binary (Rust, Go, Swift, Kotlin/Native)
  • Audio: Piper TTS (build-time) + Opus (runtime)
  • Sensors: Platform APIs directly—no wrappers, no abstractions
  • Story: JSON + pure functions
  • Tests: Hypothesis / proptest / quickcheck
  • Distribution: App stores + F-Droid / GitHub Releases

Start this weekend.

  1. Write 10 nodes in story_graph.json.
  2. Run build_audio.sh.
  3. Write the 200-line main loop.
  4. Walk.

The trail doesn't care about your tech stack. It only cares that you're on it.


Touch grass. The code will wait.

Top comments (0)