DEV Community

Cover image for We Built a 3D Game Where Your Voice Breaks Reality, As Our Hacker House Submission
spideyy
spideyy

Posted on AI-assisted

We Built a 3D Game Where Your Voice Breaks Reality, As Our Hacker House Submission

"How much of the world are you willing to destroy to create the world you want?"

Our team — Alt F4 — built this as our submission for Hacker House Goa 2026 in collaboration with Wispr Flow. And we decided to break every rule we set for ourselves.

No game engine. No Unity. No Godot. Just React, TypeScript, Three.js, and a slightly deranged idea:
What if your voice was the controller?

This is the story of ECHO — a cel-shaded dark fantasy RPG where you speak commands into the browser and the world obeys. Soldiers retreat. Bridges collapse. Rain drowns out torches. And every edit to reality leaves a scar.


The Premise (and Why It's Dangerous)

The player wakes up in the Meadowlands with no memory — but they have a power called The Echo. Whatever they speak aloud manifests in the physical world.

Say "Call the rain" → storm clouds roll in, torches extinguish, raiders break rank and flee.
Say "Destroy the bridge" → the stone bridge shatters into rubble, cutting off the enemy vanguard.
Say "Protect Rowan" → a shimmering Chronal Barrier envelops your miller friend, making him immune to blades.

Cool, right? Here's the catch:

Rewriting reality doesn't erase what happened. It fractures the world into alternate branches.

Every command leaves a silent scar. Mira the Seer carries the agonizing memory of every erased timeline. The ancient Architect tried to build a world without sorrow — and cracked the foundations of reality. You're next.


🧠 The Architecture That Makes It Tick

Here's a bird's eye view of the pipeline:

       Microphone Input
             │
             ▼
   Web Speech API (streaming)
             │
             ▼
   CommandParser
   ├── Local keyword fallback (zero-latency, offline)
   └── Gemini Flash API (structured JSON, natural language)
             │
             ▼
   Command Dispatcher
   ├── WorldState (entities, weather, physics)
   ├── CampaignSystem (missions, objectives, dialogue)
   └── TimelineSystem (snapshots, branches, rewind engine)
Enter fullscreen mode Exit fullscreen mode

No bloated middleware. No event bus spaghetti. Each layer is a Zustand store, and commands flow downward like water.


🎙️ The Voice Engine: Two Modes, One Experience

The hardest design constraint I imposed on myself: the game must feel instant.

If Gemini is slow, you can't feel it. So I built a two-layer system:

Layer 1 — Zero-Latency Local Fallback

Mission-critical commands (retreat, shield, rain, rewind) are matched locally via keyword parsing. No network call. No latency. Fires in the same animation frame as the mic input.

Layer 2 — Gemini Structured JSON

For nuanced natural language ("Tell the soldiers of Shadowfang to go back where they came from") — Gemini parses it and returns a clean structured command object:

{
  type: "INTERACT_ENTITY",
  entityName: "soldiers",
  action: "retreat"
}
Enter fullscreen mode Exit fullscreen mode

The dispatcher doesn't care how it was parsed. It just runs.

The real magic is the fallback ordering: try local first, call Gemini only when you need it. Zero perceived latency, full natural language support.


🌳 The Echo Tree — One Physical Anchor for All of Time

This was the wildest design decision.

Timeline manipulation in most games is a menu. In ECHO, it's a place. The Echo Tree is a singular location in the world at coordinates (-13, -1.5) — a glade west of the forest. You physically walk to it. You stand in its radius. You interact with it.

export const ECHO_TREE_COORDS = {
  x: -13.0,
  z: -1.5,
  interactionRadius: 4.2,
} as const;

checkProximity: (pos) => {
  const dist = Math.hypot(pos.x - ECHO_TREE_COORDS.x, pos.z - ECHO_TREE_COORDS.z);
  const isNear = dist <= ECHO_TREE_COORDS.interactionRadius;
  set({ isNear });
  return isNear;
},
Enter fullscreen mode Exit fullscreen mode

The tree maintains deep immutable world snapshots — every entity state, weather condition, chaos score, and dynamic prop. You can fork realities, jump between branches, or rewind to narrative junctions. And critically: control configurations persist through timeline rewrites. Your key bindings survive the apocalypse.


⚔️ Mission 3: Where Simulation Gets Real

Mission 3 is the first major test of the system. Warlord Vorn's Shadowfang Vanguard is marching across the bridge toward Rowan's Mill with lit torches.

The player has 5 solutions, all of them real physics events:

Voice Command What Actually Happens
"Protect Rowan" Chronal barrier activates. Entity gains aiState: 'cheering', full HP restore
"Call the rain" Weather flips to storm. Torch entities extinguished. Raiders enter fleeing state
"Destroy the bridge" Bridge mesh collapses. setBridgeDestroyed(true). Raiders halt or reroute
"Make the soldiers retreat" All Shadowfang entities get aiState: 'fleeing', despawn after 3.5s
Direct combat Melee/spell physics, knockback, stagger, health depletion

Rowan's fate is never faked through dialogue. It derives directly from his simulated health state at the end of the mission. Die at 0 HP? He collapses. Survive at 59 HP? He bandages his wounds. The game acknowledges the struggle. No cutscene tricks.

And here's what the interaction code actually looks like when a player shields Rowan:

store.updateEntity(shieldTarget.id, {
  health: shieldTarget.maxHealth,
  dialogBark: 'The air around me... it holds! The strikes cannot touch me!',
  aiState: 'cheering',
});

record({
  type: 'player_helped',
  actorId: 'player',
  actorName: 'The Voice',
  targetId: shieldTarget.id,
  description: `The Voice manifested an impenetrable chronal shield around ${shieldTarget.name}.`,
  significance: 3,
});

TimelineSystem.createCheckpoint({
  name: `Shielded ${shieldTarget.name}`,
  significance: 'story',
});
Enter fullscreen mode Exit fullscreen mode

One action. Three systems updated. Memory ledger written. Timeline checkpoint created. State broadcasted.


🔊 Audio: Zero External Assets, 100% Procedural

This was a point of pride. ECHO has zero audio files. Every sound is synthesized at runtime using the Web Audio API.

River Ambience

A continuous 6-second pink noise buffer runs through parallel dual-band filters:

  • Deep flow path → Lowpass at 420 Hz (the body of rushing water)
  • Surface trickle path → Bandpass at 920 Hz, Q=1.1 (water rippling over stones)
  • Distance attenuation → Logarithmic falloff based on player proximity to the riverbank

Shoreline Hysteresis

This one's my favorite detail. The water entry/exit logic uses a deadband to prevent oscillation at the shoreline:

  • Entry triggers at ≥ 0.08m depth
  • Exit triggers only at < 0.03m depth (6cm deadband)
  • 450ms debounce timer prevents rapid re-triggering

No more "splash splash splash" when walking along the water's edge.

Underwater Acoustics

A master lowpass filter ramps down to 320 Hz to muffle the world. A 52 Hz sub-aquatic pressure drone hums while submerged. Surfacing restores the full 22 kHz spectrum with a crisp air-breach emergence splash.

All of it synthesized. All of it reactive. No audio files.


🎨 Cel-Shading in React Three Fiber

The visual direction is stylized dark fantasy cel-shading. Custom toon color ramps (getToonGradient3, getToonGradient4) applied to procedurally textured meshes: stone masonry, timber grain, heraldic tabards, character attire.

The UI follows a cinematic dark fantasy palette across all 18 UI surfaces:

  • Typography: Warm ivory/parchment #E8E3D8
  • Highlights: Muted antique gold #B59A4A / #8F7836
  • Panels: Near-black charcoal #070B12 / #0C1119
  • Borders: Dark bronze/stone #292923 / #3A3628

Cohesion across Title Screen, Dialogue Box, Memory Panel, Timeline Tree, Save/Load modal — all 18 surfaces. When a UI element appears, it belongs in this world.


🏗️ The Full Stack

Layer Technologies
Frontend Core React 19, TypeScript, Vite
3D Rendering Three.js, React Three Fiber, Drei
State Management Zustand (modular: WorldState, CampaignSystem, TimelineSystem, EchoTreeState, SettingsStore)
Speech & AI Web Speech Recognition API + Google Gemini Flash (structured JSON)
Audio Procedural Web Audio API
Styling Vanilla CSS, Glassmorphic overlays
Testing Node.js + tsx TypeScript test runner

🧪 Actually Testing a Narrative Game

Most game test suites are a lie. Unit tests for "did the menu open?" don't catch whether your story systems stay coherent across timeline rewrites.

I wrote four test suites that actually matter:

# Full narrative vertical slice (M1 → M2 → M3 → M4, branches, Echo Tree)
npx tsx scripts/verify_narrative_slice.ts

# Controls & key rebinding (bindings, conflict resolution, timeline isolation)
npx tsx scripts/verify_controls_system.ts

# Water audio & acoustic physics (proximity falloff, hysteresis, splash synthesis)
npx tsx scripts/verify_water_audio.ts

# Mission 3 battle simulation (all 4 Echo solutions, raider pathing, Rowan fates)
npx tsx scripts/verify_m3_battle.ts
Enter fullscreen mode Exit fullscreen mode

Zero regressions across storyline logic, audio synthesis, and world simulation. That's the bar.



🤝 Alt F4 × Wispr Flow — The Collaboration

ECHO is Team Alt F4's submission for Hacker House Goa 2026, built in collaboration with Wispr Flow.

Wispr Flow is a voice-first input layer that makes speaking to software feel natural and frictionless — which made them the perfect partner for a game where your voice is the only weapon that matters. Their philosophy of making voice a first-class input shaped how we thought about the Echo's command pipeline: if it feels clunky to say, it feels clunky to play.

Building ECHO as a selection submission meant every architectural decision had to be a fast, confident bet:

  • Zustand over Redux (less boilerplate, same power)
  • Web Audio API over audio files (zero asset pipeline, zero loading time)
  • Local keyword fallback over pure LLM (playable even when the API is slow)
  • Simulation over scripting (Rowan's fate is real, not cinematic)

The pressure of a tight deadline forced exactly the kind of ruthless simplicity that made the systems clean.


💭 What We Learned

1. Simulation over scripting. The temptation to fake outcomes with dialogue is real. Resist it. When Rowan's fate comes from his actual HP value in the simulation, the game means something. Players feel it.

2. Voice input has two enemies: latency and ambiguity. Solve latency with local fallback. Solve ambiguity with LLMs. Use both.

3. Procedural audio is worth the complexity. The river doesn't loop awkwardly, the underwater acoustics feel real, and there are zero licensing concerns. The investment pays off.

4. The world should react, not the UI. No giant overlays when you speak a command. The rain falls. The bridge breaks. The soldiers run. The world tells you what happened.

5. React + Three.js is totally viable for games. The component model for scene graph management is actually great. Zustand handles game state better than any bespoke solution I've tried. Don't sleep on this stack.

6. Deadline pressure is a feature, not a bug. Having a hard deadline forced us to bet fast and ship. Every cut we made improved the core. If you have a wild idea, go find a challenge that forces you to build it.


🔗 Try It / Explore It

The repo is public — go break things:

👉 github.com/kishore1035/echo

git clone https://github.com/kishore1035/echo.git
cd echo
npm install
npm run dev
# Open http://localhost:5173 and start speaking
Enter fullscreen mode Exit fullscreen mode

Optionally drop a Gemini API key in .env for full natural language parsing — the game is fully playable without it using the local keyword fallback.

If you're building something at the intersection of web tech and games — or you're interested in what we're working on — hit us up. This space is criminally underexplored and the tools are better than you think.


Team Alt F4 · Submission for Hacker House Goa 2026 · In collaboration with Wispr Flow

React 19 · TypeScript · Three.js · Google Gemini · Web Audio API · and an embarrassing amount of late nights.

Top comments (0)