DEV Community

Kushal Medipally
Kushal Medipally

Posted on

JARVIS: Building a Voice-First Autonomous AI Workspace with Persistent Memory

JARVIS: Building a Voice-First Autonomous AI Workspace with Persistent Memory

For HackwithHyderabad 3.0, our team localhosters is building JARVIS, a voice-first autonomous AI workspace designed to turn high-level user intent into real technical outcomes.

Instead of treating an AI assistant as a chat interface that simply generates an answer, JARVIS is designed around a simple idea:

Tell it what you want done. Let it understand, execute, verify, and return the result.

From conversation to execution

JARVIS combines an agent core with a capability layer that lets it work with real software environments. Depending on the task, it can interact with files, terminals, browsers, Git/GitHub, HTTP services, and sandboxed environments.

The architecture separates responsibilities:

  • Agent Core — interprets objectives, plans work, chooses capabilities, and evaluates results.
  • Capability Layer — safely validates and executes tool calls through a central ToolRegistry.
  • Persistent Memory — recalls useful experiences, observations, preferences, and project knowledge.
  • Runtime / Interface — manages sessions, events, voice interaction, and user-visible execution results.

This separation lets the agent reason about what should happen without needing to know the internal implementation of every capability.

The part that makes JARVIS different: memory

The hackathon requires Hindsight to be central to the solution, so we are building JARVIS around persistent experiential memory rather than treating memory as a simple chat-history feature.

We integrate Hindsight as the experiential memory layer. JARVIS can record useful experiences and observations from previous execution, retrieve relevant experiences for future tasks, and use that information to make subsequent interactions more contextual.

We also use structured durable knowledge through an OKF-based knowledge layer for information that should remain explicitly organized, such as project context and user knowledge.

The intended learning loop is:

Interaction → Experience → Memory → Retrieval → Better future execution

For example, if JARVIS previously worked through a Python project setup, a later request to create another Python project can retrieve the earlier experience and its outcome instead of treating the request as completely new.

Safe autonomous tool execution

Autonomous execution also needs boundaries.

Our capability layer uses a central ToolRegistry for schema validation, permission checks, security policies, execution, and structured ToolResults.

The current implementation includes explicit permission boundaries for filesystem operations, terminal execution, Git/GitHub, HTTP, browser automation, and sandbox execution.

Security-sensitive operations such as private-network HTTP access are controlled through trusted runtime context rather than model-generated arguments. Terminal commands run with process isolation and timeout cleanup, while HTTP responses are streamed with size limits.

Why we are building it

The goal isn't to build another chatbot.

We want JARVIS to behave more like an AI workspace that can actually do things: understand an objective, use the right capabilities, learn from previous work, recover from failures, and present the completed result to the user.

Our team is currently working across the agent intelligence, persistent memory, capability/tool execution, and runtime/interface layers.

JARVIS is still under active development, but the direction is straightforward:

An assistant that doesn't just remember what you said — it remembers what happened, learns from it, and uses that experience when it acts next.

AI #AIAgents #AgenticAI #JARVIS

Top comments (0)