DEV Community

Huy Vu
Huy Vu

Posted on

ne-ne, what do you want to build: A Cloud Coding Agent So My Friend Can Go Touch Grass

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🀝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.

What I Built

My friend Renault wants to go for a walk while his coding agent works. He also wants to shut down his laptop.

That second part changes the problem.

He told me about wanting something that could keep a Claude Code session running independently of his laptop. It is part of why he wants a Mac mini so much: a computer that stays on while he goes somewhere else.

I kept thinking about that request. When you delegate work to an agent, it is reasonable to want to step away from the machine hosting it. But if that machine is your laptop, stepping away still means leaving it running somewhere.

So I built ne-ne (Japanese slang for "hey hey"), a mobile-friendly web interface that gives a coding agent a workspace on a cloud computer. Renault can describe an app from his phone, leave the agent working, then return to review the changes and decide whether to publish and deploy them.

The current implementation runs Backboard R-CLI with Gemma. The laptop is outside the execution path; the phone is the interface to work happening on the server.

Other conversations helped me see why this could matter beyond one walk. Anthony has ideas he wants to build but little time to work on them. Friends have described tools for splitting expenses, personal finances, making men talking about vulnerability, and making sense of red flags in dating experiences, etc. Those are different applications, with different responsibilities. The shared difficulty is getting from an idea to something usable, then knowing how to change it.

For this weekend, I chose Renault's request as the concrete starting point: can a person leave their laptop off and still move a software project forward?

Renault hasn't tried ne-ne yet. The results below come from my development tests and the deployed prototype.

Demo

Open ne-ne

The intended walkthrough is:

  1. Describe a small app in the phone browser.
  2. Follow the agent's progress as it creates files and runs commands remotely.
  3. Open View Changes when the build is ready.
  4. Approve the result to publish the code to GitHub.
  5. Choose Deploy, then open the application once Render reports it live.
  6. Ask for an update in the same conversation.

During development, I opened the hosted interface on an iPhone and asked for a simple restaurant voting website with three choices. While the agent was working, I shut down the MacBook completely. The task continued and the phone received the result.

That was the test I cared about most. The browser was observing the job; it wasn't keeping the coding process alive.

This is a small prototype with one coding worker. The interface currently uses shared backend access, so it is intended for a limited demonstration rather than an unrestricted service for many users.

A look at the interface

The phone screenshot shows a restaurant-voting prompt. The desktop screenshots show a heading-only update to an existing project.

ne-ne on a phone with a restaurant voting app prompt

Describe the app from the phone browser; execution happens on the cloud backend.

Approved revision published to GitHub and deployed on Render

The completed revision exposes its code and app links, with publication and deployment reported separately.

More screenshots: project history, conversation, and revision summary

ne-ne desktop home with a live project in recent projects

The home screen lists the existing project and its live app link.

Project conversation showing the requested heading update and completed build

The request and result stay together in the project conversation, with a composer for follow-up changes.

Revision summary showing one modified file and one added and removed line

The revision summary identifies the changed file before opening its patch.

Code

ne-ne backend

ne-ne builds and updates applications from a browser or phone while the user's laptop is off. This Express/TypeScript service runs Backboard R-CLI in disposable Docker containers, uses DigitalOcean Inference for gemma-4-31B-it, persists state in MongoDB, and streams progress over server-sent events (SSE). A person reviews and approves a build before the backend publishes it to GitHub. Deploy is a separate action that hosts the generated application on Render.

The companion frontend is nene-web. This guide covers backend development, a complete first deployment on a new Linux host, and operation of the existing VM.

Choose your setup

Goal Start here Requirements
Run the interface against the existing VM Frontend README Git, Node 22, npm, backend URL and privately supplied API token
Edit, compile and test backend code on your machine Local development Git, Node 22 and npm; no production credentials for unit tests
Run build β†’ approval
…

Both main repositories include setup guides checked against the current source and running infrastructure. They cover local development, configuration, deployment, and the limits of the prototype.

How I Built It

Move the work off the laptop

The frontend is a Next.js application hosted on Render. Its server routes proxy requests to an Express/TypeScript backend on a DigitalOcean VM, keeping the backend API token out of browser JavaScript.

The backend starts Backboard R-CLI inside a disposable Docker container. Backboard calls DigitalOcean Inference's gemma-4-31B-it model to read, write, edit, and execute code in the workspace. Progress comes back through server-sent events, with polling for status reconciliation.

High-level ne-ne workflow: phone interface, cloud backend, coding workspace, hosted Gemma inference, saved artifact, human review, GitHub publication, and Render deployment

The laptop is outside the execution path. Progress returns to the phone while coding runs in the cloud. The person reviews the saved changes, chooses Approve to publish, and then chooses Deploy. Follow-up builds start from the latest approved source and reuse the project's deployment service.

I deliberately started with a small host: one CPU and about 1 GB of RAM. The model runs through an external inference API. The VM hosts the coding tools, not the model weights.

On that small VM, the agent needs explicit resource limits. The same helper configures the coding container and the short-lived workspace helpers:

export function sandboxArgs(): string[] {
  return ["--read-only", "--memory=512m", "--memory-swap=640m", "--cpus=0.75",
    "--pids-limit=128", "--cap-drop=ALL", "--security-opt=no-new-privileges"];
}
Enter fullscreen mode Exit fullscreen mode

Source: src/runner.ts.

The workspace remains writable, and the container has no Docker socket. The model key is available inside it; GitHub, Render, and database credentials remain with the backend. Resource limits bound what a coding run can consume on the host.

Preserve a project across requests

A successful first build wasn't enough. Anthony's idea might take several rounds to get right, and Renault shouldn't have to start over every time he asks for a change.

MongoDB Atlas stores projects, runs, events, and locks. The backend saves the generated source as an artifact before removing the coding container. A follow-up starts from the project's latest approved successful artifact, including when the selected historical run failed.

Approved updates reuse the project's GitHub branch and Render service. The existing live application stays canonical until the new deployment reports success.

There are limits here. One worker is admitted globally, and a busy worker returns a conflict rather than adding a job to a queue. Database persistence also doesn't mean an arbitrarily killed coding process can always resume. Browser conversation history is currently local to that browser; it doesn't sync across devices.

Make approval a real boundary

I wanted the person asking for the app to remain involved in what gets shipped.

When execution succeeds, the backend saves the artifact and moves the run into a review state. This excerpt is the point where coding finishes:

transition(run, "waiting_for_approval");
run.artifactPath = artifact;
run.progress = "waiting_for_approval";
await this.store.patchRun(run.id, { status: run.status, artifactPath: artifact, progress: run.progress, updatedAt: run.updatedAt });
await this.store.releaseRun(run, true);
await this.emit(run.id, "task:progress", { stage: run.progress, message: "Build finished. Waiting for your approval." });
Enter fullscreen mode Exit fullscreen mode

Source: src/backend.ts.

The frontend makes the saved changes available for review. A separate approve() method publishes to GitHub after checking that this is the project's active revision and its artifact is available. Deploy is another action.

This also keeps generated code away from publication credentials during the coding run. The host performs the Git push after approval. An agent finishing its work is one event; a person choosing to publish it is another.

Saved revision diff showing the heading changed to Hello from ne-ne v3

The patch shows the exact heading change: one line removed and one added in app/index.js.

Inspect what the agent actually did

A progress label saying β€œTesting” is easy to display. Knowing whether a command passed takes more care.

I added Sentry Agent Tracing around workspace preparation, agent execution, model observations, tool calls, artifact capture, approval, GitHub publication, and Render deployment. Approval can happen much later, so it has its own trace correlated by project and run IDs.

Two details stood out while inspecting Backboard's source:

  • A returned tool event doesn't automatically prove a passing test. The observer recognizes controlled success/failure signals and leaves missing outcomes unknown.
  • Usage from the final model completion isn't forwarded through the same event stream as some earlier usage. The integration reports observed tokens and a labeled cost estimate, which can undercount billed usage.

Model span durations are observed intervals, too; they aren't exact provider-only latency measurements. These distinctions make the tracing more useful because I know what a number represents.

Sentry Agent Activity showing the synthetic ne-ne workflow

This is real Sentry UI displaying a synthetic fixture. Its token counts, costs, and timings are test values, not measurements of a live Gemma build. It verifies that the agent, model, and tool spans are recognized.

The observer bounds detailed tool and model spans while continuing aggregate counts. Prompts, generated code, and raw tool output are excluded from these trace payloads.

Verified production HTTP trace from the deployed backend

The production screenshot verifies tracing from the running backend. Its 109.20 ms request was a read-only lookup returning an expected 404, not an agent performance benchmark.

For verification, the local backend build and 30 tests passed, with the optional real MongoDB transaction test skipped locally. Frontend lint, 13 tests, and the production webpack build passed. The separate tracing evidence documents VM transaction testing and synthetic workflow coverage. I also checked the deployed configuration against the repositories so this write-up describes what is running.

What does it cost to keep the agent working?

Renault's Mac mini idea also raised a cost question: could I give the agent an always-on workspace without buying another computer?

The prototype separates the cost of keeping the worker available from the cost of model calls. These are USD list prices checked on October 5, 2026, before tax or promotional credits:

Component Cost basis
Backend VM and coding workspace $6/month for the matching Basic Droplet: 1 vCPU, 1 GiB RAM, 25 GiB SSD. Docker containers share that host. DigitalOcean pricing.
Gemma 4 inference $0.18 per million uncached input tokens, $0.50 per million output tokens, and $0.036 per million cached input tokens. Inference pricing.
Generated app hosting The deployment code requests Render's Free compute plan. Free services share 750 running instance hours per workspace per month and spin down after 15 minutes idle; bandwidth and build allowances also apply. Render limits.
Frontend, database, and monitoring Render frontend hosting, MongoDB Atlas, and Sentry must be budgeted according to their actual account plans and usage. Atlas offers an M0 free tier, but I haven't verified all account tiers or invoices for this analysis. Atlas pricing.

The inference estimate treats cached tokens as a subset of input:

USD = ((input - cached) Γ— 0.18
       + cached Γ— 0.036
       + output Γ— 0.50) / 1,000,000
Enter fullscreen mode Exit fullscreen mode

For a hypothetical run totaling 20,000 input tokens and 5,000 output tokens across all model calls, with no cache hits, inference would cost $0.0061. At that same usage, 100 runs would cost $0.61 in inference, or 1,000 runs $6.10, in addition to infrastructure and other services. These are arithmetic examples, not measured average build costs; retries and repeated context count too.

The practical tradeoff is a small recurring VM cost and usage-based inference, with one worker and limited hosting capacity. The $6 figure is the VM component, not the complete system bill. Failed runs can still consume tokens, and container CPU/memory limits don't impose a token-spend cap. Sentry's estimate also uses incomplete observed usage, so provider billing remains the source of truth.

Why Does Open Innovation Matter?

Backboard R-CLI is the open-source agent harness at the core of ne-ne, and Gemma is the open-weight model doing the coding reasoning. ne-ne builds on that project's existing tools and execution loop.

That openness affected the work directly. I could build a pinned Backboard version from source, configure its custom provider to use DigitalOcean's inference endpoint, and inspect how it emitted tool results and token usage. The tracing limitations above came from reading that implementation. They would have been much harder to understand from progress messages alone.

The provider configuration is explicit in the repository:

{
  "provider": "digitalocean",
  "model": "gemma-4-31B-it"
}
Enter fullscreen mode Exit fullscreen mode

The model and harness are meaningful parts of the execution path, and the integration is something I can inspect and change as the project develops.

Another friend of mine - Maesha - requested to give that future work a specific direction. She wants a local LLM for security reasons. However, the current ne-ne deployment doesn't meet that requirement: (Sorry, I could not fulfil that request of yours πŸ₯²) it sends inference requests to a hosted provider. Hosting the backend myself doesn't make those requests local, and excluding prompts from tracing doesn't change where inference happens.

A local inference option would need appropriate hardware, configuration, and verification of the data path. The open pieces give me a practical starting point, but I still have to do that work.

The next test is simpler: give Renault the prototype and see whether he can start a task, leave it working, and comfortably review the result when he returns. I built the infrastructure around the walk he wanted to take. Now I need to find out whether the experience works for him.

Prize Categories

I am entering the categories supported by the current implementation:

  • Best Use of Render β€” hosts the phone interface and deploys approved generated applications, with subsequent updates using the same service.
  • Best Use of DigitalOcean β€” hosts the backend and Docker coding workspace, and provides the inference endpoint.
  • Best Use of Gemma β€” gemma-4-31B-it is the coding model used through hosted inference.
  • Best Use of Backboard β€” R-CLI provides the agent's coding and tool execution loop inside the workspace.
  • Best Use of MongoDB Atlas β€” persists project/run state, events, and transactional locks across requests.
  • Best Use of Sentry Agent Tracing β€” instruments the workflow, tools, observed model usage, and labeled cost estimates; the screenshots above distinguish synthetic agent evidence from production HTTP evidence.

Top comments (0)