DEV Community

Cover image for Your Code for $0.10/M Comes With a Tax: Your Training Data
xxxn3m3s1sxxx
xxxn3m3s1sxxx

Posted on AI-assisted

Your Code for $0.10/M Comes With a Tax: Your Training Data

Meta's Contributor tier charges about $0.10 per million input tokens — up to 20x less than the standard API. The terms include a line you should read twice: your prompts and completions may be used to improve Meta's models.

The Case

We run a self-hosted swarm where agents spawn background research jobs. Cheapest provider wins, right? A free tier would drop our inference costs to zero — and every one of those spawned sessions would feed somebody's training run.

What Actually Happened

We stopped asking "can we afford this?" and started asking "what are we paying with?":

  • Cost discipline without a gate silently routes onto the free-tier default — the decision disappears into the config instead of into the review.
  • Every background spawn becomes potential training material: prompts, completions, function calls, the codebase itself.
  • A discount on tokens is not a discount on the currency you actually pay in.
  • Instead of paying silently: SWARM_GATE_FREE — a gate inside the swarm itself.

Proof (Swarm Finding)

Real code, 2026-09-18, our swarm. resolve_spawn_model() (opencode_run.py, lines 114–149) implements the free-tier gate route: SWARM_GATE_FREE=1 blocks silent free-tier default spawns — background research runs through local Ollama or BYOK routes, and an invalid target spawn is rejected with a clear error instead of starting silently. Tested in tests/test_opencode_run.py (SWARM_GATE_FREE cases); consensus documented in team-constitution.md § Free-Tier-Provider-Gate. See the full system story in the Pillar post.

The Lesson

  • Every price is a trade. Read the currency, not the zeros behind the dollar sign: $0.10/M is a training-data price, not an infrastructure price.
  • Saving money means knowing what you pay with. If your agent generates code on a contributor discount, your codebase is the input good.
  • Gates over wishes: deterministic blocking (error instead of spawn) is cheaper than "we'll keep an eye on it" in every single spawn.
  • Decide once, enforce everywhere — an env flag plus tests beats a rule label in the config jungle.

The Question

What's the cheapest inference tier you run in production — and what are you actually paying with?

Video

More about our swarm, the crashes and the fixes: YouTube

Built by the ERR.SYS / 0xRAGE404 team — see the full system in the Pillar post.

Top comments (0)

Some comments have been hidden by the post's author - find out more