DEV Community

Shivam Kumar
Shivam Kumar

Posted on

I Built a 188M Mixture-of-Experts LLM From Scratch on a Free GPU

By Shivam Kumar, founder of VisionQuantech. This is the honest version — what's proven, what's measured, and what's still running.

Why a tiny MoE?

Most mixture-of-experts research happens at billion-parameter scale. DeepSeekMoE, Mixtral, GLaM — all brilliant, all far beyond what a single free Colab GPU can touch. But the core idea of MoE is beautifully simple: only a fraction of the model needs to be awake for each token. A dense 188M model fires all 188M parameters on every word. A sparse one can carry 188M worth of knowledge while only spending ~51M of compute per token.

That's the bet I made: get big-model capacity at small-model cost, on a free Tesla T4, for $0. Not to beat billion-parameter models — that would be absurd — but to prove you can design a credible tiny MoE architecture principally, not by vibes.

The research system I used

I have a methodology I call the Main Researcher System v4 — recursive pattern-combination for discovering new algorithms. Instead of hand-tuning an architecture by intuition, I ran the process:

  1. Pattern extraction. I pulled 12 MoE primitives from the literature (fine-grained segmentation, shared experts, token-choice vs expert-choice routing, aux-loss balancing, router z-loss, SwiGLU experts, dropless training...) and 11 meta-patterns (TRIZ, morphological analysis, MAP-Elites, DreamCoder abstraction, and more).

  2. TRIZ contradiction resolution. The central contradiction: more experts = more capacity BUT more routing cost, imbalance risk, and data dilution; tiny model BUT wants excellence in every category. The resolution was structural, not a compromise: Segmentation (split coarse experts into 64 micro-experts — capacity without extra per-token compute), Taking Out (extract common knowledge into a shared expert), Local Quality (specialization emerges from the router, never hard-coded).

  3. Combinatorial engine. A morphological Zwicky box over 8 parameters screened out inconsistent combos (expert-choice routing leaks causally in decoder LMs — out; top-1 with no aux loss — guaranteed collapse). Documented GP/SCAMPER operators then generated 5 candidates: 3 sourced-track (DeepSeekMoE-style, Switch-style, Mixtral-style), 1 bio-inspired (a fatigue-homeostasis balancer with no aux loss), and 1 moonshot (a domain-hint gate with dropout).

  4. Dual evaluation. Sourced candidates were scored on Function ÷ active-parameters using published evidence. The moonshot had to pass five falsifiable logical checks. Then all four tested candidates ran real 300-step CPU experiments measuring loss decrease, dead experts, and routing balance.

The winner: DeepSeekMoE-tiny

  • 188,269,568 total params, 51,430,400 active per token — the compute of a ~51M dense model, ~3.7× its capacity
  • 8 layers, d_model 512, 8 heads, context 1024
  • 64 routed SwiGLU micro-experts (dim 192) + 1 always-on shared expert per layer
  • Token-choice top-6 routing (renormalized), expert-level aux loss (α=0.01) + router z-loss (1e-3), dropless, fp16

Why not 9 domain experts — one per field? I tried that design first and rejected it. The model has to be best across every category, not 9 buckets. DeepSeekMoE's combinatorics (~75M routing combinations per layer at identical FLOPs) let 64 micro-experts cover unlimited fine patterns — syntax, facts, reasoning styles — instead of 9 coarse buckets. Specialization emerges; it's never assigned.

The real numbers

Sourced value scores (Function ÷ active-params): A 0.175 > B 0.161 ≈ D 0.161 > E 0.156 > C 0.097. The Mixtral-style coarse design was worst — most active compute for the coarsest routing.

Empirical (300 CPU steps each, real runs): all four candidates trained stably (loss 41.5 → 7.6, no NaN, no collapse). The winner balanced best — aux loss 2.03, zero dead experts. The bio-inspired balancer almost worked but left 1 expert starving (utilization 0.002) — gated out by the hard robustness rule: no babysitting, no starving experts. The moonshot degraded gracefully but didn't beat the winner, so it's archived, not selected.

And here's the honest part: cluster↔expert mutual information was ~0 for all candidates. 300 toy steps cannot measure specialization. The experiments prove routing balance and stability, not quality. Quality evidence stays sourced from the literature. I claim no more than that.

The candidates that didn't win (and why they matter)

The archive keeps the losers, because losers are reusable parts. The Switch-style design (top-1 routing, no shared expert) was the efficiency champion on paper — only 37.3M active per token — but top-1 routing at tiny scale starves the router of gradient signal, and it lost on sourced value. It stays as the efficiency-cell elite.

The bio-inspired design is my favorite failure. Instead of an auxiliary loss term fighting the language-modeling objective, each expert tracks a "fatigue" exponential moving average of its own usage; routing scores get penalized by fatigue, like neural homeostasis — or ants laying pheromones. It almost balanced perfectly with zero loss-term overhead. Almost. One expert starved at 0.002 utilization, and my hard rule is no starving experts, no babysitting. So it's gated out as the primary — but the mechanism is viable, and it's the obvious crossover partner for the next generation.

The moonshot — a domain-hint gate with 50% hint dropout — passed all five logical consistency checks, including graceful degradation (no-hint inference was actually slightly better than with-hint, gap −0.061). It just didn't beat the winner on value. Archived, not deleted. That's the whole point of the moonshot track: novelty is allowed to win, but only if it earns it.

Every known failure mode has a guardrail

Inversion — "how would I guarantee this fails?" — produced 8 guardrails baked into the trainer: routing-collapse monitor (loud warning if >25% experts go near-dead), aux-loss dominance warning, auto-halve micro-batch on OOM, NaN abort with diagnostics, checkpoints every 1,000 steps with --resume for Colab preemption, and more.

What's next

Training is running right now on a free Colab T4: Phase 1 on TinyStories (~200M tokens) for fluent language, then Phase 2 on mixed instruction data (~15M tokens, including my Shivacon domain sets and Dolly-15k). Total ~215M tokens, ~2.5–3 hours, $0.

When it finishes, the real test begins: does the full-scale model utilize its experts well, and does it beat a dense-51M baseline? I'll publish those numbers whatever they say — including if they're bad.

The full research trace, architecture docs, model, and trainer are in the project repo. This is what $0 and a methodology can buy you: not a miracle, but a design you can defend.


Shivam Kumar is an AI researcher and founder of VisionQuantech, working on AGI/ASI research funded by an AI services business. All 9 Shivacon domain adapters and this MoE model are built on free infrastructure.

Top comments (0)