<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Shivam Kumar</title>
    <description>The latest articles on DEV Community by Shivam Kumar (@shivaysinghrajput).</description>
    <link>https://dev.to/shivaysinghrajput</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167710%2F5ea88fa2-d58b-4c0d-87dc-acc9ef297194.jpg</url>
      <title>DEV Community: Shivam Kumar</title>
      <link>https://dev.to/shivaysinghrajput</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shivaysinghrajput"/>
    <language>en</language>
    <item>
      <title>NEXUS: Why the Next AI Architecture Won't Be Just One Thing</title>
      <dc:creator>Shivam Kumar</dc:creator>
      <pubDate>Wed, 07 Oct 2026 05:03:13 +0000</pubDate>
      <link>https://dev.to/shivaysinghrajput/nexus-why-the-next-ai-architecture-wont-be-just-one-thing-16i</link>
      <guid>https://dev.to/shivaysinghrajput/nexus-why-the-next-ai-architecture-wont-be-just-one-thing-16i</guid>
      <description>&lt;p&gt;After watching enterprise AI deployments for a while, I've come to a conclusion that shapes everything I build: the systems that win in 2026–2030 won't be "just RAG" or "just agents" or "just fine-tuned." They'll be unified systems that compose all of these under one orchestration layer — with self-improvement loops built in.&lt;/p&gt;

&lt;p&gt;I call the template NEXUS: Neural EXecution &amp;amp; Understanding System. Here's the honest version of what it is and isn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The seven layers
&lt;/h2&gt;

&lt;p&gt;NEXUS is a 7-layer design where each layer has a defined contract:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Intent Decomposition&lt;/strong&gt; — parses a request into a typed task graph, not a free-text prompt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Adaptive Temporal Memory&lt;/strong&gt; — three knowledge tiers (hot/parametric, warm/vector, cold/archival) with relevance decay.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-Specialist Agent Pool&lt;/strong&gt; — a registry of typed agents dispatched by an orchestrator that learns which agents perform best per task class.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Neurosymbolic Verification&lt;/strong&gt; — every output passes a cascade: domain rule engine, a small fine-tuned verifier model, and a contradiction detector. Outputs carry a "verification passport."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contextual State Bus&lt;/strong&gt; — a shared typed event stream; this is what kills the "lost in the middle" problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Continual Self-Distillation&lt;/strong&gt; — verified high-quality outputs become training triplets; the system gets smarter from its own production runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Adaptive Output Formatter&lt;/strong&gt; — format separated from generation, so JSON compliance issues never corrupt the reasoning loop.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;This is an architecture white paper, not an empirical result. I haven't built the full system or run the benchmarks. What I am confident about: the component choices are sound, the layer contracts are well-defined, and the gap it addresses — no verification, no freshness, no self-improvement in current production systems — is real.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shivam Kumar is an AI researcher and founder of VisionQuantech, working on unified AI architectures and recursive methods for algorithm discovery.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>architecture</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>What Five Cancer-Fighting Technologies Taught Me About Invention</title>
      <dc:creator>Shivam Kumar</dc:creator>
      <pubDate>Wed, 07 Oct 2026 04:59:25 +0000</pubDate>
      <link>https://dev.to/shivaysinghrajput/what-five-cancer-fighting-technologies-taught-me-about-invention-3803</link>
      <guid>https://dev.to/shivaysinghrajput/what-five-cancer-fighting-technologies-taught-me-about-invention-3803</guid>
      <description>&lt;p&gt;I spent the last few weeks doing something unusual: instead of studying one cancer technology deeply, I mapped five of them side by side — radioisotope theranostics, bacteriophages, DNA-origami nanorobots, bacteria microrobots, and macrophage engineering — and asked what engineering logic they had independently converged on. The result surprised me enough that I wanted to share it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern nobody talks about
&lt;/h2&gt;

&lt;p&gt;Every one of these fields is trying to solve the same problem: chemotherapy hits the whole body, so how do you make the killing &lt;em&gt;conditional on location&lt;/em&gt;? And independently, three unrelated labs reinvented the same answer — pH-gated activation. Tumors sit around pH 6.5 while healthy tissue is 7.4, and that tiny chemical difference became a trigger.&lt;/p&gt;

&lt;p&gt;When three fields that never talk to each other invent the same trick, that's not coincidence. That's a validated pattern.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where TRIZ comes in
&lt;/h2&gt;

&lt;p&gt;Running these through TRIZ contradiction analysis, the pairings got interesting. The big one: &lt;strong&gt;potency vs. precision&lt;/strong&gt;. That led me to two original combinations that, as far as I could find, nobody is pursuing:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RadioPhage Origami Capsule.&lt;/strong&gt; A magnetotactic bacterium carrying a DNA-origami capsule chelated to a radioisotope. The capsule stays sealed until tumor pH unfolds it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CD47-Blocking PhageCloak.&lt;/strong&gt; One phage capsid displaying a tumor-receptor ligand &lt;em&gt;and&lt;/em&gt; a peptide that blocks the cancer cell's "don't-eat-me" CD47 signal, so the patient's own macrophages do the killing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;Both ideas are conceptual — TRL 1 or 2. No experiments, no data, and I want to be explicit about that. This is ideation, not a lab protocol. But the method — extract patterns from sourced literature, run them through contradiction analysis, recombine — is a machine for generating &lt;em&gt;inventable&lt;/em&gt; ideas.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shivam Kumar is an AI researcher and founder of VisionQuantech, working on recursive pattern-combination methods for algorithm and architecture discovery.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>science</category>
      <category>ai</category>
      <category>discuss</category>
    </item>
    <item>
      <title>U-STACK: What If Apps Were Guests and Users Owned Everything?</title>
      <dc:creator>Shivam Kumar</dc:creator>
      <pubDate>Wed, 07 Oct 2026 04:57:41 +0000</pubDate>
      <link>https://dev.to/shivaysinghrajput/u-stack-what-if-apps-were-guests-and-users-owned-everything-1k0f</link>
      <guid>https://dev.to/shivaysinghrajput/u-stack-what-if-apps-were-guests-and-users-owned-everything-1k0f</guid>
      <description>&lt;p&gt;Every app you install re-solves the same four problems: who you are, how money moves, where your data lives, and who can access what. Your identity in one app has nothing to do with your identity in another. Your data lives inside each app's walls. And every AI agent acting on your behalf today does it by borrowing your credentials — which is both insecure and unauditable.&lt;/p&gt;

&lt;p&gt;I wrote a working paper formalizing an alternative. It's called U-STACK, and its thesis is simple: &lt;strong&gt;identity, value, data, and permission should be user-side infrastructure, and applications should be guests operating under explicit, revocable, auditable grants.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The seven layers
&lt;/h2&gt;

&lt;p&gt;The stack has seven layers, each depending only on the ones below it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;U-ID&lt;/strong&gt; — one portable identity for humans, organizations, devices, and AI agents. Agents are first-class citizens with a declared principal — never credential-borrowing hacks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;U-WALLET&lt;/strong&gt; — one value layer per identity: ledgers, escrow, subscriptions, royalty splits, with compliance as pluggable policy rather than an afterthought.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;U-VAULT&lt;/strong&gt; — your data, encrypted, versioned, living with you. Apps hold grants, never ownership. Revoking a grant provably ends access.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;U-PERM&lt;/strong&gt; — the consent engine: fine-grained, time-bound, legible permissions with an append-only audit log. Default-deny everywhere; delegation can only narrow, never widen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;U-WORK&lt;/strong&gt; — an event-driven workflow OS. Approvals are permission grants, so separation of duties is enforced by architecture, not convention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;U-DEV&lt;/strong&gt; — the developer platform: app registry, serverless runtime, and AI-agent deployment with declared capability bounds. Revenue settles automatically from the ledger.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;U-SECTOR&lt;/strong&gt; — pre-wired industry bundles (finance, healthcare, education, government, commerce, logistics, AI economy) that collapse adoption cost for entire verticals.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The composition argument, compressed: layers 1–4 are &lt;strong&gt;control&lt;/strong&gt; (whoever defines the primitives defines the game), layer 5 is &lt;strong&gt;lock-in&lt;/strong&gt; (you can export data, but not operations), layer 6 is &lt;strong&gt;scale&lt;/strong&gt; (two-sided network effects), layer 7 is &lt;strong&gt;sector capture&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The closest thing that already works
&lt;/h2&gt;

&lt;p&gt;I'm not claiming this is unprecedented. India's digital public infrastructure is the existence proof: Aadhaar for identity, UPI for real-time payments, DigiLocker for documents, ONDC for open commerce, and the Account Aggregator framework for consent-based data sharing — which is a direct ancestor of U-PERM's consent engine. U-STACK can be read as DPI generalized: the same layered, protocol-first philosophy, extended to all data, all permissions, workflows, and developers — with AI agents as first-class citizens, which no DPI stack yet provides.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why agents change everything
&lt;/h2&gt;

&lt;p&gt;There's a reason this architecture puts AI agents at the center. Right now, an agent that books your flights or manages your calendar does it by holding your passwords — a security model that collapses the moment agents become capable enough to matter. U-STACK gives every agent its own identity with a declared principal, its own wallet with spending bounds, and permissions scoped to exactly what it needs. When the agent acts, the audit log shows which agent, whose authority, and what grant — not a shared login.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;This is a specification, not an implementation. No layer has running code. No wire formats, no security proofs, no performance measurements. The paper says this explicitly in a section I call the honesty ledger, because a target architecture should never be mistaken for shipped software.&lt;/p&gt;

&lt;p&gt;What I am confident about: the layering is right, the invariants are the right ones (default-deny, attenuation-only delegation, auditability by construction), and the direction — users owning primitives, apps as guests, agents as citizens — is where the puck is going.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Shivam Kumar is an AI researcher and founder of VisionQuantech. U-STACK is working paper #2 in the VisionQuantech architecture series.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>ai</category>
      <category>security</category>
      <category>privacy</category>
    </item>
    <item>
      <title>I Built a 188M Mixture-of-Experts LLM From Scratch on a Free GPU</title>
      <dc:creator>Shivam Kumar</dc:creator>
      <pubDate>Wed, 07 Oct 2026 04:52:54 +0000</pubDate>
      <link>https://dev.to/shivaysinghrajput/i-built-a-188m-mixture-of-experts-llm-from-scratch-on-a-free-gpu-5b30</link>
      <guid>https://dev.to/shivaysinghrajput/i-built-a-188m-mixture-of-experts-llm-from-scratch-on-a-free-gpu-5b30</guid>
      <description>&lt;p&gt;&lt;em&gt;By Shivam Kumar, founder of VisionQuantech. This is the honest version — what's proven, what's measured, and what's still running.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a tiny MoE?
&lt;/h2&gt;

&lt;p&gt;Most mixture-of-experts research happens at billion-parameter scale. DeepSeekMoE, Mixtral, GLaM — all brilliant, all far beyond what a single free Colab GPU can touch. But the core idea of MoE is beautifully simple: only a fraction of the model needs to be awake for each token. A dense 188M model fires all 188M parameters on every word. A sparse one can carry 188M worth of knowledge while only spending ~51M of compute per token.&lt;/p&gt;

&lt;p&gt;That's the bet I made: &lt;strong&gt;get big-model capacity at small-model cost&lt;/strong&gt;, on a free Tesla T4, for $0. Not to beat billion-parameter models — that would be absurd — but to prove you can design a credible tiny MoE architecture &lt;em&gt;principally&lt;/em&gt;, not by vibes.&lt;/p&gt;

&lt;h2&gt;
  
  
  The research system I used
&lt;/h2&gt;

&lt;p&gt;I have a methodology I call the &lt;strong&gt;Main Researcher System v4&lt;/strong&gt; — recursive pattern-combination for discovering new algorithms. Instead of hand-tuning an architecture by intuition, I ran the process:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pattern extraction.&lt;/strong&gt; I pulled 12 MoE primitives from the literature (fine-grained segmentation, shared experts, token-choice vs expert-choice routing, aux-loss balancing, router z-loss, SwiGLU experts, dropless training...) and 11 meta-patterns (TRIZ, morphological analysis, MAP-Elites, DreamCoder abstraction, and more).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;TRIZ contradiction resolution.&lt;/strong&gt; The central contradiction: &lt;em&gt;more experts = more capacity BUT more routing cost, imbalance risk, and data dilution; tiny model BUT wants excellence in every category.&lt;/em&gt; The resolution was structural, not a compromise: &lt;strong&gt;Segmentation&lt;/strong&gt; (split coarse experts into 64 micro-experts — capacity without extra per-token compute), &lt;strong&gt;Taking Out&lt;/strong&gt; (extract common knowledge into a shared expert), &lt;strong&gt;Local Quality&lt;/strong&gt; (specialization emerges from the router, never hard-coded).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Combinatorial engine.&lt;/strong&gt; A morphological Zwicky box over 8 parameters screened out inconsistent combos (expert-choice routing leaks causally in decoder LMs — out; top-1 with no aux loss — guaranteed collapse). Documented GP/SCAMPER operators then generated &lt;strong&gt;5 candidates&lt;/strong&gt;: 3 sourced-track (DeepSeekMoE-style, Switch-style, Mixtral-style), 1 bio-inspired (a fatigue-homeostasis balancer with no aux loss), and 1 moonshot (a domain-hint gate with dropout).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Dual evaluation.&lt;/strong&gt; Sourced candidates were scored on Function ÷ active-parameters using published evidence. The moonshot had to pass five falsifiable logical checks. Then all four tested candidates ran &lt;strong&gt;real 300-step CPU experiments&lt;/strong&gt; measuring loss decrease, dead experts, and routing balance.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The winner: DeepSeekMoE-tiny
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;188,269,568 total params, 51,430,400 active per token&lt;/strong&gt; — the compute of a ~51M dense model, ~3.7× its capacity&lt;/li&gt;
&lt;li&gt;8 layers, d_model 512, 8 heads, context 1024&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;64 routed SwiGLU micro-experts (dim 192) + 1 always-on shared expert&lt;/strong&gt; per layer&lt;/li&gt;
&lt;li&gt;Token-choice &lt;strong&gt;top-6&lt;/strong&gt; routing (renormalized), expert-level aux loss (α=0.01) + router z-loss (1e-3), dropless, fp16&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Why not 9 domain experts — one per field? I tried that design first and rejected it. The model has to be best across &lt;em&gt;every&lt;/em&gt; category, not 9 buckets. DeepSeekMoE's combinatorics (~75M routing combinations per layer at identical FLOPs) let 64 micro-experts cover unlimited fine patterns — syntax, facts, reasoning styles — instead of 9 coarse buckets. Specialization emerges; it's never assigned.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real numbers
&lt;/h2&gt;

&lt;p&gt;Sourced value scores (Function ÷ active-params): &lt;strong&gt;A 0.175&lt;/strong&gt; &amp;gt; B 0.161 ≈ D 0.161 &amp;gt; E 0.156 &amp;gt; C 0.097. The Mixtral-style coarse design was worst — most active compute for the coarsest routing.&lt;/p&gt;

&lt;p&gt;Empirical (300 CPU steps each, real runs): all four candidates trained stably (loss 41.5 → 7.6, no NaN, no collapse). The winner balanced best — &lt;strong&gt;aux loss 2.03, zero dead experts&lt;/strong&gt;. The bio-inspired balancer &lt;em&gt;almost&lt;/em&gt; worked but left 1 expert starving (utilization 0.002) — gated out by the hard robustness rule: no babysitting, no starving experts. The moonshot degraded gracefully but didn't beat the winner, so it's archived, not selected.&lt;/p&gt;

&lt;p&gt;And here's the honest part: &lt;strong&gt;cluster↔expert mutual information was ~0 for all candidates.&lt;/strong&gt; 300 toy steps cannot measure specialization. The experiments prove routing &lt;em&gt;balance and stability&lt;/em&gt;, not quality. Quality evidence stays sourced from the literature. I claim no more than that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The candidates that didn't win (and why they matter)
&lt;/h2&gt;

&lt;p&gt;The archive keeps the losers, because losers are reusable parts. The &lt;strong&gt;Switch-style&lt;/strong&gt; design (top-1 routing, no shared expert) was the efficiency champion on paper — only 37.3M active per token — but top-1 routing at tiny scale starves the router of gradient signal, and it lost on sourced value. It stays as the efficiency-cell elite.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;bio-inspired&lt;/strong&gt; design is my favorite failure. Instead of an auxiliary loss term fighting the language-modeling objective, each expert tracks a "fatigue" exponential moving average of its own usage; routing scores get penalized by fatigue, like neural homeostasis — or ants laying pheromones. It &lt;em&gt;almost&lt;/em&gt; balanced perfectly with zero loss-term overhead. Almost. One expert starved at 0.002 utilization, and my hard rule is no starving experts, no babysitting. So it's gated out as the primary — but the mechanism is viable, and it's the obvious crossover partner for the next generation.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;moonshot&lt;/strong&gt; — a domain-hint gate with 50% hint dropout — passed all five logical consistency checks, including graceful degradation (no-hint inference was actually slightly &lt;em&gt;better&lt;/em&gt; than with-hint, gap −0.061). It just didn't beat the winner on value. Archived, not deleted. That's the whole point of the moonshot track: novelty is allowed to win, but only if it earns it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every known failure mode has a guardrail
&lt;/h2&gt;

&lt;p&gt;Inversion — "how would I guarantee this fails?" — produced 8 guardrails baked into the trainer: routing-collapse monitor (loud warning if &amp;gt;25% experts go near-dead), aux-loss dominance warning, auto-halve micro-batch on OOM, NaN abort with diagnostics, checkpoints every 1,000 steps with &lt;code&gt;--resume&lt;/code&gt; for Colab preemption, and more.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Training is running right now on a free Colab T4: Phase 1 on TinyStories (~200M tokens) for fluent language, then Phase 2 on mixed instruction data (~15M tokens, including my Shivacon domain sets and Dolly-15k). Total ~215M tokens, ~2.5–3 hours, $0.&lt;/p&gt;

&lt;p&gt;When it finishes, the real test begins: does the full-scale model utilize its experts well, and does it beat a dense-51M baseline? I'll publish those numbers whatever they say — including if they're bad.&lt;/p&gt;

&lt;p&gt;The full research trace, architecture docs, model, and trainer are in the project repo. This is what $0 and a methodology can buy you: not a miracle, but a design you can defend.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Shivam Kumar is an AI researcher and founder of VisionQuantech, working on AGI/ASI research funded by an AI services business. All 9 Shivacon domain adapters and this MoE model are built on free infrastructure.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>llm</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
