I’d split this decision into two questions: do I want to use a coding agent, or do I want to own and modify the runtime around it?
Claude Code is the more opinionated product: integrated workflows, mature permissions, and a closed agent loop optimized for Claude. DeepSeek Harness (dsh) exposes that machinery through an MIT-licensed, plugin-based runtime.
That distinction matters more than whether both can edit files or launch sub-agents. Their tool lists overlap substantially. Their maintenance burden, auditability, model flexibility, and billing do not.
This comparison reflects the reported August 2026 developer-preview state of DeepSeek Harness. Performance figures below are individual reported runs, not reproducible guarantees.
My Default Choice Depends on the Deliverable
For client-facing work under a deadline, I’d favor Claude Code. The reported advantage is not simply code correctness; it is context awareness, verification, and the amount of cleanup needed before shipping.
For agent infrastructure, model experiments, bulk work, or custom execution policies, I’d evaluate DeepSeek Harness first. Its main attraction is the ability to replace parts of the runtime without maintaining a fork.
| Decision point | DeepSeek Harness | Claude Code |
|---|---|---|
| License | MIT, open source | Proprietary commercial product |
| August 2026 status | Developer preview; breaking changes expected | Stable production product |
| Model access | OpenAI-compatible endpoints | Primarily Claude; Bedrock and Vertex options |
| Main interfaces | Local web UI, headless CLI, Python SDK | Terminal, VS Code, JetBrains, desktop, and other integrations |
| Extension model | Cordis plugins throughout the runtime | Skills, hooks, MCP, preview plugins |
| Internal auditability | Source and trajectory logs | Closed implementation |
| Execution controls | Configurable sandbox plugins; restricted defaults | Built-in permissions, approval prompts, and sandbox controls |
| Billing | Free runtime license; underlying model costs | Subscription or API billing |
| Session persistence | Searchable local storage | Supported |
| Best fit | Runtime ownership and multi-model workflows | Production workflow polish |
I don’t see a strong reason to force an exclusive choice. DSH can invoke Claude Code or Codex as a sub-agent, so the boundary can be architectural rather than organizational.
What You Actually Own With DSH
DeepSeek Harness is described as having launched on August 13, 2026, alongside DeepSeek V4-Pro. Its framing is straightforward: Agent = Model + Harness.
The model generates decisions and code. The harness supplies tools, filesystem access, execution, memory, planning, permissions, sub-agent scheduling, and the interface.
The interesting part is the Cordis architecture. “Everything is a plugin” extends beyond model adapters and custom tools:
- The agent loop and planning machinery
- Tool registration and skills
- Sandboxes and permission policies
- Session storage and memory
- Scheduling and sub-agents
- The UI itself
That is a materially different extension boundary from adding a hook to a fixed agent loop. I can see its value for teams that need proprietary memory, internal security policies, specialized execution environments, or reproducible agent experiments.
Profiles and bundles compose configurations without requiring source changes. The available runtime profiles include Standard, Code/PTC, Minimal for benchmarking, and Creator.
Inspectability is part of the architecture
DSH includes append-only trajectory/event logs for inspection and replay. The reported trajectory view exposes system prompts, tool calls, reasoning records, and sub-agent scheduling.
That matters when a run fails in a way that ordinary application logs cannot explain. Was the tool definition misleading? Did the sandbox reject an operation? Did retries consume the budget? Did a sub-agent receive the wrong context?
The default interface is a local web UI at http://127.0.0.1:3080. Headless and CLI modes support scripting and CI, and a Python SDK provides another integration surface.
The project also reportedly attracted well over 100,000 stars, with counts exceeding 190,000 within days. I would treat that as ecosystem interest—not evidence that a preview runtime is production-ready.
What Claude Code Keeps Opinionated
Claude Code packages the agent loop as a product rather than as replaceable infrastructure.
Its built-in capabilities include file reads and edits, shell execution, planning, sub-agents, skills, hooks, MCP client support, and permission prompts. Integration extends beyond the terminal to VS Code, JetBrains, desktop apps, browser, mobile, Slack, and GitHub workflows, including Actions, PR review, and issue-to-PR flows.
The model layer is tightly coupled to Claude’s Opus, Sonnet, and Haiku families. That limits provider freedom but gives Anthropic control over the relationship between model behavior, tools, context management, and verification.
Extensions operate through defined surfaces. I can add skills, hook into lifecycle events, or connect MCP tools; I cannot independently inspect or replace the proprietary core loop.
For many teams, that is an acceptable trade. Maintaining an agent runtime is work, even when the license is free.
The execution model also affects throughput
One architectural approach discussed in these comparisons is Code Mode: generating TypeScript programs through a Code Mode SDK to orchestrate multiple rounds of tool calls.
Rather than issuing every operation separately, an agent can express a higher-level procedure. That is relevant to repetitive automation, long workflows, and custom orchestration.
I would evaluate that separately from the model choice. Tool orchestration changes latency and failure behavior even when the underlying model stays the same.
The Reported Speed Gap Needs Context
The headline results favor DSH on wall-clock time, but the individual runs tell a more useful story.
| Reported scenario | DeepSeek Harness | Claude Code |
|---|---|---|
| Complete website/build task | About 11 minutes | Still running after 30 minutes |
| Same-model, Claude Opus-class comparison | About 3 minutes | About 17 minutes |
| Token usage in one comparison | 483k tokens | 48k tokens |
| Subjective quality rating in one detailed comparison | About 7/10 | About 9/10 |
These numbers should not be collapsed into “DSH is faster and cheaper.” The faster runtime sometimes consumed substantially more tokens.
The detailed comparison favored Claude Code on visual polish, business-context awareness, UX, and client-ready presentation. The testers described moving some tasks to DSH rather than replacing Claude Code entirely.
Other same-model tests reportedly produced byte-identical patches that passed test suites. In one Windows run, Claude Code was faster after approvals. That is a useful counterexample to any universal speed ranking.
I would separate three measurements
- Time to first completed run
- Cost of that run
- Time and cost to reach an acceptable deliverable
The third is the one I care about for production work. A quick implementation that needs another verification pass is not necessarily the faster workflow.
Claude Code reportedly remains stronger for long sessions, deep work, and visual verification. DSH’s v0.1 preview has shown bugs, context-management issues, and weaker verification with some models. Its more convincing early uses are fast research, parallel work, overflow tasks, and inexpensive experimentation.
Model benchmarks add another layer. The reported pattern puts Claude ahead on SWE-bench-style engineering and overall coding indices, while DeepSeek is competitive on LiveCodeBench and algorithmic tasks at lower prices.
Neither pattern isolates the harness. Tool definitions, retry policies, sandbox behavior, and context handling can materially change the result.
Cost: Compare the Whole Run, Not the License
DSH has no license fee. Model inference still costs money unless the chosen deployment avoids provider billing, and operating the runtime is not free of engineering effort.
The quoted DeepSeek V4-Flash rates are approximately $0.14 per million input tokens and $0.28 per million output tokens, subject to current pricing and caching. V4-Pro is higher-priced but still positioned as a low-cost option.
Reported substantial builds have cost around five cents. Token-price comparisons put DeepSeek anywhere from 10–57× cheaper, depending on model selection and caching.
Those are not interchangeable measurements:
- A token-price ratio does not describe total run cost.
- Total run cost depends on token consumption and retries.
- Subscription economics depend on how much usable work fits within the plan.
Claude Code supports subscription and API billing. The cited subscription range is $20–$200 per month, with Pro at $20 and higher Max tiers aimed at heavier usage. Usage pools are shared across Claude surfaces.
One analysis projected roughly $1,200 annually for a heavy solo Claude Code user versus $240–$600 in tokens for DeepSeek-based workflows. Heavy solo or team usage can also reach hundreds or thousands of dollars monthly. I would treat these as workload-specific estimates, not a budget calculator.
Routing is where the flexibility becomes useful
I would reserve premium inference for tasks where mistakes or weak context handling are expensive, then test cheaper models on bulk transformations, research, and parallel subtasks.
A unified multi-model gateway such as CometAPI can simplify that experiment: its advertised offering includes one key, an OpenAI-compatible endpoint at https://api.cometapi.com/v1, and 500+ models, including DeepSeek V4-Pro/Flash and Claude families. Advertised effective discounts of 20–40% still need checking against the actual model, caching, and billing terms.
DSH can target compatible endpoints directly. Claude Code routing requires an appropriate Anthropic-compatible endpoint and base-URL configuration; an OpenAI-compatible URL alone is not equivalent.
Security: Inspectability Is Not a Permission Policy
DSH’s restricted sandbox defaults are described as bwrap/Landlock-style, and local credentials use restrictive permissions such as 0600. Its plugin architecture lets teams replace sandboxes, storage, tools, and internal security controls.
That gives organizations with runtime-audit requirements a structural advantage: they can inspect and rebuild the implementation.
It does not automatically make every configuration safe. Replacing the sandbox or tool layer also means owning the consequences.
Claude Code offers mature approval prompts and execution controls without exposing the full implementation. That is less suitable for independent internal auditing, but potentially simpler for teams that want a maintained security UX rather than a custom one.
My distinction is control versus operational responsibility, not open source versus secure.
How I Would Start an Evaluation
For DSH, the supplied launch command is:
npx @deepseek-ai/dsh web
The Node.js/TypeScript-based runtime supports npm/pnpm installation. Launch the UI, supply a DeepSeek API key or another compatible endpoint, and configure the required plugins.
I would expect configuration churn while the developer preview evolves.
Claude Code’s described setup uses a global npm installation of Anthropic’s package, followed by running the agent inside a project directory. Its official documentation and IDE extensions cover the product workflow.
For a meaningful comparison, I would hold the repository, prompt, acceptance tests, and model constant where possible. Then I would repeat with the cheaper model to measure the combined runtime-and-model trade-off.
Where I Would Use Each
Claude Code would be my default for:
- Client-facing deliverables and production deadlines
- Long, context-heavy coding sessions
- Workflows dependent on IDE integration
- Teams that prefer maintained behavior over runtime customization
DeepSeek Harness would be my first candidate for:
- Auditable, modifiable agent infrastructure
- Multi-provider or local-model experiments
- Custom tools, storage, execution policies, and agent loops
- Cost-sensitive parallel work where preview friction is acceptable
A hybrid would make sense when:
- Cheap models can handle bulk work
- Critical paths benefit from Claude Code’s verification and polish
- DSH orchestration is itself useful enough to justify maintaining it
The deciding question for me is not which agent finishes a demo first. It is whether runtime ownership improves the workflow enough to justify owning another piece of infrastructure.
Top comments (0)