DEV Community

Cover image for One Model, Many Roles: Specializing LLM Inference Without Training More Models
Yuri Pocepaev
Yuri Pocepaev

Posted on Originally published at badbat4560.hashnode.dev

One Model, Many Roles: Specializing LLM Inference Without Training More Models

I ran the same ten tool-selection tasks with four different tool catalogues against one shared model backend: 40 requests in total.

With five tools, the requests used 6,208 input tokens in total. With fifty tools, they used 38,288. Every condition returned the expected tool names and arguments, including correctly making no calls when tools were unnecessary.

That is 83.79% fewer input tokens with no selection errors in either condition on this small test. It is not evidence that fewer tools made the model more accurate. The larger catalogue also passed every task.

This result gives a concrete starting point for a broader architecture: one model, multiple inference roles.

Consider a user asking an assistant to investigate a failing integration, check the documentation, and propose a patch.

One approach puts everything into one request: the conversation history, repository context, search tools, execution tools, memory tools, and instructions for every possible task. The model must decide what matters, which tools to use, and when to stop.

Another approach gives each stage a narrower job. A research stage collects evidence. A coding stage inspects the relevant implementation. A final stage reconciles their findings and answers the user.

All three can use the same model weights.

The architectural unit is an inference profile: the instructions, context, tools, budgets, and output contract surrounding a model call. Specialization can happen at this boundary without training another model or loading another copy of its weights.

The experiment below isolates one part of that design: the tool schemas supplied to a request. The role architecture is the design under discussion; the measured result is a tool-selection pilot, not a benchmark of a complete routed agent system.

One checkpoint, several operating roles

A model checkpoint and an application role describe different things.

The checkpoint provides the model's learned capabilities. An inference profile determines which of those capabilities a request asks it to exercise, what information it receives, and which actions the application permits.

Consider four illustrative profiles:

Profile Context supplied Available actions Expected result
Chat Relevant conversation and authorized user context A small set of everyday capabilities A direct user-facing answer
Reasoning The problem, constraints, and selected evidence Only tools needed to resolve remaining uncertainty An analyzed answer within an explicit budget
Code Relevant files, interfaces, errors, and test results Scoped repository operations and sandboxed execution A patch, diagnosis, or verification report
Mediator A bounded research question and source constraints Authorized retrieval and evidence collection Structured evidence for another stage

These are design examples, not a prescription for four permanent agents.

User request
     |
     v
Routing and authorization
     |
     +--> Chat profile --------+
     +--> Reasoning profile ---+
     +--> Code profile --------+--> Shared inference server
     +--> Mediator profile ----+          |
                                          v
                                   One checkpoint
Enter fullscreen mode Exit fullscreen mode

The routing layer can select one profile or arrange several calls into a workflow. It need not send every request through every role.

Keep three identifiers distinct: the user-facing entry point, the worker's responsibility, and the model endpoint or alias. A “code” entry point might need a research worker before a coding worker. Several workers might call the same endpoint with different request parameters. Naming four aliases does not, by itself, create four different behaviors.

Make the inference profile an executable contract

A useful profile specifies more than a personality prompt:

Component Decision to make
Instructions What task should this stage complete, and what is outside its responsibility?
Context policy Which messages, files, memories, and retrieved evidence belong in this request?
Tool policy Which schemas are exposed, and which operations may execute?
Generation policy What sampling settings and supported reasoning controls apply?
Resource budget How much input, output, elapsed time, and how many iterations are allowed?
Output contract Should the result be prose, tool calls, a patch, or validated structured data?
Admission policy How many requests may this role submit, and what happens under overload?

Configuration has to reach the actual call. A declared output limit is ineffective if a worker adapter never passes it to the model request. A tool allowlist is ineffective if another path attaches the global catalogue afterward.

Resolve the effective profile before dispatch and record its version. Include the actual tool set, context size, generation limits, and timeout in request-level diagnostics without logging private content unnecessarily.

Reasoning controls also depend on model and serving support. Raising an output-token limit can allow a longer response; it does not automatically enable a distinct reasoning mode.

A smaller action space is a testable hypothesis

Suppose the model needs to inspect a repository. Alongside repository tools, it receives schemas for messaging, calendars, customer records, web search, and document management.

Every schema adds information to process. Some tools introduce plausible but inappropriate alternatives. Closely related tools can be especially difficult to distinguish: searching a file, searching its commit history, and searching an external knowledge base all sound relevant to “find where this behavior came from.”

The hypothesis is straightforward:

Accuracy may improve because the action space becomes smaller, even when the model weights stay unchanged.

“May” matters. Fewer tools do not guarantee better answers. A router can exclude the necessary tool, or a small set can still contain highly ambiguous choices. Conversely, a large collection of clearly distinct tools may be easier than a small collection with overlapping descriptions.

Tool count is therefore only one variable. Measure semantic similarity, schema length, description quality, and whether the required capabilities survive selection.

MetaTool treats deciding whether to use tools and selecting the appropriate tools as evaluation problems, including single-tool and multi-tool requests. That distinction is useful here: choosing a valid tool is not enough when the correct action is to answer without one.

There is external evidence that selective loading can help. In its advanced tool-use report, Anthropic describes reducing context consumption from roughly 77K to 8.7K tokens with tool search. It also reports improved accuracy on internal MCP evaluations. Those are vendor results for its models and evaluation setup, not measurements of this architecture or a guarantee for another checkpoint.

Static roles and deferred discovery solve different problems

A static role exposes a known subset from the beginning. This is easy to inspect and useful when task boundaries are stable: a repository diagnosis normally needs repository tools, not a calendar API.

Deferred discovery begins with a small discovery interface and loads schemas when needed. It can handle a larger and less predictable catalogue, but introduces another decision: finding the right capability before using it.

AnyTool explores hierarchical API retrieval, a solver operating on selected candidates, and self-reflection when an initial solution is unsuccessful. It is relevant as an example of separating capability discovery from task execution, rather than treating every API definition as permanent prompt content.

A practical combination is a small core tool set plus a controlled discovery path. If a task needs something outside the current role, the worker can request a handoff or additional authorized capability.

That fallback needs to appear in evaluation results. A system that succeeds only after repeatedly expanding back to the full catalogue has different costs from one that completes the task within its original subset.

Tool visibility is not authorization. The execution layer must independently enforce permissions, even when a schema was omitted from the prompt. A model-generated tool name must never bypass those checks.

The mediator should return evidence, not another confident opinion

The mediator makes the separation of responsibilities concrete. Its job is to collect support for another stage's decision.

An illustrative pipeline is:

Research question
  -> Select relevant domains
  -> Retrieve authorized source material
  -> Expose suitable evidence-collection tools
  -> Produce and validate a structured plan
  -> Execute permitted actions
  -> Return evidence, gaps, and source references
Enter fullscreen mode Exit fullscreen mode

Its response can distinguish supported claims, unresolved questions, and conflicting sources. For each claim, preserve the source identifier and enough evidence for the next stage to check it. A concise evidence package should retain a path back to the original material.

Schema-constrained JSON helps enforce the response shape. It does not establish that a claim is true or that its citation supports it. Those checks remain separate.

If retrieval finds nothing useful, the contract should permit “insufficient evidence.” Otherwise, a stage designed to gather evidence may fill the gap with its own unsupported answer.

The final answering stage should also be able to request more evidence. A mediator that aggressively compresses away qualifications can make downstream answers faster and less accurate at the same time.

Run independent stages concurrently

In the integration example, reading the relevant documentation and inspecting the existing implementation can often start independently. Producing the final patch may depend on both. Testing the patch depends on the patch existing.

                  +--> Documentation evidence --+
Request and scope |                              +--> Patch --> Tests --> Answer
                  +--> Repository inspection ---+
Enter fullscreen mode Exit fullscreen mode

This is a dependency graph. Execute independent stages concurrently and preserve ordering where one stage needs another's output. Check side effects too: two operations without a data dependency may still conflict if they write the same resource.

Do not equate asynchronous application code with parallel model execution. Concurrent submissions reach a shared scheduler, which may batch, interleave, or queue them. Network retrieval can overlap with inference, but several generation requests still compete for accelerator capacity.

LLMCompiler studies planned parallel function execution and reports latency improvements of up to 3.7× over ReAct in its evaluated tasks. That supports studying dependency-aware execution; it does not establish the same speedup for several inference roles sharing one backend.

Shared weights also mean shared constraints

One serving process avoids keeping a separate weight copy for every logical role. It does not eliminate the memory and compute costs of their requests.

Active sequences still need state. Long generations occupy capacity. Large prompts consume prefill work. A research burst can increase queueing for interactive chat.

Different service targets therefore require enforcement: per-role concurrency caps, admission control, deadlines, cancellation, bounded fan-out, and overload behavior. A role label alone does not reserve capacity or guarantee an SLA.

The server is also a shared failure domain. If it becomes unavailable, every role depending on it is affected. Hard isolation requirements may justify separate serving instances even when those instances use identical weights.

One checkpoint simplifies weight management, but each profile still needs evaluation. Different prompts, context policies, tools, and budgets produce different behavior. A checkpoint upgrade can regress coding, retrieval, or abstention independently.

What I actually tested

The pilot used ten synthetic requests: five in English and five in Russian. Six requested a single tool, two required no tool, and two requested two independent operations.

The core catalogue exposed five repository-related capabilities: reading current file contents, searching source text, reading commit history, running a named test target, and listing files beneath a directory. All tools were schemas only. No repository was read, no test target executed, and no external service modified by a generated call.

Each task ran once in each condition:

Condition Schemas exposed Calls
A Five core tools 10
B The core tools plus fifteen repository-related alternatives 10
C The core tools plus fifteen tools from unrelated domains 10
D Fifty mixed tools, including the same core set 10

Repository-related alternatives included historical file reads, commit-message search, documentation search, and reading an existing test report. The unrelated tools covered domains such as weather, calendars, music, inventory, and travel. Their descriptions were synthetic and often made the distinction explicit.

For each task, the system instruction and user prompt stayed the same across conditions. The calls used the same endpoint and requested model alias, temperature zero, automatic tool choice, an output limit of 256 tokens, and a fixed seed. The request explicitly disabled thinking through the serving template option. Both condition order and schema order were shuffled deterministically.

This was one run per task and condition, not repeated trials across several seeds. The responses identified the same backend model alias and serving fingerprint; an immutable weight digest was not captured. Endpoint identity is therefore the recorded control, not an independently verified checkpoint revision.

Requests ran sequentially, with pauses of at least fifteen seconds, on a backend already serving other traffic. This avoided a concurrency stress test but did not create a controlled latency environment.

Results: the same selections with less input

Tool catalogue Expected tool selection Exact arguments Median input tokens Total input tokens
5 core tools 10/10 10/10 619.5 6,208
20, with repository alternatives 10/10 10/10 1,691.5 16,928
20, with cross-domain alternatives 10/10 10/10 1,759.5 17,608
50 mixed tools 10/10 10/10 3,827.5 38,288

Input tokens are the server's reported prompt_tokens for the complete request, including instructions, the user message, and tool definitions. They are not a separate measurement of schema tokens or uncached compute.

Across the same ten tasks, the reduction from the fifty-tool condition to the five-tool condition was:

1 − 6,208 / 38,288 = 83.79% fewer reported input tokens
Enter fullscreen mode Exit fullscreen mode

The fifty-tool condition used about 6.17 times as many input tokens. This is a request-context result, not a measured reduction in GPU time, billed cost, or end-to-end agent cost.

The scorer compared the returned tool names and parsed arguments against the expected calls, ignoring call order for the two independent-operation tasks. It also checked allowed names, required argument keys, and string value types. For no-tool tasks, success meant returning no tool calls; the prose answer was not independently scored.

All forty responses met those checks. None hit the output limit. Multi-tool requests returned both expected calls in the first response; the experiment did not run a tool loop afterward.

The honest conclusion is narrow: the extra schemas were unnecessary for these tasks, and removing them reduced reported input tokens without introducing selection errors in this run.

What the pilot does not establish

The test has a ceiling effect: all conditions scored perfectly. It provides no evidence of improved selection accuracy from a smaller tool space and cannot establish equivalence on a broader workload.

The tasks are short and explicit. “Read the current contents” gives a strong clue against a historical-file tool. Several unrelated descriptions explicitly distinguish themselves from repository operations. Real requests may be ambiguous, incomplete, or require capabilities the router initially excludes.

The small catalogue was selected by the experiment author. No router had to discover it. Consequently, the result excludes routing errors and routing overhead. It also excludes deferred discovery, real execution, grounded final answers, memory-policy changes, context pollution from tool results, and parallel inference.

Recorded wall time includes shared-backend queueing and transport overhead. Requests with more tools sometimes completed sooner because they encountered different traffic. Treating those times as a catalogue-size speed comparison would be misleading. Streaming TTFT and isolated prefill time were not measured.

The next useful experiment should add ambiguous near-neighbor tools, less explicit requests, missing-capability cases, and repeated shuffled trials. A real router and a deferred-discovery path should then be evaluated separately, counting their failures, fallback expansion, tokens, and latency.

Context policy and scheduling need their own comparisons. Changing tools, prompts, memory, reasoning limits, and concurrency together can measure an overall system difference, but cannot identify which change caused it.

Measure successful work, including the extra stages

Use several levels of correctness:

  • Selection: Did the model choose an acceptable tool or correctly abstain?
  • Arguments: Did the call satisfy both the schema and the task's actual constraints?
  • Execution: Did the authorized operation succeed against a controlled fixture?
  • Answer: Did the final response solve the task using the returned evidence?

The Berkeley Function Calling evaluation distinguishes multiple candidate functions, parallel calls, relevance detection, and executable checks. These categories help prevent a single tool-name score from standing in for complete task success.

Also record tool-schema tokens, total input and output tokens across every stage, iteration counts, retries, time to first token, queue time, total latency, finish reasons, and failure rates. Report latency distributions and uncertainty on accuracy differences, not just one average from a small run.

For staged systems, distinguish first visible text from the time to a useful, completed answer. An early “working on it” message should not make a slow workflow look fast.

Measure cold and warm cache behavior separately. A shorter prompt can reduce uncached prefill work, but role-specific prefixes may reduce sharing between requests. Prefix-cache reuse depends on the actual serving implementation and prompt structure; it is not an automatic benefit of specialization.

Use isolated synthetic fixtures for operations with side effects. Store the corpus, schemas, ordering seeds, effective profiles, scoring rules, and raw responses with the results. Request-level usage must be collected from actual calls; placeholder zero counters are missing telemetry, not free inference.

The boundary is the design

The pattern is useful when a capable model serves tasks with different information needs, action sets, and latency expectations. It is less attractive when routing adds more overhead than it removes, or when most tasks repeatedly need the same broad context and tools.

The pilot supports a practical reason to start with a focused tool set: on these tasks, the larger catalogue consumed substantially more request context and produced no better selections. It does not settle whether role routing improves overall accuracy, cost, or latency.

A good role has a clear responsibility, sufficient authorized context, an appropriate action set, and a checkable result. That makes specialization something the application can inspect and test, instead of a label attached to another prompt.


Related: Smaller Context, Recoverable History · Prevent Context Leakage in Multi-Tenant LLM Systems

More engineering notes: Hashnode · DEV · GitHub · Hugging Face · Telegram · Instagram

Originally published on Hashnode.

Top comments (0)