XCP — Chapter 4
The previous chapters covered XVM, the open-source XCP platform, and XCP Studio.
This one goes underneath the interface.
Into the worker.
Running compute code on an Xbox Series X was not the hardest part of XCP.
The harder part was understanding exactly what execution environment we actually had.
Microsoft documents Xbox Developer Mode and the Windows application model, but that does not answer every practical question that appears when you start treating the console as a compute target.
How much memory can this process really use?
Can that memory remain allocated and intact?
Which GPU paths actually work?
How stable are they under repeated load?
Where does the sandbox stop us?
So before XCP became an execution system, a large part of the work was simply mapping the environment from inside it.
This was not about escaping the sandbox.
It was about understanding the sandbox.
We had to map the machine first
We started with small probes.
Every time one exposed a useful capability, we pushed further.
capability
↓
measure
↓
find boundary
↓
push boundary
↓
measure again
Sometimes that meant finding more usable memory.
Sometimes it meant changing the shape of a GPU workload and watching measured FP32 throughput move from roughly 4.6 TFLOP/s to around 7.7 TFLOP/s.
Sometimes we hit a hard boundary and had to redesign the runtime around it.
The system was not designed from a complete map of the machine.
The map was built experimentally, one verified boundary at a time.
That process eventually changed the central question.
At first it was:
How much can we make the Xbox do?
Later it became:
How can we prove what it actually did?
Memory: allocation is not retention
The platform can report an application memory budget.
But a reported budget is not the same thing as memory that can be allocated, retained and verified over time.
So XCP moved beyond simply reading:
AppMemoryUsageLimit
AppMemoryUsage
AppMemoryUsageLevel
The worker began allocating bounded regions, touching their pages, retaining them across heartbeat windows and checking that their contents remained intact.
One measured Developer Mode configuration exposed an application memory limit of 5.0 GiB.
But the number itself was only the beginning.
The useful question was whether memory could survive actual use.
Show the memory probe
The worker performs bounded allocations using VirtualAllocFromApp:
memory = static_cast(
VirtualAllocFromApp(
nullptr,
static_cast(committedBytes),
MEM_RESERVE | MEM_COMMIT,
PAGE_READWRITE));
if (!memory)
{
Fail(
L"topology.memory_allocation_failed",
L"the bounded foreground memory allocation failed");
}
The allocation is checked with VirtualQuery.
Then every page receives a deterministic generation marker:
for (uint64_t page = 0; page < pageCount; ++page)
{
memory[page * pageBytes] =
Marker(page, currentGeneration);
}
Those markers are verified immediately and again across later heartbeat windows:
for (uint64_t page = 0; page < pageCount; ++page)
{
if (memory[page * pageBytes] !=
Marker(page, currentGeneration))
{
++mismatchCount;
}
}
The retained state is also hashed.
So the test becomes:
allocate
↓
commit
↓
touch pages
↓
verify
↓
hash
↓
wait
↓
verify retention
↓
repeat
allocation succeeded and allocation remained valid are deliberately treated as different claims.
Then the GPU opened a much deeper path
The early GPU experiments existed primarily to answer a basic question:
Is there meaningful compute available to this application inside the sandbox?
The answer was yes.
A calibrated FP32 repeatability path averaged roughly:
4.63 TFLOP/s
But changing the shader shape exposed substantially more throughput.
The later fp32_alu_128 soak averaged:
7.705 TFLOP/s
across repeated measured windows with zero result mismatches.
Historical physical-Xbox GPU measurements from the XCP development programme. These are bounded measurements of the tested Developer Mode worker path, not a claim of access to the console's full theoretical 12.1 TFLOPS.
That graph is useful, but the performance number was not the endpoint.
In fact, it exposed the next problem.
A TFLOPS figure can tell me how quickly a deliberately constructed shader executed.
It cannot tell me:
- whether the intended program actually ran;
- whether the correct shader produced the result;
- whether the resulting state is valid;
- whether repeated execution remains consistent;
- whether CPU and GPU agree.
Performance showed that the compute surface existed.
Now we had to make it trustworthy.
Show the measured GPU results
The historical physical-Xbox worker measurements include:
- measured application memory limit: 5.0 GiB
- calibrated FP32 repeatability: 773,094,113,280 timed operations
- average: 4,632.654 GFLOP/s
- heavier shader-shape matrix best row: 7,679.569 GFLOP/s
- targeted
fp32_alu_128run: 7,708.187 GFLOP/s - longer soak average: 7,705.76 GFLOP/s
- soak total: 2,473,901,162,496 FP32 operations
- soak min/max: 7,699.702–7,711.882 GFLOP/s
- spread: 0.158064%
- mismatch count: 0
- violations: 0
- measured memory-copy peak in the mixed matrix: 306,550.022 MiB/s
These measurements characterize bounded public-UWP/Developer Mode worker paths.
They do not establish unrestricted GPU access, GDK/Game Mode entitlement or access to the full theoretical 12.1 TFLOPS of the console.
Performance was no longer enough
Once the GPU path existed, XCP started putting explicit identity and execution contracts around it.
The worker loads precompiled shader bytecode and verifies that the shader matches the profile that was admitted for execution.
Conceptually:
workload
↓
admission
↓
expected shader identity
↓
actual installed shader
↓
match?
A successful dispatch from the wrong shader is not a valid result.
This sounds obvious, but it changes the meaning of execution.
The GPU is no longer just a device receiving work.
It becomes one component in an evidence chain.
Moving XVM state through the GPU
XVM workloads can be represented using structured GPU resources.
Parameters, input data, program state and — for some profiles — microtrace data are exposed to the compute path as structured buffers.
The output lives in a GPU state buffer.
Then comes the part that became more important than the dispatch itself:
read the state back.
parameters / input / program
↓
StructuredBuffers
↓
compute shader
↓
GPU state
↓
staging buffer
↓
CPU readback
↓
decode
↓
verify
A successful Dispatch() is therefore not the end of the operation.
It is the middle.
Show the D3D readback path
The result state is exposed as an unordered access resource.
A CPU-readable staging buffer is created:
D3D11_BUFFER_DESC stagingDesc{};
stagingDesc.ByteWidth =
stateWordCount * sizeof(uint32_t);
stagingDesc.Usage =
D3D11_USAGE_STAGING;
stagingDesc.CPUAccessFlags =
D3D11_CPU_ACCESS_READ;
After execution:
context->CopyResource(
stagingBuffer.get(),
resultBuffer.get());
context->Flush();
The worker also checks for device loss.
Then the staging buffer is mapped:
D3D11_MAPPED_SUBRESOURCE mapped{};
auto mapHr = context->Map(
stagingBuffer.get(),
0,
D3D11_MAP_READ,
0,
&mapped);
std::memcpy(
words.data(),
mapped.pData,
words.size() * sizeof(uint32_t));
context->Unmap(
stagingBuffer.get(),
0);
Now the GPU state can be decoded and validated on the CPU side.
Every new capability exposed another limit
This became the recurring pattern.
Memory had a budget.
Long-running execution needed bounds.
GPU work needed admitted shapes.
Resident execution needed epochs.
GPU state needed readback.
SPMD execution introduced lane divergence.
Device loss needed to become an explicit state.
And eventually the GPU result itself needed an independent reference.
The limits were not merely obstacles.
They became part of the architecture.
limit discovered
↓
make it explicit
↓
build a contract around it
↓
measure again
That is one of the reasons XCP became evidence-driven.
Instead of pretending a boundary did not exist, we tried to make it machine-readable.
Then we compared GPU execution with the CPU
For bounded classes of XVM execution, XCP can run a CPU reference and compare it with the GPU result.
The comparison includes more than output bytes.
It can include:
output
control state
fuel consumed
trap code
trap PC
recovery PC
trap count
So the model becomes:
XVM workload
│
┌──────┴──────┐
│ │
▼ ▼
CPU reference GPU execution
│ │
│ readback
│ │
└──────┬──────┘
▼
differential check
↓
evidence
At that point the CPU is not simply a fallback.
For the admitted execution class, it becomes a reference against which the GPU can be checked.
Show the CPU/GPU differential checks
The runtime compares output:
auto outputMatches =
input.cpuRun->output == gpu.run.output;
Control state:
auto controlMatches =
input.cpuRun->controlToken ==
gpu.run.controlToken;
Fuel:
auto fuelMatches =
input.cpuRun->fuelConsumed ==
gpu.run.fuelConsumed;
And trap state:
auto trapMatches =
input.cpuRun->trap.code ==
gpu.run.trap.code &&
input.cpuRun->trap.trapPc ==
gpu.run.trap.trapPc &&
input.cpuRun->trap.recoveryPc ==
gpu.run.trap.recoveryPc &&
input.cpuRun->trap.occurrenceCount ==
gpu.run.trap.occurrenceCount;
XCP also has bounded resident execution paths, epoch validation and SPMD lane-consistency checks.
If lanes that are required to remain synchronized diverge, the execution is rejected rather than promoted into a valid result.
The limits shaped the worker
There is a tendency to describe system constraints only as things preventing software from doing more.
That was not what happened here.
The constraints repeatedly forced the system to become more explicit.
A memory ceiling led to measured resource accounting.
Long-lived compute led to bounded execution.
GPU state led to readback.
Readback led to structural validation.
Multiple execution backends led to differential verification.
Device loss led to explicit quarantine and fallback state.
The sandbox itself became part of the design.
We did not remove the limits. We built the worker around measured limits.
From benchmark to execution system
The first experiments were naturally interested in performance.
Could meaningful GPU compute run?
How much memory was available?
How hard could the environment be pushed?
Those were necessary questions.
But eventually they became the less interesting ones.
The harder questions were:
Which program executed?
Which shader produced the result?
What resource boundary was observed?
Did memory remain intact?
Was execution bounded?
Did GPU state progress legally?
Did parallel lanes remain consistent?
Did the GPU agree with the CPU?
What exactly does the evidence allow us to claim?
That is where XCP stopped looking like a benchmark project.
It became an execution system.
The Xbox is still a sandbox
None of this bypasses Xbox Developer Mode.
XCP does not turn the console into an unrestricted PC, and it does not claim full retail-console GPU entitlement.
The interesting part was precisely the opposite.
We kept hitting boundaries.
Then we measured them.
Then we built around them.
That is why the project gradually moved from:
Can this run?
to:
What ran?
Under which boundary?
What evidence came back?
What can we safely claim from it?
One important evidence boundary
The public XCP repository preserves the implementation lineage and historical physical-Xbox evidence.
The newly reconstructed development packages are different artifacts.
They remain:
NOT_TESTED_ON_XBOX
until a new physical-device reproduction generates evidence for those exact packages.
A new binary does not automatically inherit an older binary's execution proof.
That distinction is intentional.
It is the same principle that grew out of all those early probes:
a claim should not be stronger than the measurement behind it.
What this work changed
The difficult part turned out not to be making the hardware do something.
It was constructing a system that could explain later, with enough precision, what the hardware actually did.
For XCP, that path became:
map the sandbox
↓
measure resources
↓
find a boundary
↓
push further
↓
admit bounded work
↓
execute
↓
read back state
↓
compare
↓
record evidence
That is now one of the central ideas behind the whole project.
Execution is an event.
Evidence is what makes the event usable.
XCP is open source
GitHub:
https://github.com/Daniele-Cangi/xcp-xbox
XCP Studio:
Relevant implementation areas include:
reference/xcompute-probe/src/XComputeProbe/probes/ProbeRunner.cpp
reference/xcompute-probe/src/XComputeProbe/runtime/WorkerProcessTopologyMemoryProbe.cpp
reference/xcompute-probe/src/XComputeProbe/runtime/WorkerPrecompiledD3DCommandRuntime.cpp
reference/xcompute-probe/src/XComputeProbe/runtime/WorkerGpuXvmExecutor.cpp
reference/xcompute-probe/src/XComputeProbe/runtime/WorkerGpuXvmDifferentialRuntime.cpp
Previous chapters in the XcpStudio series cover the origin of XVM, the open-source XCP platform and the Studio workflow.
AI can build the workload. Xbox can execute it. XCP's job is to make the result defensible.

Top comments (0)