DEV Community

Cover image for Inside XCP: How We Turned Xbox Memory and GPU Compute Into a Verifiable Worker
DaC
DaC

Posted on

Inside XCP: How We Turned Xbox Memory and GPU Compute Into a Verifiable Worker

XCP — Chapter 4

The previous chapters covered XVM, the open-source XCP platform, and XCP Studio.

This one goes underneath the interface.

Into the worker.

Running compute code on an Xbox Series X was not the hardest part of XCP.

The harder part was understanding exactly what execution environment we actually had.

Microsoft documents Xbox Developer Mode and the Windows application model, but that does not answer every practical question that appears when you start treating the console as a compute target.

How much memory can this process really use?

Can that memory remain allocated and intact?

Which GPU paths actually work?

How stable are they under repeated load?

Where does the sandbox stop us?

So before XCP became an execution system, a large part of the work was simply mapping the environment from inside it.

This was not about escaping the sandbox.

It was about understanding the sandbox.


We had to map the machine first

We started with small probes.

Every time one exposed a useful capability, we pushed further.

capability
    ↓
measure
    ↓
find boundary
    ↓
push boundary
    ↓
measure again
Enter fullscreen mode Exit fullscreen mode

Sometimes that meant finding more usable memory.

Sometimes it meant changing the shape of a GPU workload and watching measured FP32 throughput move from roughly 4.6 TFLOP/s to around 7.7 TFLOP/s.

Sometimes we hit a hard boundary and had to redesign the runtime around it.

The system was not designed from a complete map of the machine.

The map was built experimentally, one verified boundary at a time.

That process eventually changed the central question.

At first it was:

How much can we make the Xbox do?

Later it became:

How can we prove what it actually did?


Memory: allocation is not retention

The platform can report an application memory budget.

But a reported budget is not the same thing as memory that can be allocated, retained and verified over time.

So XCP moved beyond simply reading:

AppMemoryUsageLimit
AppMemoryUsage
AppMemoryUsageLevel
Enter fullscreen mode Exit fullscreen mode

The worker began allocating bounded regions, touching their pages, retaining them across heartbeat windows and checking that their contents remained intact.

One measured Developer Mode configuration exposed an application memory limit of 5.0 GiB.

But the number itself was only the beginning.

The useful question was whether memory could survive actual use.

Show the memory probe

The worker performs bounded allocations using VirtualAllocFromApp:

memory = static_cast(
    VirtualAllocFromApp(
        nullptr,
        static_cast(committedBytes),
        MEM_RESERVE | MEM_COMMIT,
        PAGE_READWRITE));

if (!memory)
{
    Fail(
        L"topology.memory_allocation_failed",
        L"the bounded foreground memory allocation failed");
}
Enter fullscreen mode Exit fullscreen mode

The allocation is checked with VirtualQuery.

Then every page receives a deterministic generation marker:

for (uint64_t page = 0; page < pageCount; ++page)
{
    memory[page * pageBytes] =
        Marker(page, currentGeneration);
}
Enter fullscreen mode Exit fullscreen mode

Those markers are verified immediately and again across later heartbeat windows:

for (uint64_t page = 0; page < pageCount; ++page)
{
    if (memory[page * pageBytes] !=
        Marker(page, currentGeneration))
    {
        ++mismatchCount;
    }
}
Enter fullscreen mode Exit fullscreen mode

The retained state is also hashed.

So the test becomes:

allocate
   ↓
commit
   ↓
touch pages
   ↓
verify
   ↓
hash
   ↓
wait
   ↓
verify retention
   ↓
repeat
Enter fullscreen mode Exit fullscreen mode

allocation succeeded and allocation remained valid are deliberately treated as different claims.


Then the GPU opened a much deeper path

The early GPU experiments existed primarily to answer a basic question:

Is there meaningful compute available to this application inside the sandbox?

The answer was yes.

A calibrated FP32 repeatability path averaged roughly:

4.63 TFLOP/s

But changing the shader shape exposed substantially more throughput.

The later fp32_alu_128 soak averaged:

7.705 TFLOP/s

across repeated measured windows with zero result mismatches.

Historical Xbox GPU compute measurements showing the fp32_alu_128 workload reaching about 7.7 TFLOP/s

Historical physical-Xbox GPU measurements from the XCP development programme. These are bounded measurements of the tested Developer Mode worker path, not a claim of access to the console's full theoretical 12.1 TFLOPS.

That graph is useful, but the performance number was not the endpoint.

In fact, it exposed the next problem.

A TFLOPS figure can tell me how quickly a deliberately constructed shader executed.

It cannot tell me:

  • whether the intended program actually ran;
  • whether the correct shader produced the result;
  • whether the resulting state is valid;
  • whether repeated execution remains consistent;
  • whether CPU and GPU agree.

Performance showed that the compute surface existed.

Now we had to make it trustworthy.

Show the measured GPU results

The historical physical-Xbox worker measurements include:

  • measured application memory limit: 5.0 GiB
  • calibrated FP32 repeatability: 773,094,113,280 timed operations
  • average: 4,632.654 GFLOP/s
  • heavier shader-shape matrix best row: 7,679.569 GFLOP/s
  • targeted fp32_alu_128 run: 7,708.187 GFLOP/s
  • longer soak average: 7,705.76 GFLOP/s
  • soak total: 2,473,901,162,496 FP32 operations
  • soak min/max: 7,699.702–7,711.882 GFLOP/s
  • spread: 0.158064%
  • mismatch count: 0
  • violations: 0
  • measured memory-copy peak in the mixed matrix: 306,550.022 MiB/s

These measurements characterize bounded public-UWP/Developer Mode worker paths.

They do not establish unrestricted GPU access, GDK/Game Mode entitlement or access to the full theoretical 12.1 TFLOPS of the console.


Performance was no longer enough

Once the GPU path existed, XCP started putting explicit identity and execution contracts around it.

The worker loads precompiled shader bytecode and verifies that the shader matches the profile that was admitted for execution.

Conceptually:

workload
   ↓
admission
   ↓
expected shader identity
   ↓
actual installed shader
   ↓
match?
Enter fullscreen mode Exit fullscreen mode

A successful dispatch from the wrong shader is not a valid result.

This sounds obvious, but it changes the meaning of execution.

The GPU is no longer just a device receiving work.

It becomes one component in an evidence chain.


Moving XVM state through the GPU

XVM workloads can be represented using structured GPU resources.

Parameters, input data, program state and — for some profiles — microtrace data are exposed to the compute path as structured buffers.

The output lives in a GPU state buffer.

Then comes the part that became more important than the dispatch itself:

read the state back.

parameters / input / program
            ↓
      StructuredBuffers
            ↓
      compute shader
            ↓
        GPU state
            ↓
      staging buffer
            ↓
        CPU readback
            ↓
          decode
            ↓
          verify
Enter fullscreen mode Exit fullscreen mode

A successful Dispatch() is therefore not the end of the operation.

It is the middle.

Show the D3D readback path

The result state is exposed as an unordered access resource.

A CPU-readable staging buffer is created:

D3D11_BUFFER_DESC stagingDesc{};

stagingDesc.ByteWidth =
    stateWordCount * sizeof(uint32_t);

stagingDesc.Usage =
    D3D11_USAGE_STAGING;

stagingDesc.CPUAccessFlags =
    D3D11_CPU_ACCESS_READ;
Enter fullscreen mode Exit fullscreen mode

After execution:

context->CopyResource(
    stagingBuffer.get(),
    resultBuffer.get());

context->Flush();
Enter fullscreen mode Exit fullscreen mode

The worker also checks for device loss.

Then the staging buffer is mapped:

D3D11_MAPPED_SUBRESOURCE mapped{};

auto mapHr = context->Map(
    stagingBuffer.get(),
    0,
    D3D11_MAP_READ,
    0,
    &mapped);

std::memcpy(
    words.data(),
    mapped.pData,
    words.size() * sizeof(uint32_t));

context->Unmap(
    stagingBuffer.get(),
    0);
Enter fullscreen mode Exit fullscreen mode

Now the GPU state can be decoded and validated on the CPU side.


Every new capability exposed another limit

This became the recurring pattern.

Memory had a budget.

Long-running execution needed bounds.

GPU work needed admitted shapes.

Resident execution needed epochs.

GPU state needed readback.

SPMD execution introduced lane divergence.

Device loss needed to become an explicit state.

And eventually the GPU result itself needed an independent reference.

The limits were not merely obstacles.

They became part of the architecture.

limit discovered
      ↓
make it explicit
      ↓
build a contract around it
      ↓
measure again
Enter fullscreen mode Exit fullscreen mode

That is one of the reasons XCP became evidence-driven.

Instead of pretending a boundary did not exist, we tried to make it machine-readable.


Then we compared GPU execution with the CPU

For bounded classes of XVM execution, XCP can run a CPU reference and compare it with the GPU result.

The comparison includes more than output bytes.

It can include:

output
control state
fuel consumed
trap code
trap PC
recovery PC
trap count
Enter fullscreen mode Exit fullscreen mode

So the model becomes:

                XVM workload
                     │
              ┌──────┴──────┐
              │             │
              ▼             ▼
        CPU reference    GPU execution
              │             │
              │          readback
              │             │
              └──────┬──────┘
                     ▼
             differential check
                     ↓
                  evidence
Enter fullscreen mode Exit fullscreen mode

At that point the CPU is not simply a fallback.

For the admitted execution class, it becomes a reference against which the GPU can be checked.

Show the CPU/GPU differential checks

The runtime compares output:

auto outputMatches =
    input.cpuRun->output == gpu.run.output;
Enter fullscreen mode Exit fullscreen mode

Control state:

auto controlMatches =
    input.cpuRun->controlToken ==
    gpu.run.controlToken;
Enter fullscreen mode Exit fullscreen mode

Fuel:

auto fuelMatches =
    input.cpuRun->fuelConsumed ==
    gpu.run.fuelConsumed;
Enter fullscreen mode Exit fullscreen mode

And trap state:

auto trapMatches =
    input.cpuRun->trap.code ==
        gpu.run.trap.code &&

    input.cpuRun->trap.trapPc ==
        gpu.run.trap.trapPc &&

    input.cpuRun->trap.recoveryPc ==
        gpu.run.trap.recoveryPc &&

    input.cpuRun->trap.occurrenceCount ==
        gpu.run.trap.occurrenceCount;
Enter fullscreen mode Exit fullscreen mode

XCP also has bounded resident execution paths, epoch validation and SPMD lane-consistency checks.

If lanes that are required to remain synchronized diverge, the execution is rejected rather than promoted into a valid result.


The limits shaped the worker

There is a tendency to describe system constraints only as things preventing software from doing more.

That was not what happened here.

The constraints repeatedly forced the system to become more explicit.

A memory ceiling led to measured resource accounting.

Long-lived compute led to bounded execution.

GPU state led to readback.

Readback led to structural validation.

Multiple execution backends led to differential verification.

Device loss led to explicit quarantine and fallback state.

The sandbox itself became part of the design.

We did not remove the limits. We built the worker around measured limits.


From benchmark to execution system

The first experiments were naturally interested in performance.

Could meaningful GPU compute run?

How much memory was available?

How hard could the environment be pushed?

Those were necessary questions.

But eventually they became the less interesting ones.

The harder questions were:

Which program executed?

Which shader produced the result?

What resource boundary was observed?

Did memory remain intact?

Was execution bounded?

Did GPU state progress legally?

Did parallel lanes remain consistent?

Did the GPU agree with the CPU?

What exactly does the evidence allow us to claim?
Enter fullscreen mode Exit fullscreen mode

That is where XCP stopped looking like a benchmark project.

It became an execution system.


The Xbox is still a sandbox

None of this bypasses Xbox Developer Mode.

XCP does not turn the console into an unrestricted PC, and it does not claim full retail-console GPU entitlement.

The interesting part was precisely the opposite.

We kept hitting boundaries.

Then we measured them.

Then we built around them.

That is why the project gradually moved from:

Can this run?
Enter fullscreen mode Exit fullscreen mode

to:

What ran?

Under which boundary?

What evidence came back?

What can we safely claim from it?
Enter fullscreen mode Exit fullscreen mode

One important evidence boundary

The public XCP repository preserves the implementation lineage and historical physical-Xbox evidence.

The newly reconstructed development packages are different artifacts.

They remain:

NOT_TESTED_ON_XBOX
Enter fullscreen mode Exit fullscreen mode

until a new physical-device reproduction generates evidence for those exact packages.

A new binary does not automatically inherit an older binary's execution proof.

That distinction is intentional.

It is the same principle that grew out of all those early probes:

a claim should not be stronger than the measurement behind it.


What this work changed

The difficult part turned out not to be making the hardware do something.

It was constructing a system that could explain later, with enough precision, what the hardware actually did.

For XCP, that path became:

map the sandbox
       ↓
measure resources
       ↓
find a boundary
       ↓
push further
       ↓
admit bounded work
       ↓
execute
       ↓
read back state
       ↓
compare
       ↓
record evidence
Enter fullscreen mode Exit fullscreen mode

That is now one of the central ideas behind the whole project.

Execution is an event.

Evidence is what makes the event usable.


XCP is open source

GitHub:

https://github.com/Daniele-Cangi/xcp-xbox

XCP Studio:

https://xcpstudio.com/

Relevant implementation areas include:

reference/xcompute-probe/src/XComputeProbe/probes/ProbeRunner.cpp

reference/xcompute-probe/src/XComputeProbe/runtime/WorkerProcessTopologyMemoryProbe.cpp

reference/xcompute-probe/src/XComputeProbe/runtime/WorkerPrecompiledD3DCommandRuntime.cpp

reference/xcompute-probe/src/XComputeProbe/runtime/WorkerGpuXvmExecutor.cpp

reference/xcompute-probe/src/XComputeProbe/runtime/WorkerGpuXvmDifferentialRuntime.cpp
Enter fullscreen mode Exit fullscreen mode

Previous chapters in the XcpStudio series cover the origin of XVM, the open-source XCP platform and the Studio workflow.

AI can build the workload. Xbox can execute it. XCP's job is to make the result defensible.

Top comments (0)