When a material graph grows in Unreal Engine 5, it is tempting to ask: would writing HLSL directly make it faster? Would implementing the shader in C++ make optimization easier?
Moving a shader into code is not, by itself, an optimization. Material Editor nodes already become HLSL. Epic also warns that a Custom expression can prevent Unreal's material-level constant folding and produce more instructions than an equivalent graph built with standard nodes. Material concepts · Custom expressions
This guide separates five questions that are often mixed together: GPU execution time, shader permutation costs, first-use pipeline hitches, choosing an authoring layer, and adding a compute shader with C++ and HLSL.
The engine baseline is Unreal Engine 5.8.2, as in the Japanese edition dated August 31, 2026. Version-sensitive references were rechecked for this English edition on September 30, 2026. This is a documentation baseline, not a compilation or execution result for UE 5.8.2. Epic's 5.8.2 hotfix announcement is dated August 25, 2026. Release announcement
The examples primarily target Windows, DirectX 12, and Shader Model 6. Mobile, consoles, and VR have different compiler, precision, bandwidth, and GPU constraints. Make the final decision on the actual target hardware.
Substrate: New projects use Substrate by default from UE 5.7 onward, but Epic's UE 5.8 documentation still labels it Beta. Existing projects upgraded to UE 5.7 or later retain the non-Substrate path unless they explicitly opt in. The principles below apply to both paths; the code does not depend on a particular Substrate Slab. Substrate status · Substrate overview
Start with the simplest layer that meets your needs
Do not begin by moving everything to HLSL or C++. Choose the layer that solves the actual problem.
| Goal | Start with | When to move lower |
|---|---|---|
| Ordinary surface shading | Material nodes and Material Functions | An operation is unavailable, or repeated logic is becoming difficult to maintain |
| Many parameter variations | Material Instances | Split the parent when static-switch combinations become unmanageable |
| A small custom formula | Custom Material Expression | Built-in nodes are insufficient, and generated HLSL plus measurements justify it |
| Full-screen color processing | Post Process Material | You need arbitrary UAVs or buffers, custom pass ordering, or compute dispatch |
| A compute shader | Global Shader and RDG | You need reads, writes, or dispatch control outside the material model |
| Drawing meshes in a custom pass | Mesh Material Shader or a custom mesh pass | Renderer integration is necessary and its maintenance cost is acceptable |
| Something is simply “slow” | Profiling | Change the implementation after identifying the bottleneck and affected pass |
Material nodes are not a simplified substitute for “real” shaders. Expressions visually represent shader operations, and the graph is translated into HLSL. In the Material Editor, Window > Shader Code > HLSL Code opens a read-only view of the generated source. It is not the final GPU-native instruction stream. Material concepts · Material Editor UI
The useful comparison is therefore not nodes versus code. It is the shader compiled for the target platform, multiplied by the pixels, layers, and passes that execute it. Use generated HLSL to investigate the implementation, then use target-platform statistics and GPU timings to evaluate performance.
Separate three kinds of shader cost
1. Runtime GPU work
Instruction count is only one part of GPU cost. Screen coverage, overdraw, texture access and cache behavior, arithmetic, special functions, loops, branch coherence, render-target traffic, and vertex work all matter. A material can also participate in several passes: depth, shadows, base pass, velocity, and representations used by Lumen. Rendering optimization · Shader development
A hypothetical 100-instruction shader applied once to a small prop is not equivalent to the same shader running across several full-screen translucent layers. The arithmetic is similar; the amount of work is not.
Before changing a formula, ask where it runs, how often it runs, and what it reads and writes.
2. Compilation and shader permutations
An Unreal material does not correspond to one shader binary. Rendering passes, vertex factories, feature levels, material quality, static parameters, and target platforms produce different variants. Shader development
A Static Switch can remove its unused branch at compile time, which is useful for runtime performance. The trade-off is a separate compiled variation for each required static-parameter combination. Ten independent Boolean switches allow up to 2^10 = 1,024 combinations before other permutation dimensions are considered. This is a theoretical combination count, not a claim that every project compiles all 1,024. Material parameters
A parent material that includes every possible feature can increase cooking time, Derived Data Cache (DDC) work, continuous-integration costs, shader storage, and loading overhead. Audit the combinations actually used rather than counting only parent assets. Material Analyzer
3. First-use PSO hitches
Having shader binaries available does not mean every Pipeline State Object (PSO) is ready. Creating a required pipeline on first use can still cause driver-side work and a visible hitch. Removing a few pixel-shader instructions does not directly solve that problem. PSO precaching
UE's PSO Precaching system collects candidate pipelines and compiles them asynchronously. Epic's UE 5.8 documentation includes eligible Global Compute Shader permutations in precaching. You still need to verify platform support, wait for important outstanding requests during loading, and watch for missed or late PSOs during gameplay. PSO precaching
These are different optimization objectives. Decide whether the goal is less GPU time per frame, less compilation and storage, or fewer first-use hitches before choosing a technique.
Profile outside the Material Editor
1. Establish whether the GPU is the bottleneck
Start with stat unit and compare Game, Draw, GPU, and Frame times. When another stage dominates, reducing material instructions may not improve the overall frame time. The largest reported GPU value is a clue rather than a complete diagnosis: timing can include waits. Rendering optimization
Use stat gpu and ProfileGPU to identify the expensive passes, such as Base Pass, Translucency, Post Processing, Lumen, or shadows. Timing Insights in Unreal Insights can show CPU and GPU work on a shared timeline. Rendering optimization · Timing Insights
Keep the comparison conditions fixed: resolution and screen percentage; dynamic resolution; VSync and frame limits; camera, visible objects, lights, and effects; scalability and material quality; RHI and shader model; and shader/PSO warm-up state.
Save a reproducible scene and camera. Editor-only work and additional viewports can distort a comparison, so repeat the final test in an appropriate Development or Test build on the target platform, with the required profiling tools available.
2. Read Material Stats and Platform Stats
Material Stats provide a quick starting point for instruction-count comparisons. Platform Stats show results for the selected rendering targets. Compare the same rendering API and shader model. An SM5 count and an SM6 count are not automatically comparable because the compilation path and reported metrics can differ. Material Editor UI
Treat these statistics as diagnostic evidence, not as a stopwatch. They do not capture screen coverage, all overdraw, memory-bandwidth pressure, or the effect on the rest of the frame.
3. Use Shader Complexity and Quad Overdraw as clues
Shader Complexity helps locate expensive regions, particularly translucent effects. Epic explicitly notes that instruction counts do not assign equivalent execution cost to every operation: texture fetches and simple arithmetic can have very different costs. Loops that are not unrolled can also be poorly represented by the visualization. Viewport modes
A useful sequence is to find candidates with Shader Complexity, inspect overlap with Quad Overdraw, measure the affected pass with ProfileGPU, and A/B test the material under identical conditions. Use a GPU debugger such as RenderDoc or PIX when you need to inspect individual events or shader behavior more closely. Viewport modes · Rendering optimization · Shader debugging
A red region is a reason to investigate, not an instruction to rewrite the material in HLSL.
Material optimizations worth investigating first
Check blend mode, screen coverage, and overdraw
Opaque rendering can often use depth rejection effectively. Translucency processes overlapping layers, and masked materials still need to evaluate their cutout. Several layers of foliage, fences, or large particle cards can therefore become expensive even when each individual material looks modest. Rendering optimization · Viewport modes
Before changing the math, ask whether the effect needs translucency, whether its polygon covers large transparent areas, whether too many particles overlap, whether distant LODs can drop layers, and whether several full-screen effects could be reduced or combined.
Reducing the number of shaded pixels or overlapping layers can matter more than removing a few arithmetic instructions. Measure both options instead of assuming that the shorter formula wins.
Evaluate texture access and arithmetic separately
Neither “replace textures with math” nor “bake all math into textures” is a universal rule.
Complex, static, multi-octave noise evaluated per pixel is a candidate for baking. A simple gradient may not justify adding a large texture, streaming requirements, and another sample. Consider sample count, mip selection, resolution, compression, locality, memory traffic, and whether the target workload is limited by arithmetic or bandwidth. Rendering optimization
Channel packing can reduce separate texture accesses when channels are used together, but the grouping matters. Do not assume that Base Color and scalar masks such as roughness, metallic, and ambient occlusion can share arbitrary compression and color-space settings. Preserve the intended data interpretation, and check whether unrelated channels are being fetched unnecessarily. Texture masks
The decision belongs to the measured workload, not to the appearance of the material graph.
Move suitable UV calculations to the vertex stage
Customized UVs let you evaluate UV operations in the vertex shader and interpolate the results for the pixel shader. This can help when there are substantially fewer vertices than shaded pixels. Customized UVs
An affine operation is a typical candidate:
UV * Tiling + Offset
Nonlinear operations such as sin, cos, length, or squaring generally do not survive interpolation unchanged. Moving them to the vertex stage can make the result depend on mesh density. The cost of extra interpolants also matters. Customized UVs
Do not assume that the vertex stage is always cheaper, particularly for dense geometry or a Nanite rendering path. Compare the actual target path, both visually and in GPU time.
Material Instances do not automatically remove instructions
Changing Scalar, Vector, or Texture Parameters through a Material Instance or Material Instance Dynamic (MID) normally retains the parent's shader logic. The advantage is reuse and controlled variation, not automatic removal of GPU work. Static Switch Parameters are different: they select compiled variations that can omit unused branches. Material instances · Material parameters
As a design rule, use runtime parameters for colors, strengths, thresholds, and animation speeds. Consider a static switch when removing an entire feature branch is valuable. Separate genuinely different material families instead of building one universal parent, and inspect real static-override combinations with Material Analyzer. Material Analyzer
Also review project-level rendering features. Settings for features such as stationary skylights, low-quality lightmaps, whole-scene point-light shadows, or path tracing can affect which material shaders must be compiled. Disable support only after establishing that the project does not need it. Fewer permutations primarily reduce compilation and storage; they do not guarantee a cheaper visible pixel shader. Rendering settings
Use Material Functions to improve maintainability
Material Functions organize and reuse graph logic. They are not a promise that the GPU executes a shared computation once for every caller. Evaluate the resulting compiled material. Material Functions
A practical boundary is to separate responsibilities with Material Functions, then introduce Custom HLSL only where standard nodes are insufficient or the equivalent graph becomes difficult to maintain. This keeps more of the workflow accessible to artists and preserves graph previews and material integration.
Inspect Substrate statistics and the GBuffer format
Substrate represents materials using Slabs and can simplify their composition to fit project and platform budgets. Use the Material Editor's Substrate statistics to inspect topology, enabled features, and simplification. Substrate overview
Blendable GBuffer is the default for most new Substrate-enabled projects and prioritizes predictable performance. Some templates, including Automotive and Architectural, choose Adaptive GBuffer instead. Adaptive supports richer materials but increases cost and cooking work. Compare the same GBuffer format and closure configuration, not merely whether Substrate is enabled. Substrate overview
What “writing a shader in code” means in UE5
1. Custom Material Expression
A Custom expression replaces part of a material graph with HLSL. Unreal still handles material integration, parameters, lighting, and the relevant rendering passes. It is usually the smallest step beyond ordinary graph authoring. Custom expressions
Typical reasons to use it include a missing formula, a readable fixed-count loop, bit operations, or a calculation that becomes unwieldy as nodes.
Keep the boundaries clear. Unreal cannot perform all of its material-level constant folding inside a Custom expression. Inputs are function-local because the code is wrapped in a generated function. Valid HLSL can still encounter backend-specific limitations. None of this means that the downstream shader compiler performs no optimization. Custom expressions
World-space calculations also need care. UE5's Large World Coordinates (LWC) rendering types and translated world space exist to manage precision. Blindly reducing large coordinates to float3 can introduce errors far from the origin. LWC rendering
A Custom expression remains part of a material. It does not give you arbitrary render-pass scheduling or unrestricted UAV access. Choose it for a concrete functionality, readability, or reuse benefit—not because code looks inherently faster.
2. .ush, .usf, and Global Shaders
For a Global Shader, you place HLSL in shader files and register a shader type from C++. Global Shaders are not tied to a particular material or mesh and are used for tasks such as full-screen processing, compute work, and clears. Adding Global Shaders · Shader development
The division of responsibility is important:
| Layer | Responsibility |
|---|---|
| HLSL | Vertex, pixel, or compute operations executed by the GPU |
| C++ | Shader registration, compile conditions, parameters, resources, passes, and dispatch |
| Render Dependency Graph (RDG) | Resource dependencies and lifetimes, transitions, scheduling, and unused-pass culling |
C++ configures the GPU work; it does not become the per-pixel GPU program. RDG manages work whose resource usage is declared to the graph. Render Dependency Graph
3. Mesh Material Shaders and custom mesh passes
Rendering meshes in a custom pass while using their materials, consuming vertex-factory data, or integrating closely with the base pass requires more than a stand-alone Global Shader. This territory includes Mesh Material Shaders, mesh pass processors, renderer integration, and suitable extension points such as Scene View Extensions. Shader development · Mesh drawing pipeline
The integration has a larger maintenance cost across engine upgrades. For an ordinary screen effect, first consider a Post Process Material. For independent compute work, first consider a Global Shader with RDG.
Custom HLSL example: a rounded-rectangle mask
This example creates a rounded rectangle centered at UV = (0.5, 0.5). All dimensions use ordinary UV units, not a coordinate system scaled to -1..1.
Configure the Custom expression inputs as follows:
| Input | Type | Meaning | Example |
|---|---|---|---|
UV |
float2 |
Coordinates in the 0..1 UV domain |
TextureCoordinate |
HalfSize |
float2 |
Half-width and half-height from the center | (0.40, 0.25) |
Radius |
float |
Corner radius, in the same UV units | 0.08 |
Feather |
float |
Minimum edge-transition width, in UV units | 0.002 |
Set Output Type to CMOT Float 1. Custom expression setup
float2 safeHalfSize = clamp(
abs(HalfSize),
float2(1e-5, 1e-5),
float2(0.5, 0.5));
float safeRadius = clamp(
Radius,
0.0,
min(safeHalfSize.x, safeHalfSize.y));
float2 p = abs(UV - 0.5);
float2 q = p - safeHalfSize + safeRadius;
float distanceToEdge =
length(max(q, 0.0)) +
min(max(q.x, q.y), 0.0) -
safeRadius;
float aa = max(
max(abs(Feather), 1e-5),
fwidth(distanceToEdge));
return saturate(0.5 - distanceToEdge / aa);
With HalfSize = (0.40, 0.25), the rectangle has a width of 0.80 and a height of 0.50. Those dimensions describe the zero-distance contour, where the returned mask is 0.5; the feathered transition extends to either side of that contour.
The accepted inputs are finite numbers. Each half-extent is converted to its absolute value and clamped to 1e-5..0.5. The radius is clamped to 0..min(safeHalfSize.x, safeHalfSize.y), so a negative radius becomes zero and an oversized radius cannot exceed the smaller sanitized half-extent. A negative Feather uses its magnitude, with a numerical floor and screen-space derivatives determining the final transition width.
Use this example in the pixel-shader path because it uses fwidth. Connect the result to an appropriate blend factor or Opacity Mask. A masked material still applies its clipping threshold; producing a smooth scalar does not by itself make masked rendering behave like alpha blending. Apply aspect correction when a circular corner in UV space would otherwise appear stretched on screen.
This example is not a speed claim. An equivalent graph can be built from standard nodes. Compare the graph version and Custom HLSL version using the same Platform Stats target, generated HLSL, and GPU profiling conditions. The Custom version may be useful because it is easier to maintain, even when it is not faster.
For shared code, Custom expressions expose Include File Paths and Additional Defines. Mapping a plugin shader directory to a virtual path lets you keep reusable functions in a .ush file. Avoid putting unrelated code into one universal include: changes to a widely included file can trigger a much broader recompile. Custom expressions · Shader development
Add a Global Compute Shader through a plugin
The following example shows the minimal pass implementation for writing a UV gradient into an RDG texture. In a game, add it from a render-thread integration point that already receives an FRDGBuilder, such as an appropriate Scene View Extension callback or another renderer extension. It is not a game-thread RHI example. Adding Global Shaders · Render Dependency Graph
Validation and integration limits: These snippets illustrate shader registration, RDG resources, and dispatch. They do not form a complete screen effect. You must supply the rendering callback and a downstream consumer that composites the texture or transfers the result to an appropriate external render target. The C++ and HLSL have not been compiled or executed on UE 5.8.2 in this authoring environment. Compilation on your engine branch and execution on each target RHI remain required checks.
Folder layout
Plugins/ShaderLab/
├─ ShaderLab.uplugin
├─ Shaders/
│ └─ Private/
│ └─ WriteGradient.usf
└─ Source/
└─ ShaderLab/
├─ ShaderLab.Build.cs
├─ Public/
│ └─ WriteGradientShader.h
└─ Private/
├─ ShaderLabModule.cpp
└─ WriteGradientShader.cpp
Load the module at PostConfigInit so shader registration and the virtual shader-path mapping are available early enough. Adding Global Shaders
{
"FileVersion": 3,
"Version": 1,
"VersionName": "1.0.0",
"FriendlyName": "ShaderLab",
"CanContainContent": false,
"Modules": [
{
"Name": "ShaderLab",
"Type": "Runtime",
"LoadingPhase": "PostConfigInit"
}
]
}
Build.cs
The public pass header exposes RDG types, so RenderCore is a public dependency. The other dependencies below support the plugin implementation. Public include visibility and exported binary symbols are separate concerns; both will matter when another module calls the pass function. Unreal modules · Module API specifiers
using UnrealBuildTool;
public class ShaderLab : ModuleRules
{
public ShaderLab(ReadOnlyTargetRules Target) : base(Target)
{
PCHUsage = PCHUsageMode.UseExplicitOrSharedPCHs;
PublicDependencyModuleNames.AddRange(
new[]
{
"Core",
"RenderCore"
});
PrivateDependencyModuleNames.AddRange(
new[]
{
"CoreUObject",
"Engine",
"Projects",
"Renderer",
"RHI"
});
}
}
Map the shader directory
// ShaderLabModule.cpp
#include "Interfaces/IPluginManager.h"
#include "Misc/Paths.h"
#include "Modules/ModuleManager.h"
#include "ShaderCore.h"
class FShaderLabModule final : public IModuleInterface
{
public:
virtual void StartupModule() override
{
const TSharedPtr<IPlugin> Plugin =
IPluginManager::Get().FindPlugin(TEXT("ShaderLab"));
check(Plugin.IsValid());
const FString ShaderDirectory =
FPaths::Combine(Plugin->GetBaseDir(), TEXT("Shaders"));
AddShaderSourceDirectoryMapping(
TEXT("/ShaderLab"),
ShaderDirectory);
}
virtual void ShutdownModule() override
{
}
};
IMPLEMENT_MODULE(FShaderLabModule, ShaderLab)
The physical directory Plugins/ShaderLab/Shaders can now be addressed through the virtual shader path /ShaderLab. The registration below uses that virtual path, not a machine-specific filesystem path. Shaders in plugins
Write the HLSL
// Shaders/Private/WriteGradient.usf
RWTexture2D<float4> OutputTexture;
int2 OutputSize;
[numthreads(8, 8, 1)]
void MainCS(uint3 DispatchThreadId : SV_DispatchThreadID)
{
if (DispatchThreadId.x >= (uint)OutputSize.x ||
DispatchThreadId.y >= (uint)OutputSize.y)
{
return;
}
float2 uv =
(float2(DispatchThreadId.xy) + 0.5) /
float2(OutputSize);
OutputTexture[DispatchThreadId.xy] =
float4(uv, 0.0, 1.0);
}
Each thread writes one texel. The half-texel offset evaluates the gradient at texel centers. The bounds check protects the final thread groups when the texture dimensions are not multiples of eight. The caller must supply a texture with positive dimensions.
Register the Global Shader and add the RDG pass
A public function called from another module needs the export macro SHADERLAB_API. Putting a header in Public makes it available for inclusion; that alone does not export the function across a modular-build DLL boundary. Module API specifiers
// Public/WriteGradientShader.h
#pragma once
#include "RenderGraphFwd.h"
SHADERLAB_API void AddWriteGradientPass(
FRDGBuilder& GraphBuilder,
FRDGTextureRef OutputTexture);
When a consumer calls the function only from its .cpp implementation, add the dependency in that consumer module's Build.cs:
PrivateDependencyModuleNames.Add("ShaderLab");
Use PublicDependencyModuleNames instead when the consumer's public interface itself requires the ShaderLab header. Likewise, expose the appropriate owning-module dependencies for any engine types used by that public interface. Unreal modules
The implementation registers the shader and schedules its work:
// WriteGradientShader.cpp
#include "WriteGradientShader.h"
#include "DataDrivenShaderPlatformInfo.h"
#include "GlobalShader.h"
#include "RenderGraphBuilder.h"
#include "RenderGraphUtils.h"
#include "ShaderParameterStruct.h"
class FWriteGradientCS final : public FGlobalShader
{
public:
DECLARE_GLOBAL_SHADER(FWriteGradientCS);
SHADER_USE_PARAMETER_STRUCT(FWriteGradientCS, FGlobalShader);
BEGIN_SHADER_PARAMETER_STRUCT(FParameters, )
SHADER_PARAMETER(FIntPoint, OutputSize)
SHADER_PARAMETER_RDG_TEXTURE_UAV(
RWTexture2D<float4>,
OutputTexture)
END_SHADER_PARAMETER_STRUCT()
static bool ShouldCompilePermutation(
const FGlobalShaderPermutationParameters& Parameters)
{
return IsFeatureLevelSupported(
Parameters.Platform,
ERHIFeatureLevel::SM5);
}
};
IMPLEMENT_GLOBAL_SHADER(
FWriteGradientCS,
"/ShaderLab/Private/WriteGradient.usf",
"MainCS",
SF_Compute);
void AddWriteGradientPass(
FRDGBuilder& GraphBuilder,
FRDGTextureRef OutputTexture)
{
check(OutputTexture);
const FIntPoint OutputSize = OutputTexture->Desc.Extent;
FWriteGradientCS::FParameters* Parameters =
GraphBuilder.AllocParameters<
FWriteGradientCS::FParameters>();
Parameters->OutputSize = OutputSize;
Parameters->OutputTexture =
GraphBuilder.CreateUAV(
FRDGTextureUAVDesc(OutputTexture));
TShaderMapRef<FWriteGradientCS> ComputeShader(
GetGlobalShaderMap(GMaxRHIFeatureLevel));
FComputeShaderUtils::AddPass(
GraphBuilder,
RDG_EVENT_NAME("ShaderLab.WriteGradient"),
ComputeShader,
Parameters,
FComputeShaderUtils::GetGroupCount(
OutputSize,
FIntPoint(8, 8)));
}
The SM5 feature-level check is a compile filter, not a claim that every admitted platform has been tested. The intended DX12/SM6 target is not made SM5-only by this minimum-feature-level check. Keep the shader map and compile policy consistent with the rendering path used by your integration. Shader development
The output texture must be created by, or registered with, the same FRDGBuilder and support the required UAV usage. This example expects a non-multisampled 2D texture at mip zero, with a format compatible with its float4 UAV. It does not handle texture arrays or arbitrary mip views. The following caller-side fragment assumes an existing graph and a positive FIntPoint OutputSize; it is not a complete rendering callback:
const FRDGTextureDesc OutputDesc =
FRDGTextureDesc::Create2D(
OutputSize,
PF_FloatRGBA,
FClearValueBinding::Black,
TexCreate_ShaderResource | TexCreate_UAV);
FRDGTextureRef OutputTexture =
GraphBuilder.CreateTexture(
OutputDesc,
TEXT("ShaderLab.GradientOutput"));
AddWriteGradientPass(GraphBuilder, OutputTexture);
Within the graph, another pass can read the result as an SRV and composite it. Alternatively, use the appropriate RDG extraction or external-resource path, such as QueueTextureExtraction, to retain the underlying resource after graph execution. Do not retain the FRDGTextureRef itself beyond its graph lifetime. A pass with no live output consumer may be culled; allocating a texture is not the same as displaying it. Render Dependency Graph
Splitting the caller into another class or module does not change those lifetime rules. Keep the builder and RDG handles within the graph's lifetime. Do not capture game-thread UObject state or raw pointers without a valid render-thread ownership and synchronization design. Render Dependency Graph
What to optimize in this example
The five-argument FComputeShaderUtils::AddPass above schedules an ordinary compute pass on the graphics pipe. Writing a compute shader does not automatically run it on the Async Compute pipe or overlap it with graphics work. The pass-flag overload is separate, and ERDGPassFlags::Compute and ERDGPassFlags::AsyncCompute describe different execution pipes. AddPass API · Pass flags
For this pass, investigate group size, texture format, UAV traffic, unnecessary dispatch area, graph dependencies and lifetimes, permitted shader permutations, and first-use PSO behavior. An 8 x 8 group is a starting point, not a universal optimum. Register use, shared memory, wave size, and access patterns can change the result. Render Dependency Graph · Shader development
Only when testing Async Compute as an optional A/B variant, replace the shown call with the overload that accepts pass flags and specify ERDGPassFlags::AsyncCompute. That is a candidate to measure, not a default improvement. AddPass API
An async flag does not guarantee useful overlap. Resource dependencies, queue synchronization, RHI and hardware support, and competition for compute or memory bandwidth can eliminate the benefit or make the frame slower. Inspect the actual schedule and overlap with RDG Insights and GPU profiling tools. Render Dependency Graph
Decide between a Custom expression and a Global Shader
Stay with a Custom expression when the material model fits
A Custom expression is a good fit when you want to compute material inputs such as Base Color, Normal, Roughness, or Opacity Mask; keep Unreal's normal material and mesh-pass integration; express a formula more clearly; and continue controlling values through Material Instances. Custom expressions · Material parameters
You do not need a custom rendering pass merely because a material contains HLSL.
Use a Global Shader when you need independent GPU work
Move to a Global Shader when the operation needs writes to resources such as RWTexture2D or RWBuffer, arbitrary compute dispatch dimensions, a material-independent full-screen operation, or explicit RDG dependencies between intermediate textures and buffers. Adding Global Shaders · Render Dependency Graph
The benefit is control over GPU work and resources, not an automatic reduction in execution time.
Extend the renderer when the mesh pipeline itself must change
Custom mesh drawing, vertex-factory-specific data, deep base-pass or shadow integration, Nanite integration, new shading models, and new GBuffer representations belong to a more tightly coupled layer. Some changes require engine modifications rather than a self-contained plugin. Mesh drawing pipeline · Shader development
Budget for engine-upgrade maintenance, validation on every target RHI, and pipeline-cache coverage. Do not choose this layer merely to make an ordinary material graph look shorter.
Design permutations and PSOs deliberately
The same permutation discipline applies to code-authored shaders and materials.
Use ShouldCompilePermutation to exclude unsupported or unnecessary platforms. When using a permutation domain, avoid generating combinations that runtime code cannot reach. A runtime parameter can be preferable to another compile-time define when the variation does not justify a separate shader. Conversely, a small, controlled set of permutations may be worthwhile when it removes substantial runtime work. Shader development
On the material side, Material Analyzer groups instances with matching static overrides and helps identify candidates for a shared intermediate parent. Review the interaction of project features, Static Switches, material quality, and feature levels rather than treating each setting in isolation. Material Analyzer · Rendering settings
For PSOs, verify distinct milestones: the shader was cooked; the pipeline is covered by precaching; important requests finish before gameplay needs them; actual play does not reveal missed or late pipelines; and any fallback material or delayed draw behavior is acceptable. PSO precaching
A smooth editor session is not proof of a smooth first launch. Test the packaged target with an appropriately clean cache state. Shader compilation and DDC behavior, pipeline precaching, and the graphics driver's cache are related but are not the same cache or the same test.
Recompile and debug during development
For work on .usf and .ush files, enable shader development diagnostics in ConsoleVariables.ini:
r.ShaderDevelopmentMode=1
This enables shader-development logging and support for retrying compilation after errors. After saving, recompileshaders changed or Ctrl+Shift+. recompiles changed shaders. A change to a widely included file can invalidate a large set of dependent shaders, so keep shared includes focused. Shader development
For GPU debugging, use the platform-appropriate symbol settings, including r.Shaders.Symbols and r.Shaders.WriteSymbols where supported by your workflow. Symbol generation affects cooking time and storage. Enable it deliberately for the platforms and investigations that need it rather than assuming it is free. Shader debugging workflows
A practical optimization checklist
Before changing anything
- [ ] Fix the target hardware, RHI, shader model, and resolution.
- [ ] Express the frame-rate target as a time budget in milliseconds.
- [ ] Use
stat unitand GPU profiling to establish the bottleneck. - [ ] Identify the affected pass with
ProfileGPU. - [ ] Save the comparison scene, camera, quality settings, and warm-up conditions.
When changing a material
- [ ] Check blend mode, screen coverage, and overdraw before editing the formula.
- [ ] Measure GPU time; do not rely on Shader Complexity alone.
- [ ] Evaluate texture samples, arithmetic, and affected pixel count separately.
- [ ] Test whether suitable UV calculations can move to Customized UVs without changing the image.
- [ ] Estimate permutation growth before adding a Static Switch.
- [ ] Use Material Instances and Material Functions for controlled reuse and maintainability.
- [ ] Compare Custom HLSL against equivalent standard nodes on the same target.
- [ ] For Substrate, compare statistics, simplification, and GBuffer format.
When adding a Global Shader
- [ ] Explain why a Material or Post Process Material cannot meet the requirement.
- [ ] Load shader registration and virtual-path mapping early enough.
- [ ] Match C++ parameter names and types with the HLSL declarations.
- [ ] Export cross-module functions or types with the correct
MODULENAME_APImacro. - [ ] Declare consumer module dependencies at the correct public/private boundary.
- [ ] Keep RDG textures and buffers within the lifetime of their graph.
- [ ] Match resource usage and format to the required UAV, SRV, or render-target access.
- [ ] Check dispatch bounds and use positive resource dimensions.
- [ ] Restrict shader compilation with
ShouldCompilePermutation. - [ ] Name GPU events so the pass can be identified in profiling tools.
- [ ] Distinguish ordinary compute from Async Compute and measure actual overlap for the async variant.
- [ ] Verify PSO coverage and first use with clean caches.
- [ ] Compile, inspect the output, and measure performance on every target RHI and device.
Conclusion
A useful order of investigation is:
Measure → reduce shaded area and overdraw → simplify material structure → use Custom HLSL where justified → add Global Shaders or renderer extensions when the requirements demand them.
Material graphs already become HLSL. Moving work into HLSL or C++ does not automatically make it faster. Start with screen coverage, translucency, texture access, unused features, and unnecessary static variations. Material concepts · Rendering optimization
Use Custom expressions for material-local formulas that benefit from code. Use Global Shaders and RDG for independent resources and dispatch. Use mesh-pipeline integration when the rendering architecture actually needs to change.
The right layer is not the lowest-level one. It is the simplest layer that meets the requirement, can be measured on the target platform, and remains maintainable by the team.
References
The following links point to Epic's primary documentation and release announcement. Version-dependent behavior should be checked against the engine branch used by the project.
- Unreal Engine 5.8.2 Hotfix Released
- Essential Unreal Engine Material Concepts
- Unreal Engine Material Editor UI
- Custom Material Expressions
- Material Parameter Expressions
- Material Instances
- Material Functions
- Using Texture Masks
- Customized UVs
- Material Analyzer
- Viewport Modes
- Guidelines for Optimizing Rendering for Real-Time
- Rendering Project Settings
- Timing Insights
- Shader Development
- Adding Global Shaders
- Module API Specifiers
- Unreal Engine Modules
- Overview of Shaders in Plugins
- FComputeShaderUtils::AddPass
- ERDGPassFlags
- Render Dependency Graph
- Mesh Drawing Pipeline
- PSO Precaching
- Shader Debugging Workflows
- Large World Coordinates Rendering
- Substrate Materials: feature status
- Overview of Substrate Materials
Top comments (0)