18 KiB
Shader Architecture Patterns for Node-Based Image Compositing
Research conducted 2026-07-13.
1. Single Shader vs Multiple Shaders
How Professional Compositors Handle It
Blender Compositor (GPU backend) — The most relevant case study:
- Blender's GPU compositor collapses multiple connected nodes into a "compile unit" and generates a single compute shader per unit.
- The
ShaderOperationclass iterates through acompile_unit_(a set of nodes) and links their GLSL logic into one shader:source/blender/compositor/intern/shader_operation.cc:122-135. - Simple per-pixel operations (Math, Color Mix, Invert, etc.) are fused into a single pass. Operations that can't be expressed as shaders fall back to
MultiFunctionProcedureOperationon CPU. - Key insight: Blender uses a hybrid approach — fuse what you can into single shaders, fall back to separate passes for complex operations (blur, glare, convolution).
Natron — CPU-based compositor using OpenFX plugins:
- Each node is a separate processing unit (separate plugin call).
- Multi-threaded tile-based processing per node.
- Not GPU-accelerated; no shader fusion.
DaVinci Resolve / Fusion — Proprietary:
- Uses a node graph where each node can have internal multi-pass processing.
- Fusion's "Flow Region" system groups nodes for optimization.
- Effectively separate shaders per node, with internal optimization.
Recommended Approach for Prism
Use separate shaders per node, with optional fusion of simple nodes. Rationale:
- Nodes in a compositing graph have diverse operations (blur vs. blend vs. color grade). An uber-shader would have massive register pressure and poor occupancy.
- Simple per-pixel operations (math, color mix, gamma) can be fused into chains as an optimization.
- Complex operations (blur, convolutions, warps) need their own shader passes anyway.
2. Texture Ping-Ponging
The Pattern
Texture ping-ponging is the fundamental technique for chaining GPU image operations:
- Allocate two textures (A and B) at the target resolution.
- Bind texture A as input, render to texture B.
- Swap: bind texture B as input, render to texture A.
- Repeat for as many passes as needed.
Pass 1: Read(A) → Write(B) [e.g., blur]
Pass 2: Read(B) → Write(A) [e.g., color grade]
Pass 3: Read(A) → Write(B) [e.g., blend]
Final: Display(B)
How It Works in Practice
WebGL/Fragment Shader approach (from multiple sources):
- Create Framebuffer Objects (FBOs) with texture attachments.
- Bind FBO → render fullscreen quad → output goes to texture.
- Bind different FBO or default framebuffer → read from that texture.
Vulkan approach:
- Use
VkImageobjects as both sampler inputs and render targets. - Between passes, issue a pipeline barrier (
VK_PIPELINE_STAGE_FRAGMENT_SHADER_BIT → VK_PIPELINE_STAGE_COLOR_ATTACHMENT_OUTPUT_BIT). - Manage image layouts:
SHADER_READ_ONLY_OPTIMAL→COLOR_ATTACHMENT_OPTIMAL→SHADER_READ_ONLY_OPTIMAL.
Metal approach (from Kosikowski's article):
- Compute shaders read from
inTextureand write tooutTexture. - Swap the texture references between passes.
Important Considerations
- Image layout transitions are critical in Vulkan. Each pass requires the texture to be in the correct layout.
- Load/store ops: For intermediate textures, use
VK_ATTACHMENT_LOAD_OP_DONT_CAREandVK_ATTACHMENT_STORE_OP_DONT_CAREwhen contents aren't needed — saves bandwidth. - Resolution management: Different nodes may operate at different resolutions. The compositor must manage a texture pool and handle up/downsampling.
- On tile-based GPUs (mobile): Multiple passes that write/read intermediate textures to external memory is expensive. Use Vulkan subpasses or
VK_KHR_dynamic_rendering_local_readto keep data on-chip.
3. Shader Composition Strategies
3a. Runtime Shader Generation
Blender's approach (most relevant):
- The compositor has a
gpu_shader_compositor_code_generation.glsllibrary. ShaderOperationgenerates GLSL code by iterating through a compile unit's nodes and concatenating their shader code contributions.- The generated code is compiled via Blender's
GPUMaterialsystem. - Node settings are passed as UBOs; images are bound as
image2D/sampler2D.
Godot's compositor approach:
- Uses a template + injection pattern:
const template_shader = """ #version 450 layout(local_size_x = 8, local_size_y = 8, local_size_z = 1) in; layout(rgba16f, set = 0, binding = 0) uniform image2D color_image; void main() { // ... boilerplate ... vec4 color = imageLoad(color_image, uv); #COMPUTE_CODE imageStore(color_image, uv, color); } """ - User shader code replaces
#COMPUTE_CODEat runtime. - Compiled via
rd.shader_create_from_spirv()at runtime.
OGRE's RTSS (Run Time Shader System):
- Not an uber-shader. Manages a set of opaque
SubRenderStatecomponents. - Each component implements a specific effect.
- Components are composed and code-generated at runtime.
- Avoids the "exploding
#ifdef" problem of uber-shaders.
3b. Shader Permutations vs Branching
The permutation problem (from MJP's detailed analysis):
- Each feature combination = separate compiled shader.
- Exponential growth: N binary features = 2^N permutations.
- Costs: compilation time, memory, PSO creation, binding overhead, instruction cache pressure.
- Register pressure: Uber-shaders with many features need more registers, reducing occupancy even for materials that don't use all features.
Branching rules for GPUs:
- Uniform branches (same path for all pixels in a warp): Essentially free. The driver compiles both paths and selects one.
- Divergent branches (different paths within a warp): Both paths execute serially, wasting cycles.
- Branches on uniforms/constant data: OK and performant.
- Branches based on per-pixel data: Expensive when pixels in the same warp diverge.
Best practice: Use Vulkan specialization constants for compile-time branching (uber-shader with static branching). This gives you permutation-like performance with fewer actual shader binaries. The driver can optimize away dead code paths.
3c. Compute Shaders vs Fragment Shaders
Fragment shaders are generally faster for simple image processing:
- Fragment shaders benefit from hardware texture prefetch and caching optimized for 2D spatial locality.
- For simple per-pixel operations (passthrough, basic color transforms): fragment shaders ~30% faster than compute (Leadwerks benchmarks: 770 FPS vs 600 FPS).
- For multi-pass chained operations: fragment shaders maintain advantage (670 FPS vs 180 FPS at 10 passes).
Compute shaders are better when:
- You need shared memory access within workgroups (e.g., local convolution, shared reductions).
- You need read-write access to the same texture (e.g., iterative algorithms like Jump Flood).
- You're doing operations that aren't naturally per-pixel (histogram, reduction, sorting).
- You want explicit control over workgroup dispatch.
On tile-based GPUs (mobile): Arm documentation explicitly warns: "Compute shaders can be slower and less energy-efficient than fragment shaders for simple post-processing workloads."
For compositing: Use fragment shaders for per-pixel operations (blend, color grade, transform). Use compute for multi-pass algorithms that need shared memory (blur separable passes, glare FFT, flood fill).
3d. Bindless Textures and Descriptor Arrays
The concept: Instead of binding one texture per descriptor set, bind a large array of descriptors once. Access textures by integer index in shaders.
Vulkan implementation (from VK_EXT_descriptor_indexing, core since Vulkan 1.2):
// GLSL
#extension GL_EXT_nonuniform_qualifier : enable
layout(set = 1, binding = 10) uniform sampler2D textures[];
vec4 color = texture(textures[albedo_id], uv);
Key features:
VK_DESCRIPTOR_BINDING_UPDATE_AFTER_BIND_BIT: Update descriptors after binding.VK_DESCRIPTOR_BINDING_PARTIALLY_BOUND_BIT: Not all slots need valid descriptors.NonUniformResourceIndex: For divergent indexing within a warp.
For a compositor: Bindless is extremely useful. All input textures from the graph can live in one descriptor set. Each node shader indexes into the set by texture ID. This avoids re-binding descriptor sets per node.
Trade-off: Indirect memory loads can be slower on some mobile GPUs. Desktop GPUs handle this well.
4. Slang-Specific Patterns
Overview
Slang is a Khronos-hosted, open-source shading language. HLSL-like syntax with modern features:
- Targets: SPIR-V (Vulkan), DXIL (D3D12), Metal, CUDA, WGSL, CPU.
- Hosted by Khronos with broad industry governance.
- Based on years of NVIDIA/CMU/Stanford/MIT research.
Key Features Relevant to Compositing
Modules: Slang supports module and import for separate compilation. Modules compile to a custom IR and can be linked at runtime to produce SPIR-V or DXIL. This is exactly what a node compositor needs — each node type can be a module, and compositions are linked at runtime.
Generics and Interfaces: Instead of #ifdef permutations, use generics:
interface IImageOp {
float4 evaluate(float4 input, PixelContext ctx);
}
struct BlendOp : IImageOp {
float4 evaluate(float4 input, PixelContext ctx) { ... }
}
// Generic function specialized at compile time
T evaluateGraph<T : IImageOp>(T op, float4 input) {
return op.evaluate(input, ctx);
}
Runtime code generation: Slang supports runtime compilation and linking. From the docs: "Slang modules can be independently compiled offline to a custom IR and then linked at runtime to generate code in formats such as DXIL or SPIR-V." This means you can:
- Compile each node's shader as a Slang module.
- At graph edit time, link modules together.
- Generate the final SPIR-V/DXIL for the composed graph.
Reflection API: TypeReflection, VariableReflection, getLayout() allow querying shader structure at runtime — useful for automatically creating descriptor layouts.
Automatic Differentiation: fwd_diff and bwd_diff for gradient-based operations (relevant for differentiable compositing or learned operations).
Slang vs GLSL/HLSL for Compositing
| Feature | Slang | GLSL | HLSL |
|---|---|---|---|
| Separate compilation | ✅ Modules | ❌ Single TU | ⚠️ Limited |
| Runtime linking | ✅ | ❌ | ❌ |
| Generics/interfaces | ✅ | ❌ | ⚠️ Templates (limited) |
| Cross-platform | ✅ (Vulkan/Metal/DX/CUDA) | ⚠️ (OpenGL/Vulkan) | ⚠️ (DX only) |
| Vulkan SPIR-V | ✅ First-class | ✅ via glslc | ⚠️ via dxc |
| Runtime compilation | ✅ | ❌ | ❌ |
| HLSL compatibility | ✅ Most HLSL compiles out-of-box | ❌ | ✅ |
Recommendation
Slang is the ideal choice for a Vulkan-based compositor. Its module system directly solves the "runtime shader composition" problem. Each node type = a Slang module. Graph composition = module linking. No need for runtime string-based code generation.
5. Vulkan-Specific Considerations
Multi-Pass Image Processing
Render Pass approach (traditional):
// Pass 1: Blur
VkRenderPassBeginInfo rp1 = { .renderPass = blurPass, .framebuffer = blurFBO };
vkCmdBeginRenderPass(cmd, &rp1, VK_SUBPASS_CONTENTS_INLINE);
vkCmdBindPipeline(cmd, VK_PIPELINE_BIND_POINT_GRAPHICS, blurPipeline);
vkCmdDraw(cmd, 4, 1, 0, 0); // fullscreen quad
vkCmdEndRenderPass(cmd);
// Barrier between passes
VkImageMemoryBarrier barrier = {
.srcAccessMask = VK_ACCESS_COLOR_ATTACHMENT_WRITE_BIT,
.dstAccessMask = VK_ACCESS_SHADER_READ_BIT,
.oldLayout = VK_IMAGE_LAYOUT_COLOR_ATTACHMENT_OPTIMAL,
.newLayout = VK_IMAGE_LAYOUT_SHADER_READ_ONLY_OPTIMAL,
};
vkCmdPipelineBarrier(cmd, ...);
// Pass 2: Color grade
VkRenderPassBeginInfo rp2 = { .renderPass = gradePass, .framebuffer = gradeFBO };
vkCmdBeginRenderPass(cmd, &rp2, VK_SUBPASS_CONTENTS_INLINE);
vkCmdBindDescriptorSets(cmd, ..., gradeDescriptorSet); // binds blur result as texture
vkCmdBindPipeline(cmd, VK_PIPELINE_BIND_POINT_GRAPHICS, gradePipeline);
vkCmdDraw(cmd, 4, 1, 0, 0);
vkCmdEndRenderPass(cmd);
Dynamic Rendering approach (Vulkan 1.3 / VK_KHR_dynamic_rendering):
- Skip
VkRenderPassandVkFramebufferobjects entirely. - Use
vkCmdBeginRenderingwithVkRenderingInfospecifying attachments directly. - Simpler API, fewer objects to manage.
Descriptor Management Best Practices
From ARM and NVIDIA guidelines:
- Don't allocate descriptor sets on hot paths. Pre-allocate pools.
- Use
VK_DESCRIPTOR_TYPE_UNIFORM_BUFFER_DYNAMIC/VK_DESCRIPTOR_TYPE_STORAGE_BUFFER_DYNAMICfor per-draw offsets instead of new descriptor sets. - Pack descriptor bindings as tightly as possible. No holes.
- Reuse descriptor sets — update them rather than reallocating.
- For a compositor with bindless: create ONE large descriptor set with all textures. Bind once, index by ID.
Pipeline Layout Optimization
- Keep pipeline layouts consistent across similar shaders to reduce pipeline switches.
- Use push constants for small, per-pass data (resolution, time, parameters) — cheaper than UBOs for small data.
- Pre-create pipeline cache and use
VkPipelineCacheto speed up PSO creation.
Synchronization for Multi-Pass
- Use pipeline barriers between passes that read/write the same images.
- For independent passes (operating on different textures), no barrier needed — can even record in parallel.
- Use events for fine-grained synchronization within a command buffer.
- Timeline semaphores (Vulkan 1.2+) for more flexible GPU-GPU synchronization.
Tile-Based GPU Optimization (Mobile)
- Use subpasses to keep intermediate data in tile memory (on-chip).
VK_KHR_dynamic_rendering_local_readallows subpass-like behavior with dynamic rendering.- Set
loadOp = VK_ATTACHMENT_LOAD_OP_DONT_CAREandstoreOp = VK_ATTACHMENT_STORE_OP_DONT_CAREfor transient intermediates. - Merge subpasses when they share attachments (ARM: ≤8 unique attachments).
6. Industry Best Practices
The Render Graph Pattern
Modern engines use a frame graph / render graph (DAG) for multi-pass rendering:
- Declare passes and their resource inputs/outputs.
- Analyze dependencies — build execution order automatically.
- Infer synchronization — barriers are generated from resource usage.
- Alias resources — textures with non-overlapping lifetimes can share memory.
- Cull unused passes — if an output isn't used, skip the pass.
This is the most mature pattern for managing multi-pass image processing. Referenced in:
- Vulkan Tutorial: "Engine Architecture: Rendering Pipeline"
- Cat Game's "Advanced Vulkan Rendering: Building a Modern Frame Graph"
- Frostbite's "FrameGraph" (EA/DICE)
Fusing Operations
From TFLite GPU and Blender compositor:
- Fuse element-wise operations with computationally expensive ones (activations + convolution, color transforms + blend).
- Inline parameters directly into shader code instead of passing via uniforms (bakes constants, reduces memory I/O).
- Bake uniforms into source code when they don't change per-pixel.
Texture Pool Management
For a compositor with potentially many intermediate textures:
- Pre-allocate a pool of textures at common resolutions.
- Reference-count or track lifetime of each texture.
- Reuse textures with matching format/resolution once their producer is done.
- On mobile, prefer smaller intermediate formats (RGBA16F over RGBA32F when precision allows).
Papers and References
- "Performance Implications of Node Graph Complexity in Real-Time Compositing" (IEEE, 2024) — Studies Blender EEVEE's node graph rendering performance vs. structural complexity.
- "Compute Shader in Image Processing Development" (CEUR Workshop, 2020) — Compares CPU, fragment, compute, and Vulkan fragment for image processing. Found compute shader overhead makes it slower for simple operations.
- Blender Real-time Compositor (code.blender.org, 2022) — GPU-accelerated compositor architecture with operation graph, domain system, and shader-based execution.
- "The Shader Permutation Problem" (MJP, 2021) — Comprehensive analysis of uber-shader vs. permutation trade-offs.
- "GPU Rendering Pipeline: Blend Modes, Porter-Duff Compositing" (Lucio Durán, 2025) — Browser rendering pipeline compositing patterns.
- "High-Performance Software Rasterization on GPUs" (NVIDIA Research, 2011) — Software GPU pipeline, relevant for understanding GPU architecture.
- Vulkan Samples (Khronos) — Descriptor management, subpasses, async compute, tile-based rendering best practices.
Recommended Architecture for Prism
Based on all research:
- DAG-based execution: Topological sort the node graph. Execute in dependency order.
- Separate shaders per node type: Each node type (Blend, ColorGrade, Blur, etc.) has a dedicated Slang shader module.
- Runtime composition via Slang modules: Simple chains of per-pixel operations can be fused into single compute/fragment passes by linking their Slang modules.
- Texture pool: Pre-allocated RGBA16F textures. Reference-counted. Reuse when possible.
- Ping-pong for chains: Two textures alternating for sequential per-pixel chains.
- Fragment shaders for per-pixel ops, compute shaders for operations needing shared memory (blur, convolution, reduction).
- Bindless descriptors: One large descriptor set with all input textures. Node shaders index by texture ID.
- Push constants for per-pass uniforms (resolution, parameters).
- Pipeline barriers between passes on the same texture. No barriers for independent passes.
- Render graph for automatic dependency tracking and synchronization.