From 96c89dcb86e49303e112651921590d67574784fb Mon Sep 17 00:00:00 2001 From: Abdelrahman Said Date: Tue, 14 Jul 2026 00:03:40 +0100 Subject: [PATCH] Initial research on shaders --- .../research/shader-architecture-patterns.md | 346 ++++++++++++++++++ 1 file changed, 346 insertions(+) create mode 100644 documents/research/shader-architecture-patterns.md diff --git a/documents/research/shader-architecture-patterns.md b/documents/research/shader-architecture-patterns.md new file mode 100644 index 0000000..924a861 --- /dev/null +++ b/documents/research/shader-architecture-patterns.md @@ -0,0 +1,346 @@ +# Shader Architecture Patterns for Node-Based Image Compositing + +Research conducted 2026-07-13. + +--- + +## 1. Single Shader vs Multiple Shaders + +### How Professional Compositors Handle It + +**Blender Compositor (GPU backend)** — The most relevant case study: +- Blender's GPU compositor collapses multiple connected nodes into a **"compile unit"** and generates a **single compute shader** per unit. +- The `ShaderOperation` class iterates through a `compile_unit_` (a set of nodes) and links their GLSL logic into one shader: `source/blender/compositor/intern/shader_operation.cc:122-135`. +- Simple per-pixel operations (Math, Color Mix, Invert, etc.) are fused into a single pass. Operations that can't be expressed as shaders fall back to `MultiFunctionProcedureOperation` on CPU. +- **Key insight**: Blender uses a **hybrid approach** — fuse what you can into single shaders, fall back to separate passes for complex operations (blur, glare, convolution). + +**Natron** — CPU-based compositor using OpenFX plugins: +- Each node is a separate processing unit (separate plugin call). +- Multi-threaded tile-based processing per node. +- Not GPU-accelerated; no shader fusion. + +**DaVinci Resolve / Fusion** — Proprietary: +- Uses a node graph where each node can have internal multi-pass processing. +- Fusion's "Flow Region" system groups nodes for optimization. +- Effectively separate shaders per node, with internal optimization. + +### Recommended Approach for Prism + +**Use separate shaders per node, with optional fusion of simple nodes.** Rationale: +- Nodes in a compositing graph have diverse operations (blur vs. blend vs. color grade). An uber-shader would have massive register pressure and poor occupancy. +- Simple per-pixel operations (math, color mix, gamma) can be fused into chains as an optimization. +- Complex operations (blur, convolutions, warps) need their own shader passes anyway. + +--- + +## 2. Texture Ping-Ponging + +### The Pattern + +Texture ping-ponging is the fundamental technique for chaining GPU image operations: + +1. Allocate two textures (A and B) at the target resolution. +2. Bind texture A as input, render to texture B. +3. Swap: bind texture B as input, render to texture A. +4. Repeat for as many passes as needed. + +``` +Pass 1: Read(A) → Write(B) [e.g., blur] +Pass 2: Read(B) → Write(A) [e.g., color grade] +Pass 3: Read(A) → Write(B) [e.g., blend] +Final: Display(B) +``` + +### How It Works in Practice + +**WebGL/Fragment Shader approach** (from multiple sources): +- Create Framebuffer Objects (FBOs) with texture attachments. +- Bind FBO → render fullscreen quad → output goes to texture. +- Bind different FBO or default framebuffer → read from that texture. + +**Vulkan approach**: +- Use `VkImage` objects as both sampler inputs and render targets. +- Between passes, issue a pipeline barrier (`VK_PIPELINE_STAGE_FRAGMENT_SHADER_BIT → VK_PIPELINE_STAGE_COLOR_ATTACHMENT_OUTPUT_BIT`). +- Manage image layouts: `SHADER_READ_ONLY_OPTIMAL` → `COLOR_ATTACHMENT_OPTIMAL` → `SHADER_READ_ONLY_OPTIMAL`. + +**Metal approach** (from Kosikowski's article): +- Compute shaders read from `inTexture` and write to `outTexture`. +- Swap the texture references between passes. + +### Important Considerations + +- **Image layout transitions** are critical in Vulkan. Each pass requires the texture to be in the correct layout. +- **Load/store ops**: For intermediate textures, use `VK_ATTACHMENT_LOAD_OP_DONT_CARE` and `VK_ATTACHMENT_STORE_OP_DONT_CARE` when contents aren't needed — saves bandwidth. +- **Resolution management**: Different nodes may operate at different resolutions. The compositor must manage a texture pool and handle up/downsampling. +- **On tile-based GPUs (mobile)**: Multiple passes that write/read intermediate textures to external memory is expensive. Use Vulkan subpasses or `VK_KHR_dynamic_rendering_local_read` to keep data on-chip. + +--- + +## 3. Shader Composition Strategies + +### 3a. Runtime Shader Generation + +**Blender's approach** (most relevant): +- The compositor has a `gpu_shader_compositor_code_generation.glsl` library. +- `ShaderOperation` generates GLSL code by iterating through a compile unit's nodes and concatenating their shader code contributions. +- The generated code is compiled via Blender's `GPUMaterial` system. +- Node settings are passed as UBOs; images are bound as `image2D`/`sampler2D`. + +**Godot's compositor approach**: +- Uses a **template + injection** pattern: + ``` + const template_shader = """ + #version 450 + layout(local_size_x = 8, local_size_y = 8, local_size_z = 1) in; + layout(rgba16f, set = 0, binding = 0) uniform image2D color_image; + void main() { + // ... boilerplate ... + vec4 color = imageLoad(color_image, uv); + #COMPUTE_CODE + imageStore(color_image, uv, color); + } + """ + ``` +- User shader code replaces `#COMPUTE_CODE` at runtime. +- Compiled via `rd.shader_create_from_spirv()` at runtime. + +**OGRE's RTSS (Run Time Shader System)**: +- Not an uber-shader. Manages a set of opaque `SubRenderState` components. +- Each component implements a specific effect. +- Components are composed and code-generated at runtime. +- Avoids the "exploding `#ifdef`" problem of uber-shaders. + +### 3b. Shader Permutations vs Branching + +**The permutation problem** (from MJP's detailed analysis): +- Each feature combination = separate compiled shader. +- Exponential growth: N binary features = 2^N permutations. +- Costs: compilation time, memory, PSO creation, binding overhead, instruction cache pressure. +- **Register pressure**: Uber-shaders with many features need more registers, reducing occupancy even for materials that don't use all features. + +**Branching rules for GPUs**: +- **Uniform branches** (same path for all pixels in a warp): Essentially free. The driver compiles both paths and selects one. +- **Divergent branches** (different paths within a warp): Both paths execute serially, wasting cycles. +- **Branches on uniforms/constant data**: OK and performant. +- **Branches based on per-pixel data**: Expensive when pixels in the same warp diverge. + +**Best practice**: Use **Vulkan specialization constants** for compile-time branching (uber-shader with static branching). This gives you permutation-like performance with fewer actual shader binaries. The driver can optimize away dead code paths. + +### 3c. Compute Shaders vs Fragment Shaders + +**Fragment shaders are generally faster for simple image processing:** +- Fragment shaders benefit from hardware texture prefetch and caching optimized for 2D spatial locality. +- For simple per-pixel operations (passthrough, basic color transforms): fragment shaders ~30% faster than compute (Leadwerks benchmarks: 770 FPS vs 600 FPS). +- For multi-pass chained operations: fragment shaders maintain advantage (670 FPS vs 180 FPS at 10 passes). + +**Compute shaders are better when:** +- You need **shared memory** access within workgroups (e.g., local convolution, shared reductions). +- You need **read-write access** to the same texture (e.g., iterative algorithms like Jump Flood). +- You're doing operations that aren't naturally per-pixel (histogram, reduction, sorting). +- You want explicit control over workgroup dispatch. + +**On tile-based GPUs (mobile)**: Arm documentation explicitly warns: "Compute shaders can be slower and less energy-efficient than fragment shaders for simple post-processing workloads." + +**For compositing**: Use fragment shaders for per-pixel operations (blend, color grade, transform). Use compute for multi-pass algorithms that need shared memory (blur separable passes, glare FFT, flood fill). + +### 3d. Bindless Textures and Descriptor Arrays + +**The concept**: Instead of binding one texture per descriptor set, bind a large array of descriptors once. Access textures by integer index in shaders. + +**Vulkan implementation** (from `VK_EXT_descriptor_indexing`, core since Vulkan 1.2): +```glsl +// GLSL +#extension GL_EXT_nonuniform_qualifier : enable +layout(set = 1, binding = 10) uniform sampler2D textures[]; +vec4 color = texture(textures[albedo_id], uv); +``` + +**Key features**: +- `VK_DESCRIPTOR_BINDING_UPDATE_AFTER_BIND_BIT`: Update descriptors after binding. +- `VK_DESCRIPTOR_BINDING_PARTIALLY_BOUND_BIT`: Not all slots need valid descriptors. +- `NonUniformResourceIndex`: For divergent indexing within a warp. + +**For a compositor**: Bindless is extremely useful. All input textures from the graph can live in one descriptor set. Each node shader indexes into the set by texture ID. This avoids re-binding descriptor sets per node. + +**Trade-off**: Indirect memory loads can be slower on some mobile GPUs. Desktop GPUs handle this well. + +--- + +## 4. Slang-Specific Patterns + +### Overview + +Slang is a Khronos-hosted, open-source shading language. HLSL-like syntax with modern features: +- Targets: SPIR-V (Vulkan), DXIL (D3D12), Metal, CUDA, WGSL, CPU. +- Hosted by Khronos with broad industry governance. +- Based on years of NVIDIA/CMU/Stanford/MIT research. + +### Key Features Relevant to Compositing + +**Modules**: Slang supports `module` and `import` for separate compilation. Modules compile to a custom IR and can be linked at runtime to produce SPIR-V or DXIL. This is **exactly what a node compositor needs** — each node type can be a module, and compositions are linked at runtime. + +**Generics and Interfaces**: Instead of #ifdef permutations, use generics: +```slang +interface IImageOp { + float4 evaluate(float4 input, PixelContext ctx); +} + +struct BlendOp : IImageOp { + float4 evaluate(float4 input, PixelContext ctx) { ... } +} + +// Generic function specialized at compile time +T evaluateGraph(T op, float4 input) { + return op.evaluate(input, ctx); +} +``` + +**Runtime code generation**: Slang supports **runtime compilation and linking**. From the docs: "Slang modules can be independently compiled offline to a custom IR and then linked at runtime to generate code in formats such as DXIL or SPIR-V." This means you can: +1. Compile each node's shader as a Slang module. +2. At graph edit time, link modules together. +3. Generate the final SPIR-V/DXIL for the composed graph. + +**Reflection API**: `TypeReflection`, `VariableReflection`, `getLayout()` allow querying shader structure at runtime — useful for automatically creating descriptor layouts. + +**Automatic Differentiation**: `fwd_diff` and `bwd_diff` for gradient-based operations (relevant for differentiable compositing or learned operations). + +### Slang vs GLSL/HLSL for Compositing + +| Feature | Slang | GLSL | HLSL | +|---------|-------|------|------| +| Separate compilation | ✅ Modules | ❌ Single TU | ⚠️ Limited | +| Runtime linking | ✅ | ❌ | ❌ | +| Generics/interfaces | ✅ | ❌ | ⚠️ Templates (limited) | +| Cross-platform | ✅ (Vulkan/Metal/DX/CUDA) | ⚠️ (OpenGL/Vulkan) | ⚠️ (DX only) | +| Vulkan SPIR-V | ✅ First-class | ✅ via glslc | ⚠️ via dxc | +| Runtime compilation | ✅ | ❌ | ❌ | +| HLSL compatibility | ✅ Most HLSL compiles out-of-box | ❌ | ✅ | + +### Recommendation + +**Slang is the ideal choice for a Vulkan-based compositor.** Its module system directly solves the "runtime shader composition" problem. Each node type = a Slang module. Graph composition = module linking. No need for runtime string-based code generation. + +--- + +## 5. Vulkan-Specific Considerations + +### Multi-Pass Image Processing + +**Render Pass approach** (traditional): +```c +// Pass 1: Blur +VkRenderPassBeginInfo rp1 = { .renderPass = blurPass, .framebuffer = blurFBO }; +vkCmdBeginRenderPass(cmd, &rp1, VK_SUBPASS_CONTENTS_INLINE); +vkCmdBindPipeline(cmd, VK_PIPELINE_BIND_POINT_GRAPHICS, blurPipeline); +vkCmdDraw(cmd, 4, 1, 0, 0); // fullscreen quad +vkCmdEndRenderPass(cmd); + +// Barrier between passes +VkImageMemoryBarrier barrier = { + .srcAccessMask = VK_ACCESS_COLOR_ATTACHMENT_WRITE_BIT, + .dstAccessMask = VK_ACCESS_SHADER_READ_BIT, + .oldLayout = VK_IMAGE_LAYOUT_COLOR_ATTACHMENT_OPTIMAL, + .newLayout = VK_IMAGE_LAYOUT_SHADER_READ_ONLY_OPTIMAL, +}; +vkCmdPipelineBarrier(cmd, ...); + +// Pass 2: Color grade +VkRenderPassBeginInfo rp2 = { .renderPass = gradePass, .framebuffer = gradeFBO }; +vkCmdBeginRenderPass(cmd, &rp2, VK_SUBPASS_CONTENTS_INLINE); +vkCmdBindDescriptorSets(cmd, ..., gradeDescriptorSet); // binds blur result as texture +vkCmdBindPipeline(cmd, VK_PIPELINE_BIND_POINT_GRAPHICS, gradePipeline); +vkCmdDraw(cmd, 4, 1, 0, 0); +vkCmdEndRenderPass(cmd); +``` + +**Dynamic Rendering approach** (Vulkan 1.3 / `VK_KHR_dynamic_rendering`): +- Skip `VkRenderPass` and `VkFramebuffer` objects entirely. +- Use `vkCmdBeginRendering` with `VkRenderingInfo` specifying attachments directly. +- Simpler API, fewer objects to manage. + +### Descriptor Management Best Practices + +From ARM and NVIDIA guidelines: +- **Don't allocate descriptor sets on hot paths.** Pre-allocate pools. +- **Use `VK_DESCRIPTOR_TYPE_UNIFORM_BUFFER_DYNAMIC` / `VK_DESCRIPTOR_TYPE_STORAGE_BUFFER_DYNAMIC`** for per-draw offsets instead of new descriptor sets. +- **Pack descriptor bindings** as tightly as possible. No holes. +- **Reuse descriptor sets** — update them rather than reallocating. +- For a compositor with bindless: create ONE large descriptor set with all textures. Bind once, index by ID. + +### Pipeline Layout Optimization + +- Keep pipeline layouts consistent across similar shaders to reduce pipeline switches. +- Use **push constants** for small, per-pass data (resolution, time, parameters) — cheaper than UBOs for small data. +- Pre-create pipeline cache and use `VkPipelineCache` to speed up PSO creation. + +### Synchronization for Multi-Pass + +- Use **pipeline barriers** between passes that read/write the same images. +- For independent passes (operating on different textures), no barrier needed — can even record in parallel. +- Use **events** for fine-grained synchronization within a command buffer. +- **Timeline semaphores** (Vulkan 1.2+) for more flexible GPU-GPU synchronization. + +### Tile-Based GPU Optimization (Mobile) + +- Use **subpasses** to keep intermediate data in tile memory (on-chip). +- `VK_KHR_dynamic_rendering_local_read` allows subpass-like behavior with dynamic rendering. +- Set `loadOp = VK_ATTACHMENT_LOAD_OP_DONT_CARE` and `storeOp = VK_ATTACHMENT_STORE_OP_DONT_CARE` for transient intermediates. +- Merge subpasses when they share attachments (ARM: ≤8 unique attachments). + +--- + +## 6. Industry Best Practices + +### The Render Graph Pattern + +Modern engines use a **frame graph** / **render graph** (DAG) for multi-pass rendering: +1. **Declare passes** and their resource inputs/outputs. +2. **Analyze dependencies** — build execution order automatically. +3. **Infer synchronization** — barriers are generated from resource usage. +4. **Alias resources** — textures with non-overlapping lifetimes can share memory. +5. **Cull unused passes** — if an output isn't used, skip the pass. + +This is the most mature pattern for managing multi-pass image processing. Referenced in: +- Vulkan Tutorial: "Engine Architecture: Rendering Pipeline" +- Cat Game's "Advanced Vulkan Rendering: Building a Modern Frame Graph" +- Frostbite's "FrameGraph" (EA/DICE) + +### Fusing Operations + +From TFLite GPU and Blender compositor: +- **Fuse element-wise operations** with computationally expensive ones (activations + convolution, color transforms + blend). +- **Inline parameters** directly into shader code instead of passing via uniforms (bakes constants, reduces memory I/O). +- **Bake uniforms into source code** when they don't change per-pixel. + +### Texture Pool Management + +For a compositor with potentially many intermediate textures: +- Pre-allocate a pool of textures at common resolutions. +- Reference-count or track lifetime of each texture. +- Reuse textures with matching format/resolution once their producer is done. +- On mobile, prefer smaller intermediate formats (RGBA16F over RGBA32F when precision allows). + +### Papers and References + +1. **"Performance Implications of Node Graph Complexity in Real-Time Compositing"** (IEEE, 2024) — Studies Blender EEVEE's node graph rendering performance vs. structural complexity. +2. **"Compute Shader in Image Processing Development"** (CEUR Workshop, 2020) — Compares CPU, fragment, compute, and Vulkan fragment for image processing. Found compute shader overhead makes it slower for simple operations. +3. **Blender Real-time Compositor** (code.blender.org, 2022) — GPU-accelerated compositor architecture with operation graph, domain system, and shader-based execution. +4. **"The Shader Permutation Problem"** (MJP, 2021) — Comprehensive analysis of uber-shader vs. permutation trade-offs. +5. **"GPU Rendering Pipeline: Blend Modes, Porter-Duff Compositing"** (Lucio Durán, 2025) — Browser rendering pipeline compositing patterns. +6. **"High-Performance Software Rasterization on GPUs"** (NVIDIA Research, 2011) — Software GPU pipeline, relevant for understanding GPU architecture. +7. **Vulkan Samples** (Khronos) — Descriptor management, subpasses, async compute, tile-based rendering best practices. + +### Recommended Architecture for Prism + +Based on all research: + +1. **DAG-based execution**: Topological sort the node graph. Execute in dependency order. +2. **Separate shaders per node type**: Each node type (Blend, ColorGrade, Blur, etc.) has a dedicated Slang shader module. +3. **Runtime composition via Slang modules**: Simple chains of per-pixel operations can be fused into single compute/fragment passes by linking their Slang modules. +4. **Texture pool**: Pre-allocated RGBA16F textures. Reference-counted. Reuse when possible. +5. **Ping-pong for chains**: Two textures alternating for sequential per-pixel chains. +6. **Fragment shaders for per-pixel ops**, compute shaders for operations needing shared memory (blur, convolution, reduction). +7. **Bindless descriptors**: One large descriptor set with all input textures. Node shaders index by texture ID. +8. **Push constants** for per-pass uniforms (resolution, parameters). +9. **Pipeline barriers** between passes on the same texture. No barriers for independent passes. +10. **Render graph** for automatic dependency tracking and synchronization.