Initial research on shaders
This commit is contained in:
@@ -0,0 +1,346 @@
|
||||
# Shader Architecture Patterns for Node-Based Image Compositing
|
||||
|
||||
Research conducted 2026-07-13.
|
||||
|
||||
---
|
||||
|
||||
## 1. Single Shader vs Multiple Shaders
|
||||
|
||||
### How Professional Compositors Handle It
|
||||
|
||||
**Blender Compositor (GPU backend)** — The most relevant case study:
|
||||
- Blender's GPU compositor collapses multiple connected nodes into a **"compile unit"** and generates a **single compute shader** per unit.
|
||||
- The `ShaderOperation` class iterates through a `compile_unit_` (a set of nodes) and links their GLSL logic into one shader: `source/blender/compositor/intern/shader_operation.cc:122-135`.
|
||||
- Simple per-pixel operations (Math, Color Mix, Invert, etc.) are fused into a single pass. Operations that can't be expressed as shaders fall back to `MultiFunctionProcedureOperation` on CPU.
|
||||
- **Key insight**: Blender uses a **hybrid approach** — fuse what you can into single shaders, fall back to separate passes for complex operations (blur, glare, convolution).
|
||||
|
||||
**Natron** — CPU-based compositor using OpenFX plugins:
|
||||
- Each node is a separate processing unit (separate plugin call).
|
||||
- Multi-threaded tile-based processing per node.
|
||||
- Not GPU-accelerated; no shader fusion.
|
||||
|
||||
**DaVinci Resolve / Fusion** — Proprietary:
|
||||
- Uses a node graph where each node can have internal multi-pass processing.
|
||||
- Fusion's "Flow Region" system groups nodes for optimization.
|
||||
- Effectively separate shaders per node, with internal optimization.
|
||||
|
||||
### Recommended Approach for Prism
|
||||
|
||||
**Use separate shaders per node, with optional fusion of simple nodes.** Rationale:
|
||||
- Nodes in a compositing graph have diverse operations (blur vs. blend vs. color grade). An uber-shader would have massive register pressure and poor occupancy.
|
||||
- Simple per-pixel operations (math, color mix, gamma) can be fused into chains as an optimization.
|
||||
- Complex operations (blur, convolutions, warps) need their own shader passes anyway.
|
||||
|
||||
---
|
||||
|
||||
## 2. Texture Ping-Ponging
|
||||
|
||||
### The Pattern
|
||||
|
||||
Texture ping-ponging is the fundamental technique for chaining GPU image operations:
|
||||
|
||||
1. Allocate two textures (A and B) at the target resolution.
|
||||
2. Bind texture A as input, render to texture B.
|
||||
3. Swap: bind texture B as input, render to texture A.
|
||||
4. Repeat for as many passes as needed.
|
||||
|
||||
```
|
||||
Pass 1: Read(A) → Write(B) [e.g., blur]
|
||||
Pass 2: Read(B) → Write(A) [e.g., color grade]
|
||||
Pass 3: Read(A) → Write(B) [e.g., blend]
|
||||
Final: Display(B)
|
||||
```
|
||||
|
||||
### How It Works in Practice
|
||||
|
||||
**WebGL/Fragment Shader approach** (from multiple sources):
|
||||
- Create Framebuffer Objects (FBOs) with texture attachments.
|
||||
- Bind FBO → render fullscreen quad → output goes to texture.
|
||||
- Bind different FBO or default framebuffer → read from that texture.
|
||||
|
||||
**Vulkan approach**:
|
||||
- Use `VkImage` objects as both sampler inputs and render targets.
|
||||
- Between passes, issue a pipeline barrier (`VK_PIPELINE_STAGE_FRAGMENT_SHADER_BIT → VK_PIPELINE_STAGE_COLOR_ATTACHMENT_OUTPUT_BIT`).
|
||||
- Manage image layouts: `SHADER_READ_ONLY_OPTIMAL` → `COLOR_ATTACHMENT_OPTIMAL` → `SHADER_READ_ONLY_OPTIMAL`.
|
||||
|
||||
**Metal approach** (from Kosikowski's article):
|
||||
- Compute shaders read from `inTexture` and write to `outTexture`.
|
||||
- Swap the texture references between passes.
|
||||
|
||||
### Important Considerations
|
||||
|
||||
- **Image layout transitions** are critical in Vulkan. Each pass requires the texture to be in the correct layout.
|
||||
- **Load/store ops**: For intermediate textures, use `VK_ATTACHMENT_LOAD_OP_DONT_CARE` and `VK_ATTACHMENT_STORE_OP_DONT_CARE` when contents aren't needed — saves bandwidth.
|
||||
- **Resolution management**: Different nodes may operate at different resolutions. The compositor must manage a texture pool and handle up/downsampling.
|
||||
- **On tile-based GPUs (mobile)**: Multiple passes that write/read intermediate textures to external memory is expensive. Use Vulkan subpasses or `VK_KHR_dynamic_rendering_local_read` to keep data on-chip.
|
||||
|
||||
---
|
||||
|
||||
## 3. Shader Composition Strategies
|
||||
|
||||
### 3a. Runtime Shader Generation
|
||||
|
||||
**Blender's approach** (most relevant):
|
||||
- The compositor has a `gpu_shader_compositor_code_generation.glsl` library.
|
||||
- `ShaderOperation` generates GLSL code by iterating through a compile unit's nodes and concatenating their shader code contributions.
|
||||
- The generated code is compiled via Blender's `GPUMaterial` system.
|
||||
- Node settings are passed as UBOs; images are bound as `image2D`/`sampler2D`.
|
||||
|
||||
**Godot's compositor approach**:
|
||||
- Uses a **template + injection** pattern:
|
||||
```
|
||||
const template_shader = """
|
||||
#version 450
|
||||
layout(local_size_x = 8, local_size_y = 8, local_size_z = 1) in;
|
||||
layout(rgba16f, set = 0, binding = 0) uniform image2D color_image;
|
||||
void main() {
|
||||
// ... boilerplate ...
|
||||
vec4 color = imageLoad(color_image, uv);
|
||||
#COMPUTE_CODE
|
||||
imageStore(color_image, uv, color);
|
||||
}
|
||||
"""
|
||||
```
|
||||
- User shader code replaces `#COMPUTE_CODE` at runtime.
|
||||
- Compiled via `rd.shader_create_from_spirv()` at runtime.
|
||||
|
||||
**OGRE's RTSS (Run Time Shader System)**:
|
||||
- Not an uber-shader. Manages a set of opaque `SubRenderState` components.
|
||||
- Each component implements a specific effect.
|
||||
- Components are composed and code-generated at runtime.
|
||||
- Avoids the "exploding `#ifdef`" problem of uber-shaders.
|
||||
|
||||
### 3b. Shader Permutations vs Branching
|
||||
|
||||
**The permutation problem** (from MJP's detailed analysis):
|
||||
- Each feature combination = separate compiled shader.
|
||||
- Exponential growth: N binary features = 2^N permutations.
|
||||
- Costs: compilation time, memory, PSO creation, binding overhead, instruction cache pressure.
|
||||
- **Register pressure**: Uber-shaders with many features need more registers, reducing occupancy even for materials that don't use all features.
|
||||
|
||||
**Branching rules for GPUs**:
|
||||
- **Uniform branches** (same path for all pixels in a warp): Essentially free. The driver compiles both paths and selects one.
|
||||
- **Divergent branches** (different paths within a warp): Both paths execute serially, wasting cycles.
|
||||
- **Branches on uniforms/constant data**: OK and performant.
|
||||
- **Branches based on per-pixel data**: Expensive when pixels in the same warp diverge.
|
||||
|
||||
**Best practice**: Use **Vulkan specialization constants** for compile-time branching (uber-shader with static branching). This gives you permutation-like performance with fewer actual shader binaries. The driver can optimize away dead code paths.
|
||||
|
||||
### 3c. Compute Shaders vs Fragment Shaders
|
||||
|
||||
**Fragment shaders are generally faster for simple image processing:**
|
||||
- Fragment shaders benefit from hardware texture prefetch and caching optimized for 2D spatial locality.
|
||||
- For simple per-pixel operations (passthrough, basic color transforms): fragment shaders ~30% faster than compute (Leadwerks benchmarks: 770 FPS vs 600 FPS).
|
||||
- For multi-pass chained operations: fragment shaders maintain advantage (670 FPS vs 180 FPS at 10 passes).
|
||||
|
||||
**Compute shaders are better when:**
|
||||
- You need **shared memory** access within workgroups (e.g., local convolution, shared reductions).
|
||||
- You need **read-write access** to the same texture (e.g., iterative algorithms like Jump Flood).
|
||||
- You're doing operations that aren't naturally per-pixel (histogram, reduction, sorting).
|
||||
- You want explicit control over workgroup dispatch.
|
||||
|
||||
**On tile-based GPUs (mobile)**: Arm documentation explicitly warns: "Compute shaders can be slower and less energy-efficient than fragment shaders for simple post-processing workloads."
|
||||
|
||||
**For compositing**: Use fragment shaders for per-pixel operations (blend, color grade, transform). Use compute for multi-pass algorithms that need shared memory (blur separable passes, glare FFT, flood fill).
|
||||
|
||||
### 3d. Bindless Textures and Descriptor Arrays
|
||||
|
||||
**The concept**: Instead of binding one texture per descriptor set, bind a large array of descriptors once. Access textures by integer index in shaders.
|
||||
|
||||
**Vulkan implementation** (from `VK_EXT_descriptor_indexing`, core since Vulkan 1.2):
|
||||
```glsl
|
||||
// GLSL
|
||||
#extension GL_EXT_nonuniform_qualifier : enable
|
||||
layout(set = 1, binding = 10) uniform sampler2D textures[];
|
||||
vec4 color = texture(textures[albedo_id], uv);
|
||||
```
|
||||
|
||||
**Key features**:
|
||||
- `VK_DESCRIPTOR_BINDING_UPDATE_AFTER_BIND_BIT`: Update descriptors after binding.
|
||||
- `VK_DESCRIPTOR_BINDING_PARTIALLY_BOUND_BIT`: Not all slots need valid descriptors.
|
||||
- `NonUniformResourceIndex`: For divergent indexing within a warp.
|
||||
|
||||
**For a compositor**: Bindless is extremely useful. All input textures from the graph can live in one descriptor set. Each node shader indexes into the set by texture ID. This avoids re-binding descriptor sets per node.
|
||||
|
||||
**Trade-off**: Indirect memory loads can be slower on some mobile GPUs. Desktop GPUs handle this well.
|
||||
|
||||
---
|
||||
|
||||
## 4. Slang-Specific Patterns
|
||||
|
||||
### Overview
|
||||
|
||||
Slang is a Khronos-hosted, open-source shading language. HLSL-like syntax with modern features:
|
||||
- Targets: SPIR-V (Vulkan), DXIL (D3D12), Metal, CUDA, WGSL, CPU.
|
||||
- Hosted by Khronos with broad industry governance.
|
||||
- Based on years of NVIDIA/CMU/Stanford/MIT research.
|
||||
|
||||
### Key Features Relevant to Compositing
|
||||
|
||||
**Modules**: Slang supports `module` and `import` for separate compilation. Modules compile to a custom IR and can be linked at runtime to produce SPIR-V or DXIL. This is **exactly what a node compositor needs** — each node type can be a module, and compositions are linked at runtime.
|
||||
|
||||
**Generics and Interfaces**: Instead of #ifdef permutations, use generics:
|
||||
```slang
|
||||
interface IImageOp {
|
||||
float4 evaluate(float4 input, PixelContext ctx);
|
||||
}
|
||||
|
||||
struct BlendOp : IImageOp {
|
||||
float4 evaluate(float4 input, PixelContext ctx) { ... }
|
||||
}
|
||||
|
||||
// Generic function specialized at compile time
|
||||
T evaluateGraph<T : IImageOp>(T op, float4 input) {
|
||||
return op.evaluate(input, ctx);
|
||||
}
|
||||
```
|
||||
|
||||
**Runtime code generation**: Slang supports **runtime compilation and linking**. From the docs: "Slang modules can be independently compiled offline to a custom IR and then linked at runtime to generate code in formats such as DXIL or SPIR-V." This means you can:
|
||||
1. Compile each node's shader as a Slang module.
|
||||
2. At graph edit time, link modules together.
|
||||
3. Generate the final SPIR-V/DXIL for the composed graph.
|
||||
|
||||
**Reflection API**: `TypeReflection`, `VariableReflection`, `getLayout()` allow querying shader structure at runtime — useful for automatically creating descriptor layouts.
|
||||
|
||||
**Automatic Differentiation**: `fwd_diff` and `bwd_diff` for gradient-based operations (relevant for differentiable compositing or learned operations).
|
||||
|
||||
### Slang vs GLSL/HLSL for Compositing
|
||||
|
||||
| Feature | Slang | GLSL | HLSL |
|
||||
|---------|-------|------|------|
|
||||
| Separate compilation | ✅ Modules | ❌ Single TU | ⚠️ Limited |
|
||||
| Runtime linking | ✅ | ❌ | ❌ |
|
||||
| Generics/interfaces | ✅ | ❌ | ⚠️ Templates (limited) |
|
||||
| Cross-platform | ✅ (Vulkan/Metal/DX/CUDA) | ⚠️ (OpenGL/Vulkan) | ⚠️ (DX only) |
|
||||
| Vulkan SPIR-V | ✅ First-class | ✅ via glslc | ⚠️ via dxc |
|
||||
| Runtime compilation | ✅ | ❌ | ❌ |
|
||||
| HLSL compatibility | ✅ Most HLSL compiles out-of-box | ❌ | ✅ |
|
||||
|
||||
### Recommendation
|
||||
|
||||
**Slang is the ideal choice for a Vulkan-based compositor.** Its module system directly solves the "runtime shader composition" problem. Each node type = a Slang module. Graph composition = module linking. No need for runtime string-based code generation.
|
||||
|
||||
---
|
||||
|
||||
## 5. Vulkan-Specific Considerations
|
||||
|
||||
### Multi-Pass Image Processing
|
||||
|
||||
**Render Pass approach** (traditional):
|
||||
```c
|
||||
// Pass 1: Blur
|
||||
VkRenderPassBeginInfo rp1 = { .renderPass = blurPass, .framebuffer = blurFBO };
|
||||
vkCmdBeginRenderPass(cmd, &rp1, VK_SUBPASS_CONTENTS_INLINE);
|
||||
vkCmdBindPipeline(cmd, VK_PIPELINE_BIND_POINT_GRAPHICS, blurPipeline);
|
||||
vkCmdDraw(cmd, 4, 1, 0, 0); // fullscreen quad
|
||||
vkCmdEndRenderPass(cmd);
|
||||
|
||||
// Barrier between passes
|
||||
VkImageMemoryBarrier barrier = {
|
||||
.srcAccessMask = VK_ACCESS_COLOR_ATTACHMENT_WRITE_BIT,
|
||||
.dstAccessMask = VK_ACCESS_SHADER_READ_BIT,
|
||||
.oldLayout = VK_IMAGE_LAYOUT_COLOR_ATTACHMENT_OPTIMAL,
|
||||
.newLayout = VK_IMAGE_LAYOUT_SHADER_READ_ONLY_OPTIMAL,
|
||||
};
|
||||
vkCmdPipelineBarrier(cmd, ...);
|
||||
|
||||
// Pass 2: Color grade
|
||||
VkRenderPassBeginInfo rp2 = { .renderPass = gradePass, .framebuffer = gradeFBO };
|
||||
vkCmdBeginRenderPass(cmd, &rp2, VK_SUBPASS_CONTENTS_INLINE);
|
||||
vkCmdBindDescriptorSets(cmd, ..., gradeDescriptorSet); // binds blur result as texture
|
||||
vkCmdBindPipeline(cmd, VK_PIPELINE_BIND_POINT_GRAPHICS, gradePipeline);
|
||||
vkCmdDraw(cmd, 4, 1, 0, 0);
|
||||
vkCmdEndRenderPass(cmd);
|
||||
```
|
||||
|
||||
**Dynamic Rendering approach** (Vulkan 1.3 / `VK_KHR_dynamic_rendering`):
|
||||
- Skip `VkRenderPass` and `VkFramebuffer` objects entirely.
|
||||
- Use `vkCmdBeginRendering` with `VkRenderingInfo` specifying attachments directly.
|
||||
- Simpler API, fewer objects to manage.
|
||||
|
||||
### Descriptor Management Best Practices
|
||||
|
||||
From ARM and NVIDIA guidelines:
|
||||
- **Don't allocate descriptor sets on hot paths.** Pre-allocate pools.
|
||||
- **Use `VK_DESCRIPTOR_TYPE_UNIFORM_BUFFER_DYNAMIC` / `VK_DESCRIPTOR_TYPE_STORAGE_BUFFER_DYNAMIC`** for per-draw offsets instead of new descriptor sets.
|
||||
- **Pack descriptor bindings** as tightly as possible. No holes.
|
||||
- **Reuse descriptor sets** — update them rather than reallocating.
|
||||
- For a compositor with bindless: create ONE large descriptor set with all textures. Bind once, index by ID.
|
||||
|
||||
### Pipeline Layout Optimization
|
||||
|
||||
- Keep pipeline layouts consistent across similar shaders to reduce pipeline switches.
|
||||
- Use **push constants** for small, per-pass data (resolution, time, parameters) — cheaper than UBOs for small data.
|
||||
- Pre-create pipeline cache and use `VkPipelineCache` to speed up PSO creation.
|
||||
|
||||
### Synchronization for Multi-Pass
|
||||
|
||||
- Use **pipeline barriers** between passes that read/write the same images.
|
||||
- For independent passes (operating on different textures), no barrier needed — can even record in parallel.
|
||||
- Use **events** for fine-grained synchronization within a command buffer.
|
||||
- **Timeline semaphores** (Vulkan 1.2+) for more flexible GPU-GPU synchronization.
|
||||
|
||||
### Tile-Based GPU Optimization (Mobile)
|
||||
|
||||
- Use **subpasses** to keep intermediate data in tile memory (on-chip).
|
||||
- `VK_KHR_dynamic_rendering_local_read` allows subpass-like behavior with dynamic rendering.
|
||||
- Set `loadOp = VK_ATTACHMENT_LOAD_OP_DONT_CARE` and `storeOp = VK_ATTACHMENT_STORE_OP_DONT_CARE` for transient intermediates.
|
||||
- Merge subpasses when they share attachments (ARM: ≤8 unique attachments).
|
||||
|
||||
---
|
||||
|
||||
## 6. Industry Best Practices
|
||||
|
||||
### The Render Graph Pattern
|
||||
|
||||
Modern engines use a **frame graph** / **render graph** (DAG) for multi-pass rendering:
|
||||
1. **Declare passes** and their resource inputs/outputs.
|
||||
2. **Analyze dependencies** — build execution order automatically.
|
||||
3. **Infer synchronization** — barriers are generated from resource usage.
|
||||
4. **Alias resources** — textures with non-overlapping lifetimes can share memory.
|
||||
5. **Cull unused passes** — if an output isn't used, skip the pass.
|
||||
|
||||
This is the most mature pattern for managing multi-pass image processing. Referenced in:
|
||||
- Vulkan Tutorial: "Engine Architecture: Rendering Pipeline"
|
||||
- Cat Game's "Advanced Vulkan Rendering: Building a Modern Frame Graph"
|
||||
- Frostbite's "FrameGraph" (EA/DICE)
|
||||
|
||||
### Fusing Operations
|
||||
|
||||
From TFLite GPU and Blender compositor:
|
||||
- **Fuse element-wise operations** with computationally expensive ones (activations + convolution, color transforms + blend).
|
||||
- **Inline parameters** directly into shader code instead of passing via uniforms (bakes constants, reduces memory I/O).
|
||||
- **Bake uniforms into source code** when they don't change per-pixel.
|
||||
|
||||
### Texture Pool Management
|
||||
|
||||
For a compositor with potentially many intermediate textures:
|
||||
- Pre-allocate a pool of textures at common resolutions.
|
||||
- Reference-count or track lifetime of each texture.
|
||||
- Reuse textures with matching format/resolution once their producer is done.
|
||||
- On mobile, prefer smaller intermediate formats (RGBA16F over RGBA32F when precision allows).
|
||||
|
||||
### Papers and References
|
||||
|
||||
1. **"Performance Implications of Node Graph Complexity in Real-Time Compositing"** (IEEE, 2024) — Studies Blender EEVEE's node graph rendering performance vs. structural complexity.
|
||||
2. **"Compute Shader in Image Processing Development"** (CEUR Workshop, 2020) — Compares CPU, fragment, compute, and Vulkan fragment for image processing. Found compute shader overhead makes it slower for simple operations.
|
||||
3. **Blender Real-time Compositor** (code.blender.org, 2022) — GPU-accelerated compositor architecture with operation graph, domain system, and shader-based execution.
|
||||
4. **"The Shader Permutation Problem"** (MJP, 2021) — Comprehensive analysis of uber-shader vs. permutation trade-offs.
|
||||
5. **"GPU Rendering Pipeline: Blend Modes, Porter-Duff Compositing"** (Lucio Durán, 2025) — Browser rendering pipeline compositing patterns.
|
||||
6. **"High-Performance Software Rasterization on GPUs"** (NVIDIA Research, 2011) — Software GPU pipeline, relevant for understanding GPU architecture.
|
||||
7. **Vulkan Samples** (Khronos) — Descriptor management, subpasses, async compute, tile-based rendering best practices.
|
||||
|
||||
### Recommended Architecture for Prism
|
||||
|
||||
Based on all research:
|
||||
|
||||
1. **DAG-based execution**: Topological sort the node graph. Execute in dependency order.
|
||||
2. **Separate shaders per node type**: Each node type (Blend, ColorGrade, Blur, etc.) has a dedicated Slang shader module.
|
||||
3. **Runtime composition via Slang modules**: Simple chains of per-pixel operations can be fused into single compute/fragment passes by linking their Slang modules.
|
||||
4. **Texture pool**: Pre-allocated RGBA16F textures. Reference-counted. Reuse when possible.
|
||||
5. **Ping-pong for chains**: Two textures alternating for sequential per-pixel chains.
|
||||
6. **Fragment shaders for per-pixel ops**, compute shaders for operations needing shared memory (blur, convolution, reduction).
|
||||
7. **Bindless descriptors**: One large descriptor set with all input textures. Node shaders index by texture ID.
|
||||
8. **Push constants** for per-pass uniforms (resolution, parameters).
|
||||
9. **Pipeline barriers** between passes on the same texture. No barriers for independent passes.
|
||||
10. **Render graph** for automatic dependency tracking and synchronization.
|
||||
Reference in New Issue
Block a user