diff --git a/documents/research/rendering-hardware-interface.md b/documents/research/rendering-hardware-interface.md new file mode 100644 index 0000000..698b5c1 --- /dev/null +++ b/documents/research/rendering-hardware-interface.md @@ -0,0 +1,489 @@ +# Rendering Hardware Interface (RHI) — Research + +## 1. What is an RHI + +A Render(ing) Hardware Interface (RHI) is the abstraction layer between a +renderer and the platform-specific graphics API (Vulkan, DirectX 12, Metal, +etc.). It allows the renderer to be completely API-independent while providing +a simpler, more explicit interface than the raw API. + +**Key goals:** +- API portability (write once, run on Vulkan, D3D12, Metal) +- Clean separation: renderer talks to RHI, RHI talks to driver +- Zero/low overhead over the native API +- Explicit control over GPU resources, synchronisation, and memory + +**What an RHI is NOT:** +- A high-level rendering engine or framework +- An automatic resource manager +- A scene graph or render graph + +--- + +## 2. Common Architecture Patterns (from real-world RHIs) + +### 2.1 Object-Based RHI (orhi, SnapRHI, tobyc11/RHI) + +Each GPU concept maps to an explicit object with a create/destroy lifecycle. +Objects are passed by handle/pointer to command recording functions. + +``` +RHI Instance → PhysicalDevice → Device → Queue +Device → CommandPool → CommandBuffer +Device → Buffer, Texture, Sampler +Device → ShaderModule, PipelineLayout, Pipeline +Device → DescriptorPool, DescriptorSetLayout, DescriptorSet +Device → Fence, Semaphore, SwapChain +``` + +**Examples:** +- [orhi](https://github.com/adriengivry/orhi) — C++20, Vulkan/D3D12/Metal, MIT +- [SnapRHI](https://github.com/Snapchat/SnapRHI) — C++20, Metal/Vulkan/OpenGL, Apache 2.0 +- [NVRHI](https://github.com/NVIDIAGameWorks/nvrhi) — C++14, Vulkan/D3D12, NVIDIA +- [RGL](https://github.com/RavEngine/RGL) — C++20, Vulkan/D3D12/Metal + +**Pros:** +- Familiar mapping to Vulkan/D3D12 concepts +- Easy to add new backends +- Each object owns its lifetime explicitly + +**Cons:** +- Boilerplate-heavy +- API surface grows with each backend quirk exposed + +### 2.2 Command-List-Oriented RHI (Adept Engine, Unreal Engine) + +Commands are recorded into command-list objects. The renderer records draws, +bindings, and state changes into command lists which are then submitted to the +GPU. The RHI thread translates these into API-specific calls. + +``` +Renderer → RHI Command List → RHI Thread → Backend (VkCmdBuf / ID3D12GraphicsCommandList) +``` + +**Pros:** +- Natural threading model (record in parallel, submit once) +- Easy to defer and reorder commands +- Close to D3D12/Vulkan command buffer semantics + +**Cons:** +- More indirection +- State shadowing complexity + +### 2.3 Immediate-Mode RHI (VRHI, NVRHI immediate mode) + +Functions execute synchronously. No command list abstraction — the API is +called directly. Simpler but less performant for multi-threaded recording. + +**Prism choice:** Object-based + command-list-oriented. For a compositing +application the graph evaluation can record node commands into per-frame +command buffers. + +--- + +## 3. Core API Surface (what every RHI needs) + +| Category | Objects | Notes | +|---|---|---| +| **Instance/Device** | `PrRhiInstance`, `PrRhiDevice`, `PrRhiPhysicalDevice` | Instance owns debug + layers; Device owns queues + memory | +| **Swap chain** | `PrRhiSwapChain` | Presentation surface + frame sync | +| **Resources** | `PrRhiBuffer`, `PrRhiTexture`, `PrRhiSampler` | GPU memory, sub-allocated via VMA-like pattern | +| **Shaders** | `PrRhiShader` | Slang → SPIR-V → `VkShaderModule` | +| **Pipeline** | `PrRhiPipelineLayout`, `PrRhiPipeline`, `PrRhiComputePipeline` | Compiled shader + vertex layout + state | +| **Descriptors** | `PrRhiDescriptorPool`, `PrRhiDescriptorSetLayout`, `PrRhiDescriptorSet` | Bindless or bindful | +| **Commands** | `PrRhiCommandPool`, `PrRhiCommandBuffer` | Per-frame recording | +| **Sync** | `PrRhiFence`, `PrRhiSemaphore` | CPU-GPU and GPU-GPU sync | +| **Query** | `PrRhiQueryPool` | Timestamps, occlusion | + +### Minimal surface for Prism MVP + +For a node-based compositor that renders images via shader passes, the minimum +API surface is: + +``` +Device +├── CommandPool → CommandBuffer +├── Buffer (vertex, index, uniform/staging) +├── Texture (read, write, render-target) +├── Sampler +├── Shader (from SPIR-V) +├── PipelineLayout + Pipeline (graphics) +├── DescriptorSetLayout + DescriptorSet (or push descriptors) +├── Fence +└── SwapChain (for output display) +``` + +--- + +## 4. Design Decisions for Prism + +### 4.1 C11 / C++11 with wapp allocators + +The project follows C11/C++11 dual-mode from `src/wapp/`. The RHI should: + +- Use `WpAllocator *` for all allocations (no `new`/`delete` or raw malloc) +- Expose opaque handle types (`PrRhiBuffer` as struct, not `VkBuffer`) +- Keep the backend implementation in separate `.c` files per API +- Use `wp_extern`/`wp_intern`/`wp_persist` conventions + +All existing wapp infrastructure (arena allocators, arrays, queues, string +types) should be used throughout. + +### 4.2 Vulkan-only for now, but design for multi-backend + +The AGENTS.md says "Vulkan, abstracted behind an RHI". The interface should be +designed so that a D3D12 or Metal backend could be added later without changing +the renderer. This means: + +- Backend-agnostic types in the public header (`pr_rhi.h`) +- Backend-specific implementations in `rhi/vulkan/`, `rhi/d3d12/` etc. +- A factory pattern or compile-time dispatch for backend selection +- No Vulkan types in the public RHI API + +API objects that need per-backend variance: +- **Object creation/teardown** (always differs) +- **Shader compilation** (SPIR-V is universal, but creation paths differ) +- **Pipeline state** (VkPipeline vs ID3D12PipelineState) +- **Command recording** (VkCmdBuf vs ID3D12GraphicsCommandList) +- **Memory management** (VkDeviceMemory vs ID3D12Heap) + +### 4.3 Explicit over implicit + +The RHI should not hide Vulkan's explicit nature. If the renderer needs to +manage descriptor sets, layout transitions, and fences, the RHI should expose +those operations — not paper over them with OpenGL-style "bind and forget." + +### 4.4 Memory management: use VMA + +Vulkan Memory Allocator (VMA) from AMD is the de-facto standard for Vulkan +memory management. Rather than writing our own sub-allocator, we should: + +- Use VMA for host+device memory allocation +- Wrap it behind the RHI so backends can swap it out +- Expose `PrRhiAllocation` as an opaque handle + +### 4.5 Descriptor management + +For a compositor, the number of unique descriptors per frame is bounded by the +node graph size. Two approaches: + +**A) Push descriptors** (Vulkan 1.0+, no pool needed) +- Limited to `maxPushDescriptors` (typically 32-256) +- Simple — inline with command recording +- Good for small numbers of parameters per node + +**B) Descriptor sets with per-frame pools** +- More flexible for many resources +- Requires pool management and reset +- Better for texture-heavy graphs + +**Recommendation:** Use push descriptors for uniforms, small descriptor set +pools for sampled textures (images). Start with descriptor set approach since +it scales better. + +### 4.6 Pipeline management + +Pipelines in Vulkan are expensive to create. Strategy: + +- Hash pipeline state (shaders, blend mode, depth, etc.) → cache +- Create pipelines lazily on first use +- Store in a lock-free hash table (or arena-backed sorted array for + single-threaded graph eval) +- Use pipeline libraries (`VK_EXT_graphics_pipeline_library`) for faster + creation when available + +For a compositor, the number of distinct pipeline configurations is small +(blend modes, colour-grade LUTs, blit, etc.), so a simple hash map suffices. + +--- + +## 5. Vulkan-Specific Considerations + +### 5.1 Queue selection + +| Queue type | Usage in compositor | +|---|---| +| Graphics | Main rendering (draw calls) | +| Compute | Image processing, convolution, colour-grade | +| Transfer | Image upload from disk, staging | + +The device should expose at least one graphics queue. If separate compute +queues are available, use them for async processing. Transfer queue is +desirable for texture loading without stalling the render loop. + +### 5.2 Command buffer strategy + +Two-level approach: +- **Per-frame primary command buffers**: one per swap-chain image, filled by + graph evaluation +- **One-shot secondary command buffers**: for transient operations (texture + upload, blits) using `immediate_submit` pattern + +Command pools should be per-frame to allow reset without synchronisation. + +### 5.3 Synchronisation + +- `VkSemaphore` for swap-chain acquire/present +- `VkFence` for CPU-GPU sync (frame completion, upload completion) +- Timeline semaphores (`VK_KHR_timeline_semaphore`) if compute queue is used + async + +### 5.4 Image layouts + +For a compositor where images flow through nodes: +- `VK_IMAGE_LAYOUT_SHADER_READ_ONLY_OPTIMAL` — node inputs +- `VK_IMAGE_LAYOUT_COLOR_ATTACHMENT_OPTIMAL` — node render targets +- `VK_IMAGE_LAYOUT_GENERAL` — storage images (compute nodes) +- `VK_IMAGE_LAYOUT_PRESENT_SRC_KHR` — final output + +Transitions happen via explicit barriers in the command buffer (or via +`VK_KHR_synchronization2`). The RHI should expose barrier helpers. + +### 5.5 Debug / validation layers + +- Load `VK_LAYER_KHRONOS_validation` in debug builds +- Use `VK_EXT_debug_utils` for object naming +- Enable GPU-assisted validation for shader issues +- Consider RenderDoc for frame debugging + +--- + +## 6. Slang Shader Integration + +### 6.1 Why Slang + +- HLSL/GLSL compatible syntax +- Module system for shared shading code (colour science, maths) +- Single source for multiple stages (vertex+fragment in one file) +- SPIR-V output (directly consumable by Vulkan) +- Rich reflection API (bindings, buffer layouts, entry points) +- Active development, Khronos exploratory forum + +### 6.2 Compilation pipeline + +``` +.slang file + → slangc (offline) or libslang (runtime) + → SPIR-V binary + → vkCreateShaderModule + → PrRhiShader +``` + +Options: +- **Offline**: Pre-compile `.slang` → `.spv` at build time. Simpler, no runtime + compiler dependency. Good for shipped shaders. +- **Runtime**: Use libslang to compile at app startup (or on first use). + Enables shader hot-reload during development. + +**Recommendation:** Offline for release, runtime for debug/dev. The +nvpro-samples `vk_slang_editor` demonstrates both approaches. + +### 6.3 Reflection-driven pipeline creation + +Slang's reflection API (`slang::ProgramLayout`) provides: +- Binding locations (set, binding, space) +- Buffer member offsets and sizes +- Entry point names and stage types +- Specialisation constant info + +The RHI can use this to automatically build: +- `VkDescriptorSetLayout` from declared bindings +- `VkPipelineLayout` from descriptor set layouts + push constants +- Push constant ranges from reflected constant buffers + +See the [Slang Reflection API docs](https://shader-slang.com/slang/user-guide/reflection) +and the [vk_slang_editor](https://github.com/nvpro-samples/vk_slang_editor) +source for concrete patterns. + +### 6.4 Shader organisation for a compositor + +``` +src/shaders/ +├── common/ +│ ├── math.slang — Matrix/vector utilities +│ ├── colour.slang — Colour space conversions +│ └── compositing.slang — Blend equations, alpha handling +├── blit.slang — Full-screen quad draw +├── blend.slang — Over/under/add blend modes +├── grade.slang — Colour grading (lift/gamma/gain) +├── blur.slang — Separable gaussian blur +└── read.slang — Simple texture passthrough +``` + +Each shader file contains both vertex and fragment stages: + +```slang +// blit.slang +[shader("vertex")] +void vs_main(...) { ... } + +[shader("fragment")] +void fs_main(...) { ... } +``` + +--- + +## 7. Reference Projects + +| Project | Language | APIs | Notable features | +|---|---|---|---| +| **[orhi](https://github.com/adriengivry/orhi)** | C++20 | Vulkan, D3D12, Metal (planned) | Clean object hierarchy, CMake, MIT | +| **[SnapRHI](https://github.com/Snapchat/SnapRHI)** | C++20 | Metal, Vulkan, OpenGL/ES | Compile-switchable validation (if constexpr), aggressive pooling | +| **[tobyc11/RHI](https://github.com/tobyc11/RHI)** | C++ | Vulkan, D3D11 | SPIR-V as common shader format, SPIRV-Cross for translation | +| **[NVRHI](https://github.com/NVIDIAGameWorks/nvrhi)** | C++14 | Vulkan, D3D12 | Production-grade, NVIDIA maintained, header-only-ish API | +| **[RGL](https://github.com/RavEngine/RGL)** | C++20 | Vulkan, D3D12, Metal | Thin wrapper, focuses on simplicity | +| **[The Forge](https://github.com/ConfettiFX/The-Forge)** | C99/C++11 | All major APIs | Cross-platform, used in shipping games, FS | +| **[O3DE Atom RHI](https://docs.o3de.org/docs/atom-guide/dev-guide/rhi/rhi/)** | C++17 | Vulkan, D3D12, Metal | Full-featured engine RHI, frame scheduler, multi-threaded | +| **[Magma](https://github.com/vcoda/magma)** | C++17 | Vulkan | C++ abstraction, uses VMA, SPIR-V reflection | +| **[rafx](https://github.com/zeozeozeo/rafx)** | C/C++ | Vulkan, D3D12 | C API (good FFI), explicit design | + +### What to borrow from each + +| Project | Lesson | +|---|---| +| **orhi** | Object hierarchy + backend-agnostic headers pattern | +| **SnapRHI** | Compile-switchable validation; per-frame resource pooling | +| **NVRHI** | Header-only-ish API with implementation in .cpp | +| **The Forge** | C99-friendly, explicit API with minimal hidden state | +| **O3DE Atom** | Frame scheduler concept (render passes as graph nodes) | +| **RGL** | Simplicity — don't over-abstract | +| **Magma** | VMA integration pattern + SPIR-V reflection | +| **rafx** | C API design (relevant since Prism is C11) | + +--- + +## 8. Proposed Architecture for Prism + +### 8.1 Directory layout + +``` +src/prism/ +├── rhi/ +│ ├── pr_rhi.h ← Public API (backend-agnostic) +│ ├── pr_rhi_types.h ← Shared types (PrRhiBufferDesc, etc.) +│ ├── vulkan/ +│ │ ├── pr_rhi_vk.h ← Vulkan backend internal header +│ │ ├── pr_rhi_vk_device.c +│ │ ├── pr_rhi_vk_buffer.c +│ │ ├── pr_rhi_vk_texture.c +│ │ ├── pr_rhi_vk_shader.c +│ │ ├── pr_rhi_vk_pipeline.c +│ │ ├── pr_rhi_vk_descriptor.c +│ │ ├── pr_rhi_vk_command.c +│ │ └── pr_rhi_vk_swapchain.c +│ └── d3d12/ ← (future) +└── ... +``` + +### 8.2 Object lifecycle pattern + +```c +// Creation: takes an allocator + device + desc, returns handle +PrRhiBuffer *prRhiCreateBuffer(PrRhiDevice *device, const PrRhiBufferDesc *desc, + WpAllocator *alloc); + +// Destruction: frees all GPU resources + backing memory +void prRhiDestroyBuffer(PrRhiBuffer *buffer, WpAllocator *alloc); + +// Usage: command buffer records operations on handles +void prRhiCmdCopyBuffer(PrRhiCommandBuffer *cb, + PrRhiBuffer *src, PrRhiBuffer *dst); +``` + +### 8.3 Backend dispatch (compile-time) + +```c +// pr_rhi.h +typedef struct PrRhiDevice PrRhiDevice; +struct PrRhiDevice { + PrRhiDeviceVtbl *vtbl; // function pointer table + void *backend; // VkDevice or ID3D12Device +}; + +// Each backend fills the vtbl +typedef struct PrRhiDeviceVtbl { + PrRhiBuffer* (*createBuffer)(PrRhiDevice*, const PrRhiBufferDesc*, WpAllocator*); + void (*destroyBuffer)(PrRhiBuffer*, WpAllocator*); + // ... etc +} PrRhiDeviceVtbl; +``` + +This gives zero-cost abstraction (pointer indirection on calls) while keeping +the API clean. The renderer calls `device->vtbl->createBuffer(...)` and the +backend resolves the call. + +### 8.4 Frame lifecycle + +``` +Loop: + 1. prRhiAcquireNextImage(swapchain) → image index, semaphore + 2. prRhiResetCommandPool(pool, frame_idx) → recycles command buffers + 3. For each node in topo-sorted graph: + a. prRhiCmdBindPipeline(cb, pipeline) + b. prRhiCmdBindDescriptorSets(cb, ...) + c. prRhiCmdPushConstants(cb, ...) + d. prRhiCmdDraw(cb, ...) + 4. prRhiQueueSubmit(queue, cb, wait_sem, signal_sem, fence) + 5. prRhiPresent(swapchain, signal_sem) + 6. prRhiWaitForFence(fence) → CPU-GPU sync +``` + +### 8.5 Static allocation strategy + +Following wapp conventions and the project's data-oriented design principles: + +- **Command pools**: one per swap-chain image (2-3), allocated once +- **Descriptor pools**: one per frame, reset each frame +- **Upload buffers**: ring buffer for staging data, bumped each frame +- **Pipeline cache**: arena-backed hash table, populated lazily +- **Scratch buffers**: arena-allocated in the per-frame scratch space + +No dynamic allocation on the hot path — all per-frame memory comes from +frame-local arena allocators that are reset at the start of each frame. + +--- + +## 9. Open Questions + +1. **Multi-queue**: Should the RHI expose separate compute/transfer queues, or + keep everything on a single graphics queue and serialise? For an MVP, single + queue is simpler and likely sufficient. + +2. **Bindless vs bindful**: Bindless descriptors (VK_EXT_descriptor_indexing) + simplify shader resource access but require higher Vulkan version. For + maximum compatibility, start with bindful descriptor sets. + +3. **Shader compilation**: Use `slangc` at build time and ship SPIR-V, or link + libslang for runtime compilation + reflection? Runtime enables hot-reload + but adds ~15MB to binary size. Recommendation: both — offline for release, + runtime for debug. + +4. **Swap chain**: Headless mode (no window) for batch/compute-only operation? + Useful for a compositor that renders to a file. The RHI should support + both windowed and headless modes. + +5. **Vulkan version**: Target Vulkan 1.3 (widely available on desktop, adds + timeline semaphores, dynamic rendering, and sync2) with fallback to 1.2. + +--- + +## References + +- [O3DE Atom RHI Overview](https://docs.o3de.org/docs/atom-guide/dev-guide/rhi/rhi/) +- [Adept Engine RHI Design](https://andrewcjp.wordpress.com/2019/11/09/designing-a-render-hardware-interface-for-explicit-multi-gpu-programming/) +- [Unreal Engine RHI Architecture](https://dev.epicgames.com/documentation/unreal-engine/parallel-rendering-overview-for-unreal-engine) +- [orhi — OpenRHI](https://github.com/adriengivry/orhi) +- [SnapRHI](https://github.com/Snapchat/SnapRHI) +- [NVRHI](https://github.com/NVIDIAGameWorks/nvrhi) +- [The Forge](https://github.com/ConfettiFX/The-Forge) +- [RGL](https://github.com/RavEngine/RGL) +- [tobyc11/RHI](https://github.com/tobyc11/RHI) +- [rafx](https://github.com/zeozeozeo/rafx) +- [Magma](https://github.com/vcoda/magma) +- [Vulkan Memory Allocator](https://github.com/GPUOpen-LibrariesAndSDKs/VulkanMemoryAllocator) +- [Slang Shading Language](https://github.com/shader-slang/slang) +- [Slang Reflection API](https://shader-slang.com/slang/user-guide/reflection) +- [vk_slang_editor](https://github.com/nvpro-samples/vk_slang_editor) +- [Vulkan in 30 minutes](https://renderdoc.org/vulkan-in-30-minutes.html) +- [Vulkan Memory Management Guide](https://docs.vulkan.org/guide/latest/memory_allocation.html) +- [Khronos Vulkan Spec — Command Buffers](https://docs.vulkan.org/spec/latest/chapters/cmdbuffers.html) diff --git a/documents/resources.md b/documents/resources.md index d0c39b7..02e3c35 100644 --- a/documents/resources.md +++ b/documents/resources.md @@ -1,8 +1,17 @@ # Useful resources +## Graphs + - [The Algorithm Design Manual](https://sureshcseit.wordpress.com/wp-content/uploads/2021/04/skienathealgorithmdesignmanual.pdf) - [Topological Sorting using BFS (Kahn's Algorithm)](https://www.geeksforgeeks.org/topological-sorting-indegree-based-solution/) - [Kahn's Algorithm Explained with Code & Examples](https://dev.to/rui_jiang/kahns-algorithm-for-topological-sorting-explained-with-code-examples-2if5) - [Understanding Kahn's Algorithm for Topological Sorting](https://blog.devgenius.io/dsa-kahns-algorithm-for-topological-sorting-33c8587985a1) - [Detect a Cycle in Directed Graph](https://takeuforward.org/data-structure/detect-a-cycle-in-directed-graph-topological-sort-kahns-algorithm-g-23) - [Kahn's Algorithm](https://leetcodethehardway.com/tutorials/graph-theory/kahns-algorithm) + +## Rendering Hardware Interface (RHI) + +- [O3DE Atom RHI Overview](https://docs.o3de.org/docs/atom-guide/dev-guide/rhi/rhi/) +- [Adept Engine RHI Design](https://andrewcjp.wordpress.com/2019/11/09/designing-a-render-hardware-interface-for-explicit-multi-gpu-programming/) +- [Unreal Engine RHI Architecture](https://dev.epicgames.com/documentation/unreal-engine/parallel-rendering-overview-for-unreal-engine) +- [NVRHI](https://github.com/NVIDIAGameWorks/nvrhi) — NVIDIA's production RHI (Vulkan/D3D12), reference API design