Start researching RHIs
This commit is contained in:
@@ -0,0 +1,489 @@
|
|||||||
|
# Rendering Hardware Interface (RHI) — Research
|
||||||
|
|
||||||
|
## 1. What is an RHI
|
||||||
|
|
||||||
|
A Render(ing) Hardware Interface (RHI) is the abstraction layer between a
|
||||||
|
renderer and the platform-specific graphics API (Vulkan, DirectX 12, Metal,
|
||||||
|
etc.). It allows the renderer to be completely API-independent while providing
|
||||||
|
a simpler, more explicit interface than the raw API.
|
||||||
|
|
||||||
|
**Key goals:**
|
||||||
|
- API portability (write once, run on Vulkan, D3D12, Metal)
|
||||||
|
- Clean separation: renderer talks to RHI, RHI talks to driver
|
||||||
|
- Zero/low overhead over the native API
|
||||||
|
- Explicit control over GPU resources, synchronisation, and memory
|
||||||
|
|
||||||
|
**What an RHI is NOT:**
|
||||||
|
- A high-level rendering engine or framework
|
||||||
|
- An automatic resource manager
|
||||||
|
- A scene graph or render graph
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Common Architecture Patterns (from real-world RHIs)
|
||||||
|
|
||||||
|
### 2.1 Object-Based RHI (orhi, SnapRHI, tobyc11/RHI)
|
||||||
|
|
||||||
|
Each GPU concept maps to an explicit object with a create/destroy lifecycle.
|
||||||
|
Objects are passed by handle/pointer to command recording functions.
|
||||||
|
|
||||||
|
```
|
||||||
|
RHI Instance → PhysicalDevice → Device → Queue
|
||||||
|
Device → CommandPool → CommandBuffer
|
||||||
|
Device → Buffer, Texture, Sampler
|
||||||
|
Device → ShaderModule, PipelineLayout, Pipeline
|
||||||
|
Device → DescriptorPool, DescriptorSetLayout, DescriptorSet
|
||||||
|
Device → Fence, Semaphore, SwapChain
|
||||||
|
```
|
||||||
|
|
||||||
|
**Examples:**
|
||||||
|
- [orhi](https://github.com/adriengivry/orhi) — C++20, Vulkan/D3D12/Metal, MIT
|
||||||
|
- [SnapRHI](https://github.com/Snapchat/SnapRHI) — C++20, Metal/Vulkan/OpenGL, Apache 2.0
|
||||||
|
- [NVRHI](https://github.com/NVIDIAGameWorks/nvrhi) — C++14, Vulkan/D3D12, NVIDIA
|
||||||
|
- [RGL](https://github.com/RavEngine/RGL) — C++20, Vulkan/D3D12/Metal
|
||||||
|
|
||||||
|
**Pros:**
|
||||||
|
- Familiar mapping to Vulkan/D3D12 concepts
|
||||||
|
- Easy to add new backends
|
||||||
|
- Each object owns its lifetime explicitly
|
||||||
|
|
||||||
|
**Cons:**
|
||||||
|
- Boilerplate-heavy
|
||||||
|
- API surface grows with each backend quirk exposed
|
||||||
|
|
||||||
|
### 2.2 Command-List-Oriented RHI (Adept Engine, Unreal Engine)
|
||||||
|
|
||||||
|
Commands are recorded into command-list objects. The renderer records draws,
|
||||||
|
bindings, and state changes into command lists which are then submitted to the
|
||||||
|
GPU. The RHI thread translates these into API-specific calls.
|
||||||
|
|
||||||
|
```
|
||||||
|
Renderer → RHI Command List → RHI Thread → Backend (VkCmdBuf / ID3D12GraphicsCommandList)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Pros:**
|
||||||
|
- Natural threading model (record in parallel, submit once)
|
||||||
|
- Easy to defer and reorder commands
|
||||||
|
- Close to D3D12/Vulkan command buffer semantics
|
||||||
|
|
||||||
|
**Cons:**
|
||||||
|
- More indirection
|
||||||
|
- State shadowing complexity
|
||||||
|
|
||||||
|
### 2.3 Immediate-Mode RHI (VRHI, NVRHI immediate mode)
|
||||||
|
|
||||||
|
Functions execute synchronously. No command list abstraction — the API is
|
||||||
|
called directly. Simpler but less performant for multi-threaded recording.
|
||||||
|
|
||||||
|
**Prism choice:** Object-based + command-list-oriented. For a compositing
|
||||||
|
application the graph evaluation can record node commands into per-frame
|
||||||
|
command buffers.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Core API Surface (what every RHI needs)
|
||||||
|
|
||||||
|
| Category | Objects | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| **Instance/Device** | `PrRhiInstance`, `PrRhiDevice`, `PrRhiPhysicalDevice` | Instance owns debug + layers; Device owns queues + memory |
|
||||||
|
| **Swap chain** | `PrRhiSwapChain` | Presentation surface + frame sync |
|
||||||
|
| **Resources** | `PrRhiBuffer`, `PrRhiTexture`, `PrRhiSampler` | GPU memory, sub-allocated via VMA-like pattern |
|
||||||
|
| **Shaders** | `PrRhiShader` | Slang → SPIR-V → `VkShaderModule` |
|
||||||
|
| **Pipeline** | `PrRhiPipelineLayout`, `PrRhiPipeline`, `PrRhiComputePipeline` | Compiled shader + vertex layout + state |
|
||||||
|
| **Descriptors** | `PrRhiDescriptorPool`, `PrRhiDescriptorSetLayout`, `PrRhiDescriptorSet` | Bindless or bindful |
|
||||||
|
| **Commands** | `PrRhiCommandPool`, `PrRhiCommandBuffer` | Per-frame recording |
|
||||||
|
| **Sync** | `PrRhiFence`, `PrRhiSemaphore` | CPU-GPU and GPU-GPU sync |
|
||||||
|
| **Query** | `PrRhiQueryPool` | Timestamps, occlusion |
|
||||||
|
|
||||||
|
### Minimal surface for Prism MVP
|
||||||
|
|
||||||
|
For a node-based compositor that renders images via shader passes, the minimum
|
||||||
|
API surface is:
|
||||||
|
|
||||||
|
```
|
||||||
|
Device
|
||||||
|
├── CommandPool → CommandBuffer
|
||||||
|
├── Buffer (vertex, index, uniform/staging)
|
||||||
|
├── Texture (read, write, render-target)
|
||||||
|
├── Sampler
|
||||||
|
├── Shader (from SPIR-V)
|
||||||
|
├── PipelineLayout + Pipeline (graphics)
|
||||||
|
├── DescriptorSetLayout + DescriptorSet (or push descriptors)
|
||||||
|
├── Fence
|
||||||
|
└── SwapChain (for output display)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Design Decisions for Prism
|
||||||
|
|
||||||
|
### 4.1 C11 / C++11 with wapp allocators
|
||||||
|
|
||||||
|
The project follows C11/C++11 dual-mode from `src/wapp/`. The RHI should:
|
||||||
|
|
||||||
|
- Use `WpAllocator *` for all allocations (no `new`/`delete` or raw malloc)
|
||||||
|
- Expose opaque handle types (`PrRhiBuffer` as struct, not `VkBuffer`)
|
||||||
|
- Keep the backend implementation in separate `.c` files per API
|
||||||
|
- Use `wp_extern`/`wp_intern`/`wp_persist` conventions
|
||||||
|
|
||||||
|
All existing wapp infrastructure (arena allocators, arrays, queues, string
|
||||||
|
types) should be used throughout.
|
||||||
|
|
||||||
|
### 4.2 Vulkan-only for now, but design for multi-backend
|
||||||
|
|
||||||
|
The AGENTS.md says "Vulkan, abstracted behind an RHI". The interface should be
|
||||||
|
designed so that a D3D12 or Metal backend could be added later without changing
|
||||||
|
the renderer. This means:
|
||||||
|
|
||||||
|
- Backend-agnostic types in the public header (`pr_rhi.h`)
|
||||||
|
- Backend-specific implementations in `rhi/vulkan/`, `rhi/d3d12/` etc.
|
||||||
|
- A factory pattern or compile-time dispatch for backend selection
|
||||||
|
- No Vulkan types in the public RHI API
|
||||||
|
|
||||||
|
API objects that need per-backend variance:
|
||||||
|
- **Object creation/teardown** (always differs)
|
||||||
|
- **Shader compilation** (SPIR-V is universal, but creation paths differ)
|
||||||
|
- **Pipeline state** (VkPipeline vs ID3D12PipelineState)
|
||||||
|
- **Command recording** (VkCmdBuf vs ID3D12GraphicsCommandList)
|
||||||
|
- **Memory management** (VkDeviceMemory vs ID3D12Heap)
|
||||||
|
|
||||||
|
### 4.3 Explicit over implicit
|
||||||
|
|
||||||
|
The RHI should not hide Vulkan's explicit nature. If the renderer needs to
|
||||||
|
manage descriptor sets, layout transitions, and fences, the RHI should expose
|
||||||
|
those operations — not paper over them with OpenGL-style "bind and forget."
|
||||||
|
|
||||||
|
### 4.4 Memory management: use VMA
|
||||||
|
|
||||||
|
Vulkan Memory Allocator (VMA) from AMD is the de-facto standard for Vulkan
|
||||||
|
memory management. Rather than writing our own sub-allocator, we should:
|
||||||
|
|
||||||
|
- Use VMA for host+device memory allocation
|
||||||
|
- Wrap it behind the RHI so backends can swap it out
|
||||||
|
- Expose `PrRhiAllocation` as an opaque handle
|
||||||
|
|
||||||
|
### 4.5 Descriptor management
|
||||||
|
|
||||||
|
For a compositor, the number of unique descriptors per frame is bounded by the
|
||||||
|
node graph size. Two approaches:
|
||||||
|
|
||||||
|
**A) Push descriptors** (Vulkan 1.0+, no pool needed)
|
||||||
|
- Limited to `maxPushDescriptors` (typically 32-256)
|
||||||
|
- Simple — inline with command recording
|
||||||
|
- Good for small numbers of parameters per node
|
||||||
|
|
||||||
|
**B) Descriptor sets with per-frame pools**
|
||||||
|
- More flexible for many resources
|
||||||
|
- Requires pool management and reset
|
||||||
|
- Better for texture-heavy graphs
|
||||||
|
|
||||||
|
**Recommendation:** Use push descriptors for uniforms, small descriptor set
|
||||||
|
pools for sampled textures (images). Start with descriptor set approach since
|
||||||
|
it scales better.
|
||||||
|
|
||||||
|
### 4.6 Pipeline management
|
||||||
|
|
||||||
|
Pipelines in Vulkan are expensive to create. Strategy:
|
||||||
|
|
||||||
|
- Hash pipeline state (shaders, blend mode, depth, etc.) → cache
|
||||||
|
- Create pipelines lazily on first use
|
||||||
|
- Store in a lock-free hash table (or arena-backed sorted array for
|
||||||
|
single-threaded graph eval)
|
||||||
|
- Use pipeline libraries (`VK_EXT_graphics_pipeline_library`) for faster
|
||||||
|
creation when available
|
||||||
|
|
||||||
|
For a compositor, the number of distinct pipeline configurations is small
|
||||||
|
(blend modes, colour-grade LUTs, blit, etc.), so a simple hash map suffices.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Vulkan-Specific Considerations
|
||||||
|
|
||||||
|
### 5.1 Queue selection
|
||||||
|
|
||||||
|
| Queue type | Usage in compositor |
|
||||||
|
|---|---|
|
||||||
|
| Graphics | Main rendering (draw calls) |
|
||||||
|
| Compute | Image processing, convolution, colour-grade |
|
||||||
|
| Transfer | Image upload from disk, staging |
|
||||||
|
|
||||||
|
The device should expose at least one graphics queue. If separate compute
|
||||||
|
queues are available, use them for async processing. Transfer queue is
|
||||||
|
desirable for texture loading without stalling the render loop.
|
||||||
|
|
||||||
|
### 5.2 Command buffer strategy
|
||||||
|
|
||||||
|
Two-level approach:
|
||||||
|
- **Per-frame primary command buffers**: one per swap-chain image, filled by
|
||||||
|
graph evaluation
|
||||||
|
- **One-shot secondary command buffers**: for transient operations (texture
|
||||||
|
upload, blits) using `immediate_submit` pattern
|
||||||
|
|
||||||
|
Command pools should be per-frame to allow reset without synchronisation.
|
||||||
|
|
||||||
|
### 5.3 Synchronisation
|
||||||
|
|
||||||
|
- `VkSemaphore` for swap-chain acquire/present
|
||||||
|
- `VkFence` for CPU-GPU sync (frame completion, upload completion)
|
||||||
|
- Timeline semaphores (`VK_KHR_timeline_semaphore`) if compute queue is used
|
||||||
|
async
|
||||||
|
|
||||||
|
### 5.4 Image layouts
|
||||||
|
|
||||||
|
For a compositor where images flow through nodes:
|
||||||
|
- `VK_IMAGE_LAYOUT_SHADER_READ_ONLY_OPTIMAL` — node inputs
|
||||||
|
- `VK_IMAGE_LAYOUT_COLOR_ATTACHMENT_OPTIMAL` — node render targets
|
||||||
|
- `VK_IMAGE_LAYOUT_GENERAL` — storage images (compute nodes)
|
||||||
|
- `VK_IMAGE_LAYOUT_PRESENT_SRC_KHR` — final output
|
||||||
|
|
||||||
|
Transitions happen via explicit barriers in the command buffer (or via
|
||||||
|
`VK_KHR_synchronization2`). The RHI should expose barrier helpers.
|
||||||
|
|
||||||
|
### 5.5 Debug / validation layers
|
||||||
|
|
||||||
|
- Load `VK_LAYER_KHRONOS_validation` in debug builds
|
||||||
|
- Use `VK_EXT_debug_utils` for object naming
|
||||||
|
- Enable GPU-assisted validation for shader issues
|
||||||
|
- Consider RenderDoc for frame debugging
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Slang Shader Integration
|
||||||
|
|
||||||
|
### 6.1 Why Slang
|
||||||
|
|
||||||
|
- HLSL/GLSL compatible syntax
|
||||||
|
- Module system for shared shading code (colour science, maths)
|
||||||
|
- Single source for multiple stages (vertex+fragment in one file)
|
||||||
|
- SPIR-V output (directly consumable by Vulkan)
|
||||||
|
- Rich reflection API (bindings, buffer layouts, entry points)
|
||||||
|
- Active development, Khronos exploratory forum
|
||||||
|
|
||||||
|
### 6.2 Compilation pipeline
|
||||||
|
|
||||||
|
```
|
||||||
|
.slang file
|
||||||
|
→ slangc (offline) or libslang (runtime)
|
||||||
|
→ SPIR-V binary
|
||||||
|
→ vkCreateShaderModule
|
||||||
|
→ PrRhiShader
|
||||||
|
```
|
||||||
|
|
||||||
|
Options:
|
||||||
|
- **Offline**: Pre-compile `.slang` → `.spv` at build time. Simpler, no runtime
|
||||||
|
compiler dependency. Good for shipped shaders.
|
||||||
|
- **Runtime**: Use libslang to compile at app startup (or on first use).
|
||||||
|
Enables shader hot-reload during development.
|
||||||
|
|
||||||
|
**Recommendation:** Offline for release, runtime for debug/dev. The
|
||||||
|
nvpro-samples `vk_slang_editor` demonstrates both approaches.
|
||||||
|
|
||||||
|
### 6.3 Reflection-driven pipeline creation
|
||||||
|
|
||||||
|
Slang's reflection API (`slang::ProgramLayout`) provides:
|
||||||
|
- Binding locations (set, binding, space)
|
||||||
|
- Buffer member offsets and sizes
|
||||||
|
- Entry point names and stage types
|
||||||
|
- Specialisation constant info
|
||||||
|
|
||||||
|
The RHI can use this to automatically build:
|
||||||
|
- `VkDescriptorSetLayout` from declared bindings
|
||||||
|
- `VkPipelineLayout` from descriptor set layouts + push constants
|
||||||
|
- Push constant ranges from reflected constant buffers
|
||||||
|
|
||||||
|
See the [Slang Reflection API docs](https://shader-slang.com/slang/user-guide/reflection)
|
||||||
|
and the [vk_slang_editor](https://github.com/nvpro-samples/vk_slang_editor)
|
||||||
|
source for concrete patterns.
|
||||||
|
|
||||||
|
### 6.4 Shader organisation for a compositor
|
||||||
|
|
||||||
|
```
|
||||||
|
src/shaders/
|
||||||
|
├── common/
|
||||||
|
│ ├── math.slang — Matrix/vector utilities
|
||||||
|
│ ├── colour.slang — Colour space conversions
|
||||||
|
│ └── compositing.slang — Blend equations, alpha handling
|
||||||
|
├── blit.slang — Full-screen quad draw
|
||||||
|
├── blend.slang — Over/under/add blend modes
|
||||||
|
├── grade.slang — Colour grading (lift/gamma/gain)
|
||||||
|
├── blur.slang — Separable gaussian blur
|
||||||
|
└── read.slang — Simple texture passthrough
|
||||||
|
```
|
||||||
|
|
||||||
|
Each shader file contains both vertex and fragment stages:
|
||||||
|
|
||||||
|
```slang
|
||||||
|
// blit.slang
|
||||||
|
[shader("vertex")]
|
||||||
|
void vs_main(...) { ... }
|
||||||
|
|
||||||
|
[shader("fragment")]
|
||||||
|
void fs_main(...) { ... }
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Reference Projects
|
||||||
|
|
||||||
|
| Project | Language | APIs | Notable features |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **[orhi](https://github.com/adriengivry/orhi)** | C++20 | Vulkan, D3D12, Metal (planned) | Clean object hierarchy, CMake, MIT |
|
||||||
|
| **[SnapRHI](https://github.com/Snapchat/SnapRHI)** | C++20 | Metal, Vulkan, OpenGL/ES | Compile-switchable validation (if constexpr), aggressive pooling |
|
||||||
|
| **[tobyc11/RHI](https://github.com/tobyc11/RHI)** | C++ | Vulkan, D3D11 | SPIR-V as common shader format, SPIRV-Cross for translation |
|
||||||
|
| **[NVRHI](https://github.com/NVIDIAGameWorks/nvrhi)** | C++14 | Vulkan, D3D12 | Production-grade, NVIDIA maintained, header-only-ish API |
|
||||||
|
| **[RGL](https://github.com/RavEngine/RGL)** | C++20 | Vulkan, D3D12, Metal | Thin wrapper, focuses on simplicity |
|
||||||
|
| **[The Forge](https://github.com/ConfettiFX/The-Forge)** | C99/C++11 | All major APIs | Cross-platform, used in shipping games, FS |
|
||||||
|
| **[O3DE Atom RHI](https://docs.o3de.org/docs/atom-guide/dev-guide/rhi/rhi/)** | C++17 | Vulkan, D3D12, Metal | Full-featured engine RHI, frame scheduler, multi-threaded |
|
||||||
|
| **[Magma](https://github.com/vcoda/magma)** | C++17 | Vulkan | C++ abstraction, uses VMA, SPIR-V reflection |
|
||||||
|
| **[rafx](https://github.com/zeozeozeo/rafx)** | C/C++ | Vulkan, D3D12 | C API (good FFI), explicit design |
|
||||||
|
|
||||||
|
### What to borrow from each
|
||||||
|
|
||||||
|
| Project | Lesson |
|
||||||
|
|---|---|
|
||||||
|
| **orhi** | Object hierarchy + backend-agnostic headers pattern |
|
||||||
|
| **SnapRHI** | Compile-switchable validation; per-frame resource pooling |
|
||||||
|
| **NVRHI** | Header-only-ish API with implementation in .cpp |
|
||||||
|
| **The Forge** | C99-friendly, explicit API with minimal hidden state |
|
||||||
|
| **O3DE Atom** | Frame scheduler concept (render passes as graph nodes) |
|
||||||
|
| **RGL** | Simplicity — don't over-abstract |
|
||||||
|
| **Magma** | VMA integration pattern + SPIR-V reflection |
|
||||||
|
| **rafx** | C API design (relevant since Prism is C11) |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Proposed Architecture for Prism
|
||||||
|
|
||||||
|
### 8.1 Directory layout
|
||||||
|
|
||||||
|
```
|
||||||
|
src/prism/
|
||||||
|
├── rhi/
|
||||||
|
│ ├── pr_rhi.h ← Public API (backend-agnostic)
|
||||||
|
│ ├── pr_rhi_types.h ← Shared types (PrRhiBufferDesc, etc.)
|
||||||
|
│ ├── vulkan/
|
||||||
|
│ │ ├── pr_rhi_vk.h ← Vulkan backend internal header
|
||||||
|
│ │ ├── pr_rhi_vk_device.c
|
||||||
|
│ │ ├── pr_rhi_vk_buffer.c
|
||||||
|
│ │ ├── pr_rhi_vk_texture.c
|
||||||
|
│ │ ├── pr_rhi_vk_shader.c
|
||||||
|
│ │ ├── pr_rhi_vk_pipeline.c
|
||||||
|
│ │ ├── pr_rhi_vk_descriptor.c
|
||||||
|
│ │ ├── pr_rhi_vk_command.c
|
||||||
|
│ │ └── pr_rhi_vk_swapchain.c
|
||||||
|
│ └── d3d12/ ← (future)
|
||||||
|
└── ...
|
||||||
|
```
|
||||||
|
|
||||||
|
### 8.2 Object lifecycle pattern
|
||||||
|
|
||||||
|
```c
|
||||||
|
// Creation: takes an allocator + device + desc, returns handle
|
||||||
|
PrRhiBuffer *prRhiCreateBuffer(PrRhiDevice *device, const PrRhiBufferDesc *desc,
|
||||||
|
WpAllocator *alloc);
|
||||||
|
|
||||||
|
// Destruction: frees all GPU resources + backing memory
|
||||||
|
void prRhiDestroyBuffer(PrRhiBuffer *buffer, WpAllocator *alloc);
|
||||||
|
|
||||||
|
// Usage: command buffer records operations on handles
|
||||||
|
void prRhiCmdCopyBuffer(PrRhiCommandBuffer *cb,
|
||||||
|
PrRhiBuffer *src, PrRhiBuffer *dst);
|
||||||
|
```
|
||||||
|
|
||||||
|
### 8.3 Backend dispatch (compile-time)
|
||||||
|
|
||||||
|
```c
|
||||||
|
// pr_rhi.h
|
||||||
|
typedef struct PrRhiDevice PrRhiDevice;
|
||||||
|
struct PrRhiDevice {
|
||||||
|
PrRhiDeviceVtbl *vtbl; // function pointer table
|
||||||
|
void *backend; // VkDevice or ID3D12Device
|
||||||
|
};
|
||||||
|
|
||||||
|
// Each backend fills the vtbl
|
||||||
|
typedef struct PrRhiDeviceVtbl {
|
||||||
|
PrRhiBuffer* (*createBuffer)(PrRhiDevice*, const PrRhiBufferDesc*, WpAllocator*);
|
||||||
|
void (*destroyBuffer)(PrRhiBuffer*, WpAllocator*);
|
||||||
|
// ... etc
|
||||||
|
} PrRhiDeviceVtbl;
|
||||||
|
```
|
||||||
|
|
||||||
|
This gives zero-cost abstraction (pointer indirection on calls) while keeping
|
||||||
|
the API clean. The renderer calls `device->vtbl->createBuffer(...)` and the
|
||||||
|
backend resolves the call.
|
||||||
|
|
||||||
|
### 8.4 Frame lifecycle
|
||||||
|
|
||||||
|
```
|
||||||
|
Loop:
|
||||||
|
1. prRhiAcquireNextImage(swapchain) → image index, semaphore
|
||||||
|
2. prRhiResetCommandPool(pool, frame_idx) → recycles command buffers
|
||||||
|
3. For each node in topo-sorted graph:
|
||||||
|
a. prRhiCmdBindPipeline(cb, pipeline)
|
||||||
|
b. prRhiCmdBindDescriptorSets(cb, ...)
|
||||||
|
c. prRhiCmdPushConstants(cb, ...)
|
||||||
|
d. prRhiCmdDraw(cb, ...)
|
||||||
|
4. prRhiQueueSubmit(queue, cb, wait_sem, signal_sem, fence)
|
||||||
|
5. prRhiPresent(swapchain, signal_sem)
|
||||||
|
6. prRhiWaitForFence(fence) → CPU-GPU sync
|
||||||
|
```
|
||||||
|
|
||||||
|
### 8.5 Static allocation strategy
|
||||||
|
|
||||||
|
Following wapp conventions and the project's data-oriented design principles:
|
||||||
|
|
||||||
|
- **Command pools**: one per swap-chain image (2-3), allocated once
|
||||||
|
- **Descriptor pools**: one per frame, reset each frame
|
||||||
|
- **Upload buffers**: ring buffer for staging data, bumped each frame
|
||||||
|
- **Pipeline cache**: arena-backed hash table, populated lazily
|
||||||
|
- **Scratch buffers**: arena-allocated in the per-frame scratch space
|
||||||
|
|
||||||
|
No dynamic allocation on the hot path — all per-frame memory comes from
|
||||||
|
frame-local arena allocators that are reset at the start of each frame.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Open Questions
|
||||||
|
|
||||||
|
1. **Multi-queue**: Should the RHI expose separate compute/transfer queues, or
|
||||||
|
keep everything on a single graphics queue and serialise? For an MVP, single
|
||||||
|
queue is simpler and likely sufficient.
|
||||||
|
|
||||||
|
2. **Bindless vs bindful**: Bindless descriptors (VK_EXT_descriptor_indexing)
|
||||||
|
simplify shader resource access but require higher Vulkan version. For
|
||||||
|
maximum compatibility, start with bindful descriptor sets.
|
||||||
|
|
||||||
|
3. **Shader compilation**: Use `slangc` at build time and ship SPIR-V, or link
|
||||||
|
libslang for runtime compilation + reflection? Runtime enables hot-reload
|
||||||
|
but adds ~15MB to binary size. Recommendation: both — offline for release,
|
||||||
|
runtime for debug.
|
||||||
|
|
||||||
|
4. **Swap chain**: Headless mode (no window) for batch/compute-only operation?
|
||||||
|
Useful for a compositor that renders to a file. The RHI should support
|
||||||
|
both windowed and headless modes.
|
||||||
|
|
||||||
|
5. **Vulkan version**: Target Vulkan 1.3 (widely available on desktop, adds
|
||||||
|
timeline semaphores, dynamic rendering, and sync2) with fallback to 1.2.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [O3DE Atom RHI Overview](https://docs.o3de.org/docs/atom-guide/dev-guide/rhi/rhi/)
|
||||||
|
- [Adept Engine RHI Design](https://andrewcjp.wordpress.com/2019/11/09/designing-a-render-hardware-interface-for-explicit-multi-gpu-programming/)
|
||||||
|
- [Unreal Engine RHI Architecture](https://dev.epicgames.com/documentation/unreal-engine/parallel-rendering-overview-for-unreal-engine)
|
||||||
|
- [orhi — OpenRHI](https://github.com/adriengivry/orhi)
|
||||||
|
- [SnapRHI](https://github.com/Snapchat/SnapRHI)
|
||||||
|
- [NVRHI](https://github.com/NVIDIAGameWorks/nvrhi)
|
||||||
|
- [The Forge](https://github.com/ConfettiFX/The-Forge)
|
||||||
|
- [RGL](https://github.com/RavEngine/RGL)
|
||||||
|
- [tobyc11/RHI](https://github.com/tobyc11/RHI)
|
||||||
|
- [rafx](https://github.com/zeozeozeo/rafx)
|
||||||
|
- [Magma](https://github.com/vcoda/magma)
|
||||||
|
- [Vulkan Memory Allocator](https://github.com/GPUOpen-LibrariesAndSDKs/VulkanMemoryAllocator)
|
||||||
|
- [Slang Shading Language](https://github.com/shader-slang/slang)
|
||||||
|
- [Slang Reflection API](https://shader-slang.com/slang/user-guide/reflection)
|
||||||
|
- [vk_slang_editor](https://github.com/nvpro-samples/vk_slang_editor)
|
||||||
|
- [Vulkan in 30 minutes](https://renderdoc.org/vulkan-in-30-minutes.html)
|
||||||
|
- [Vulkan Memory Management Guide](https://docs.vulkan.org/guide/latest/memory_allocation.html)
|
||||||
|
- [Khronos Vulkan Spec — Command Buffers](https://docs.vulkan.org/spec/latest/chapters/cmdbuffers.html)
|
||||||
@@ -1,8 +1,17 @@
|
|||||||
# Useful resources
|
# Useful resources
|
||||||
|
|
||||||
|
## Graphs
|
||||||
|
|
||||||
- [The Algorithm Design Manual](https://sureshcseit.wordpress.com/wp-content/uploads/2021/04/skienathealgorithmdesignmanual.pdf)
|
- [The Algorithm Design Manual](https://sureshcseit.wordpress.com/wp-content/uploads/2021/04/skienathealgorithmdesignmanual.pdf)
|
||||||
- [Topological Sorting using BFS (Kahn's Algorithm)](https://www.geeksforgeeks.org/topological-sorting-indegree-based-solution/)
|
- [Topological Sorting using BFS (Kahn's Algorithm)](https://www.geeksforgeeks.org/topological-sorting-indegree-based-solution/)
|
||||||
- [Kahn's Algorithm Explained with Code & Examples](https://dev.to/rui_jiang/kahns-algorithm-for-topological-sorting-explained-with-code-examples-2if5)
|
- [Kahn's Algorithm Explained with Code & Examples](https://dev.to/rui_jiang/kahns-algorithm-for-topological-sorting-explained-with-code-examples-2if5)
|
||||||
- [Understanding Kahn's Algorithm for Topological Sorting](https://blog.devgenius.io/dsa-kahns-algorithm-for-topological-sorting-33c8587985a1)
|
- [Understanding Kahn's Algorithm for Topological Sorting](https://blog.devgenius.io/dsa-kahns-algorithm-for-topological-sorting-33c8587985a1)
|
||||||
- [Detect a Cycle in Directed Graph](https://takeuforward.org/data-structure/detect-a-cycle-in-directed-graph-topological-sort-kahns-algorithm-g-23)
|
- [Detect a Cycle in Directed Graph](https://takeuforward.org/data-structure/detect-a-cycle-in-directed-graph-topological-sort-kahns-algorithm-g-23)
|
||||||
- [Kahn's Algorithm](https://leetcodethehardway.com/tutorials/graph-theory/kahns-algorithm)
|
- [Kahn's Algorithm](https://leetcodethehardway.com/tutorials/graph-theory/kahns-algorithm)
|
||||||
|
|
||||||
|
## Rendering Hardware Interface (RHI)
|
||||||
|
|
||||||
|
- [O3DE Atom RHI Overview](https://docs.o3de.org/docs/atom-guide/dev-guide/rhi/rhi/)
|
||||||
|
- [Adept Engine RHI Design](https://andrewcjp.wordpress.com/2019/11/09/designing-a-render-hardware-interface-for-explicit-multi-gpu-programming/)
|
||||||
|
- [Unreal Engine RHI Architecture](https://dev.epicgames.com/documentation/unreal-engine/parallel-rendering-overview-for-unreal-engine)
|
||||||
|
- [NVRHI](https://github.com/NVIDIAGameWorks/nvrhi) — NVIDIA's production RHI (Vulkan/D3D12), reference API design
|
||||||
|
|||||||
Reference in New Issue
Block a user