6.2 KiB
P3 Forward+ experiment: correctness and cost on localized lights
The fixed 1920×1080 P3 benchmark crossed the agreed threshold for trying
Forward+. A depth-free 16×16 tiled implementation now exists, but the measured
build + raster cost is higher than a full light scan on that benchmark's
dense lights. RendererConfig::lighting_mode = Auto therefore keeps the forward
path. Tiled is an explicit option for scenes whose projected light volumes
are localized. There is no unmeasured automatic occupancy heuristic.
This is a follow-up to the fixed-scene baseline sweep, which is being merged
as a separate study. It compares both paths in the same source revision
a0a4e29d480ed3344f19bd3565d48668ca913fed. The baseline's dense
placement remains the default. An explicit --light-layout localized changes
only point-light range from 8 to 1.75 world units; camera, nine casters,
receiver, positions, colors, light count, and output size are unchanged. The
localized fixture is a separate workload, not a replacement for the fixed
baseline gate.
Renderer behavior and safety
The compute pass builds up to 64 stable-order light indices per screen tile.
It tests each world-space range sphere against four clip-space tile planes.
It does not use depth or reject near-plane intersections. A tile with more than
64 candidates sets an overflow bit; the fragment shader then scans all
submitted lights for that tile. Zero lights, missing capability, excessive
buffer size, failed optional allocation, and Auto use the forward path. The
shader contract checks the new compute entry's descriptors and 96-byte push
constants; Direct and P2 GPU graphics still use materials at set 0, lighting
at set 1, and GPU scene data at set 2. The tile list is set 1 binding 4 in the
shared fragment shader. Sprite/UI shading returns before tile reads.
The Linux Vulkan image test compares forward and tiled output in Direct, GPU frustum, and GPU occlusion modes, including a cropped scene viewport, near-plane crossing point light and shadow, resize, an offscreen light, and 80 coincident lights that exceed tile capacity. Every overflowing tile falls back to the full list. Shader reload preserves a working tiled pipeline after invalid bytecode and recreates it after a valid reload. A separate 1920×1080 capture with 128 localized lights was byte-identical across both paths; its SHA-256 is in the provenance record.
Measurement
The device was NVIDIA GeForce RTX 2080 Ti with NVIDIA driver 595.84.0.0,
Linux Clang Release, Direct visibility, shadows off, 1920×1080. Each mode had
three independent process runs with ten warm-up and thirty measured frames.
Forward/tiled run order alternated. The table uses the median of the three
per-run medians in milliseconds. The tile build column is an actual GPU
timestamp; build + raster also includes post raster if present. The dense
and localized CSVs contain every one of the 1080 measured frames, with a
source_csv identifier. The executable and all loaded .spv/reflection
SHA-256 values are in the provenance record.
| Light layout | Lights | Forward raster | Tile build | Tiled raster | Tiled build + raster | Tiled change |
|---|---|---|---|---|---|---|
| Dense fixed scene | 32 | 0.5500 | 0.0617 | 0.5527 | 0.6144 | +0.0644 ms (11.7% slower) |
| Dense fixed scene | 64 | 1.0701 | 0.1177 | 1.0740 | 1.1921 | +0.1220 ms (11.4% slower) |
| Dense fixed scene | 128 | 2.1172 | 0.2219 | 2.1164 | 2.3388 | +0.2216 ms (10.5% slower) |
| Localized range 1.75 | 32 | 0.2336 | 0.0555 | 0.0758 | 0.1312 | −0.1025 ms (43.9% faster) |
| Localized range 1.75 | 64 | 0.4254 | 0.1060 | 0.1057 | 0.2109 | −0.2144 ms (50.4% faster) |
| Localized range 1.75 | 128 | 0.8094 | 0.2048 | 0.1643 | 0.3691 | −0.4404 ms (54.4% faster) |
At 32 dense lights, the first forward process had a 0.7405 ms run median;
the other two were 0.5488 and 0.5500 ms. A single paired run would have
incorrectly suggested a tiled win. The median of three process medians and a
separate earlier repeat both support the slower dense result. This is why
Auto remains forward despite the localized-scene gain. The total GPU frame
also includes visibility, shadow fallback, copies, and synchronous readback;
the table isolates the passes that the optimization changes. For example, at
32 localized lights the full GPU frame was 1.6494 ms forward and 1.6472 ms
tiled, essentially unchanged despite lower build + raster cost. At 128 it
was 2.2788 versus 1.7948 ms.
One diagnostic frame per layout/count copied the tile buffer after the timed draw. That copy was not enabled in the 1080 performance frames. The grid has 8160 tiles and a 64-index capacity per tile.
| Layout | Lights | Stored candidates across tiles | Overflowed tiles |
|---|---|---|---|
| Dense | 32 | 259,896 | 0 |
| Dense | 64 | 519,792 | 0 |
| Dense | 128 | 522,240 | 8,160 |
| Localized | 32 | 38,237 | 0 |
| Localized | 64 | 76,103 | 0 |
| Localized | 128 | 152,202 | 0 |
The dense 128 candidate count is capped at 64 × 8160 stored slots; all tiles overflow and correctly evaluate all 128 lights in the fragment shader. This explains why paying for tile construction cannot help that frame. The localized 128 scene averages about 19 stored candidates per tile and avoids fallback.
Raw data: all paired frames, diagnostic frames, and binary/shader provenance.
Verification and scope
At the implementation revision, Linux Debug built all targets and passed
62/63 CTests, with the compositor-dependent window lifecycle case skipped and
no failures. The pinned Linux SwiftShader ICD passed all six P3 cases, including
the tiled parity/overflow test. The lighting validation record
retains those logs. These are functional checks on Linux and software Vulkan,
not physical Windows GPU performance. The A/B numbers apply to one GPU, driver,
camera, receiver and two synthetic light layouts. They do not establish an
engine-wide speedup. A measured runtime occupancy predictor and representative
game scenes are prerequisites before changing Auto from forward.