Document Forward+ measurements and Linux lighting evidence

This commit is contained in:
Emil
2026-09-24 03:49:06 +03:00
parent 609ba632f4
commit 77404a1035
15 changed files with 1773 additions and 40 deletions
+65
View File
@@ -147,6 +147,71 @@ and reads back the full image, so `cpu_ms` is wall time including waits, not CPU
utilization. An open scene can run slower with HZB; visibility correctness and
full-frame speed are separate findings.
## Measure P3 lighting and shadows
A Player `--profile` sample includes `effective_lighting_path`, local lights
submitted/omitted, requested/effective sun cascades, requested/rasterized local
shadow faces, tile use, shadow drop reasons, caster draws, and explicit atlas
allocation bytes. `gpu_main_raster_ms`, `gpu_sun_shadow_ms`, and
`gpu_local_shadow_ms` are GPU timestamps or `null` when timestamps are
unavailable. A light can illuminate while its shadow faces are dropped. A
submitted-light count of zero is a different workload from 128 lights whose
shadows are disabled. See [Lighting](lighting.md) for the capacity policy and
[Diagnostics](diagnostics.md) for the Editor counters.
The same sample includes `effective_lighting_path` (`forward` or `tiled`),
`gpu_light_tiles_ms`, and `light_tile_count`. Stored candidate and overflow
counts are present only when visibility diagnostics readback was enabled;
`light_tile_counts_valid: false` means their `null` values are unavailable,
not zero. The normal `Auto` setting currently resolves to `forward` after the
fixed dense 1080p benchmark showed that tile construction cost outweighed its
raster savings. A C++ renderer integration can explicitly request `Tiled` for
a localized-light scene, then check the actual path before comparing timings.
The fixed-scene benchmark compares 0, 4, 16, 32, 64, and 128 local lights under
Direct, GPU frustum, and GPU occlusion visibility, with shadows on and off. Its
wrapper runs three independent 1920×1080 repetitions per configuration, each
with ten warm-up and thirty recorded frames. First inspect the planned matrix:
```sh
python3 tools/benchmark_p3_lighting.py --list-runs
```
From the repository, after a Linux Release renderer build, run one shadow setting
into a new output directory. Supply the actual device driver identity:
```sh
python3 tools/benchmark_p3_lighting.py --sweep \
--executable build/linux-release/faset_p3_lighting_benchmark \
--output .cache/p3-lighting-off \
--shadows off --driver 'REPLACE_WITH_ACTUAL_DRIVER' --validation off
```
The wrapper writes one raw CSV per run, `merged.csv`, and `summary.json`. Keep
all three with the exact source revision and device. It checks that every run
used its requested visibility mode and submitted every requested light. GPU
timestamps for the main raster isolate fragment-heavy lighting better than
renderer wall time, which includes GPU waits and synchronous readback. Shadow
time is split into sun and local GPU durations. The Forward+ decision compares
the median of three run medians against the matching zero-light configuration;
the threshold is **1.0 ms extra main raster time or 15% of the zero-light GPU
frame** at 32, 64, or 128 lights on the Linux physical reference GPU. The
[P3 lighting validation record](https://github.com/emil28092005/Faset_Engine/blob/main/docs/validation/p3-lighting-2026-09-24/README.md)
states the measured decision and scope. A software Vulkan run checks
functionality, not physical GPU performance.
For a direct comparison of the two algorithms on the same scene, invoke the
Release executable twice with `--lighting forward` and `--lighting tiled`,
using the same `--lights`, `--shadows`, `--visibility`, and output size. The
default `--light-layout dense` preserves the fixed benchmark scene;
`--light-layout localized` reduces point-light ranges to 1.75 units as a
separately labelled workload. Compare `gpu_build_plus_raster_ms`, which includes
`gpu_light_tiles_ms`, rather than raster time alone. One optional diagnostic
frame with `--tile-diagnostics on` reports candidate and overflow counts but
adds a GPU readback, so do not mix it into the timed runs. The
[Forward+ measurement](https://github.com/emil28092005/Faset_Engine/blob/main/docs/studies/23-p3-forward-plus-2026-09-24.md) retains
raw frames, shader hashes, and the decision.
## Current performance scope
The accepted MVP path uses direct draws and CPU culling; P2 adds optional GPU