Integrate Forward+ lighting with temporal rendering
Native and manual checks / native (ubuntu-24.04) (push) Failing after 34s
Native and manual checks / manual (push) Successful in 28s
Windows editor and software Vulkan / windows-graphics (push) Canceled after 0s
Native and manual checks / native (windows-2025) (push) Canceled after 0s

This commit is contained in:
Emil
2026-09-24 04:05:58 +03:00
27 changed files with 1025 additions and 85 deletions
+52 -6
View File
@@ -144,13 +144,59 @@ quality, inspect a still thin edge, a slow pan and a newly uncovered surface, an
compare the same frame against Off. See [Temporal rendering](temporal.md) for
mode controls and native C++ configuration.
## Measure P3 lighting and shadows
A Player `--profile` sample includes `effective_lighting_path`, local lights
submitted/omitted, requested/effective sun cascades, requested/rasterized local
shadow faces, tile use, shadow drop reasons, caster draws, and explicit atlas
allocation bytes. `gpu_main_raster_ms`, `gpu_sun_shadow_ms`, and
`gpu_local_shadow_ms` are GPU timestamps or `null` when timestamps are
unavailable. A light can illuminate while its shadow faces are dropped. A
submitted-light count of zero is a different workload from 128 lights whose
shadows are disabled. See [Lighting](lighting.md) for the capacity policy and
[Diagnostics](diagnostics.md) for the Editor counters.
The fixed-scene benchmark compares 0, 4, 16, 32, 64, and 128 local lights under
Direct, GPU frustum, and GPU occlusion visibility, with shadows on and off. Its
wrapper runs three independent 1920×1080 repetitions per configuration, each
with ten warm-up and thirty recorded frames. First inspect the planned matrix:
```sh
python3 tools/benchmark_p3_lighting.py --list-runs
```
From the repository, after a Linux Release renderer build, run one shadow setting
into a new output directory. Supply the actual device driver identity:
```sh
python3 tools/benchmark_p3_lighting.py --sweep \
--executable build/linux-release/faset_p3_lighting_benchmark \
--output .cache/p3-lighting-off \
--shadows off --driver 'REPLACE_WITH_ACTUAL_DRIVER' --validation off
```
The wrapper writes one raw CSV per run, `merged.csv`, and `summary.json`. Keep
all three with the exact source revision and device. It checks that every run
used its requested visibility mode and submitted every requested light. GPU
timestamps for the main raster isolate fragment-heavy lighting better than
renderer wall time, which includes GPU waits and synchronous readback. Shadow
time is split into sun and local GPU durations. The Forward+ decision compares
the median of three run medians against the matching zero-light configuration;
the threshold is **1.0 ms extra main raster time or 15% of the zero-light GPU
frame** at 32, 64, or 128 lights on the Linux physical reference GPU. The
[P3 lighting validation record](https://github.com/emil28092005/Faset_Engine/blob/main/docs/validation/p3-lighting-2026-09-24/README.md)
states the measured decision and scope. A software Vulkan run checks
functionality, not physical GPU performance.
## Current performance scope
The accepted MVP path uses direct draws and CPU culling; P2 adds optional GPU
visibility for opaque static meshes, with prepared LODs supplied by the project.
Both paths currently use one graphics queue and synchronous full-image
capture/readback. Use measurements to find the next bottleneck before introducing
parallel jobs or expanding GPU-driven rendering. Neither an offscreen capture
benchmark nor a tiny demo is a promise of a production frame budget. Observed
measurements and follow-up targets belong in the implementation acceptance report
with their source revision and method.
P3 adds local lights and bounded sun/local shadow atlases. The benchmark's
`lighting_path` and a Player profile's `effective_lighting_path` identify the
algorithm actually used. Both paths currently use one graphics queue and
synchronous full-image capture/readback. Use measurements to find the next
bottleneck before introducing parallel jobs or expanding GPU-driven rendering.
Neither an offscreen capture benchmark nor a tiny demo is a promise of a
production frame budget. Observed measurements and follow-up targets belong in
the implementation acceptance report with their source revision and method.