Native and manual checks / native (ubuntu-24.04) (push) Failing after 35s
Native and manual checks / manual (push) Successful in 27s
Native and manual checks / native (windows-2025) (push) Canceled after 0s
Windows editor and software Vulkan / windows-graphics (push) Canceled after 0s
198 lines
11 KiB
Markdown
198 lines
11 KiB
Markdown
# Profiling and measurements
|
||
|
||
Use a Release export to measure the shipping Player. Record the exact scene, hardware,
|
||
driver, build configuration and resolution with the result. Small test scenes do not
|
||
establish performance for a large game.
|
||
|
||
## Capture a bounded Player profile
|
||
|
||
From a standalone generation directory:
|
||
|
||
```sh
|
||
./faset_player --headless --frames 240 --profile profile.json
|
||
```
|
||
|
||
`--headless` here means **offscreen Vulkan rendering**. A GPU/driver is still required.
|
||
The Editor's headless authoring mode is a separate feature. Omit this flag to measure
|
||
the windowed path. `--profile` requires an explicit `--frames` between 1 and 100000,
|
||
which bounds the stored samples.
|
||
|
||
The JSON contains raw completed-frame samples and nearest-rank p50/p95 summaries.
|
||
No warm-up frames are silently removed. It records the presentation mode, device,
|
||
resolution, validation activation, fixed ticks and timestep. A bounded run advances
|
||
one synthetic fixed timestep per frame; it does not reproduce a real-time input
|
||
session. Keep that distinction when comparing runs.
|
||
|
||
Startup starts at the Player application entry after platform argument normalization
|
||
and ends at the first completed frame. OS process loading and Windows `wmain` UTF-8
|
||
argument conversion are excluded. Frame wall times exclude writing the final profile and
|
||
capture files. Simulation and scene-snapshot times are separate from the renderer
|
||
call. Renderer CPU wall duration includes GPU waits and readback; it is **not CPU
|
||
utilization**. GPU timestamps measure the submitted graphics work and can be null
|
||
when timestamps are unsupported.
|
||
|
||
Resource counters report live renderer allocations and texture count. GPU allocation
|
||
bytes include Vulkan allocation alignment and exclude driver-internal memory; they
|
||
are not a whole-process VRAM meter. The fallback white texture is included.
|
||
|
||
Use `--debug-physics` or press **F3** to show current physics box colliders. Debug
|
||
geometry increases draw count, so record whether it was enabled. The collider view
|
||
uses simulation poses; normal visuals can use interpolated poses.
|
||
|
||
## Measure Editor and C++ workflows
|
||
|
||
From the engine repository:
|
||
|
||
```sh
|
||
python3 tools/measure_workflows.py \
|
||
--editor build/linux-debug/faset_editor \
|
||
--project examples/projects/collect-3d \
|
||
--output .cache/my-workflow-measurement
|
||
```
|
||
|
||
Use a new output directory. The tool copies the project, preserving your original,
|
||
and records command startup, two-frame GUI startup/shutdown, first/cached Blender
|
||
bundle import, initial/no-change/changed Debug builds and a subsequent Player frame.
|
||
It checks that editing gameplay makes the schema stale and successful building
|
||
clears that state. The initial build uses available dependency archives and OS
|
||
caches; it is not a measurement of internet download speed.
|
||
|
||
On Linux, GNU `time` records peak RSS for each command and its waited-for children.
|
||
This is a maximum, not the sum of simultaneous compiler processes. Other platforms
|
||
report this field as null unless equivalent measurement support is added. The tool
|
||
keeps raw stdout/stderr, durations, hardware and revision information alongside its
|
||
report. A dirty source checkout is explicitly identified.
|
||
|
||
`tools/verify_playable_exports.py` separately verifies the two sample games in
|
||
relocated Release packages and records their Player profiles. Its assertions test
|
||
correct execution, not a frame-time threshold.
|
||
|
||
## Compare P2 GPU visibility modes
|
||
|
||
The Editor diagnostics panel (**F12**) can switch its current viewport between
|
||
**Direct**, **GPU frustum**, and **GPU occlusion**. Direct is the default reference.
|
||
The selector is an Editor viewport setting; it does not change the saved scene or
|
||
automatically change an exported Player. An exported Player can select a mode for a
|
||
bounded run:
|
||
|
||
```sh
|
||
./faset_player --headless --frames 240 --profile gpu-frustum.json --visibility gpu-frustum
|
||
```
|
||
|
||
Accepted values are `direct`, `gpu-frustum`, and `gpu-occlusion`; Direct is the
|
||
default. The profile records the requested `visibility_mode`, the run's and each
|
||
frame's `effective_visibility_mode`, and each frame's `gpu_visibility_active` state.
|
||
Compare requested and effective modes before interpreting a GPU run: a missing GPU
|
||
profile falls back to Direct, while missing HZB can reduce GPU occlusion to GPU
|
||
frustum. The Editor shows the **Effective path** and any **Fallback from** line. See
|
||
[Diagnostics](diagnostics.md) for the counters and HZB preview.
|
||
|
||
For a repeatable offscreen comparison, build and run the P2 benchmark harness:
|
||
|
||
```sh
|
||
build/linux-debug/faset_render_gpu_acceptance_tests --benchmark /tmp/faset-p2.csv
|
||
```
|
||
|
||
It records 10 warm-up and 30 measured frames for Direct, GPU frustum and GPU
|
||
occlusion in fixed frustum-heavy, open and occluded scenes. Run it three times and
|
||
compare median/p95 by scene and mode. Keep the raw CSV, hardware/driver, resolution,
|
||
shader bundle, validation state and source revision with any published result. The
|
||
[P2 acceptance protocol](https://github.com/emil28092005/Faset_Engine/blob/main/docs/studies/19-p2-gpu-visibility-acceptance.md)
|
||
documents the scenes and CSV columns. The
|
||
[first measured report](https://github.com/emil28092005/Faset_Engine/blob/main/docs/studies/20-p2-gpu-visibility-benchmark-2026-09-23.md)
|
||
is a **pre-optimization baseline**: its Debug/validation profile found GPU MainCull
|
||
substantially more expensive than direct GPU work. The
|
||
[optimized follow-up](https://github.com/emil28092005/Faset_Engine/blob/main/docs/studies/21-p2-gpu-visibility-optimization-2026-09-23.md)
|
||
retains three additional raw runs and isolates the effects of device-local output
|
||
buffers and bounded atomic append. MainCull p50 fell to 0.030–0.042 ms in those
|
||
synthetic scenes. That comparison is useful for diagnosis, not a guarantee that
|
||
GPU visibility speeds up a particular game or device.
|
||
|
||
The harness enables GPU visibility counters, so diagnostic readback is part of its
|
||
timings. In the Editor, opening diagnostics likewise enables these counters, and
|
||
**Show HZB** adds an on-demand image copy and preview upload. Close the panel and
|
||
disable the HZB preview for ordinary gameplay timing. GPU pass timestamps separate
|
||
MainCull, MainRaster, HZB, PostCull and PostRaster when supported; they are not a
|
||
measure of CPU extraction/upload. The renderer still waits for frame completion
|
||
and reads back the full image, so `cpu_ms` is wall time including waits, not CPU
|
||
utilization. An open scene can run slower with HZB; visibility correctness and
|
||
full-frame speed are separate findings.
|
||
|
||
## Measure P3 lighting and shadows
|
||
|
||
A Player `--profile` sample includes `effective_lighting_path`, local lights
|
||
submitted/omitted, requested/effective sun cascades, requested/rasterized local
|
||
shadow faces, tile use, shadow drop reasons, caster draws, and explicit atlas
|
||
allocation bytes. `gpu_main_raster_ms`, `gpu_sun_shadow_ms`, and
|
||
`gpu_local_shadow_ms` are GPU timestamps or `null` when timestamps are
|
||
unavailable. A light can illuminate while its shadow faces are dropped. A
|
||
submitted-light count of zero is a different workload from 128 lights whose
|
||
shadows are disabled. See [Lighting](lighting.md) for the capacity policy and
|
||
[Diagnostics](diagnostics.md) for the Editor counters.
|
||
|
||
The same sample includes `effective_lighting_path` (`forward` or `tiled`),
|
||
`gpu_light_tiles_ms`, and `light_tile_count`. Stored candidate and overflow
|
||
counts are present only when visibility diagnostics readback was enabled;
|
||
`light_tile_counts_valid: false` means their `null` values are unavailable,
|
||
not zero. The normal `Auto` setting currently resolves to `forward` after the
|
||
fixed dense 1080p benchmark showed that tile construction cost outweighed its
|
||
raster savings. A C++ renderer integration can explicitly request `Tiled` for
|
||
a localized-light scene, then check the actual path before comparing timings.
|
||
|
||
The fixed-scene benchmark compares 0, 4, 16, 32, 64, and 128 local lights under
|
||
Direct, GPU frustum, and GPU occlusion visibility, with shadows on and off. Its
|
||
wrapper runs three independent 1920×1080 repetitions per configuration, each
|
||
with ten warm-up and thirty recorded frames. First inspect the planned matrix:
|
||
|
||
```sh
|
||
python3 tools/benchmark_p3_lighting.py --list-runs
|
||
```
|
||
|
||
From the repository, after a Linux Release renderer build, run one shadow setting
|
||
into a new output directory. Supply the actual device driver identity:
|
||
|
||
```sh
|
||
python3 tools/benchmark_p3_lighting.py --sweep \
|
||
--executable build/linux-release/faset_p3_lighting_benchmark \
|
||
--output .cache/p3-lighting-off \
|
||
--shadows off --driver 'REPLACE_WITH_ACTUAL_DRIVER' --validation off
|
||
```
|
||
|
||
The wrapper writes one raw CSV per run, `merged.csv`, and `summary.json`. Keep
|
||
all three with the exact source revision and device. It checks that every run
|
||
used its requested visibility mode and submitted every requested light. GPU
|
||
timestamps for the main raster isolate fragment-heavy lighting better than
|
||
renderer wall time, which includes GPU waits and synchronous readback. Shadow
|
||
time is split into sun and local GPU durations. The Forward+ decision compares
|
||
the median of three run medians against the matching zero-light configuration;
|
||
the threshold is **1.0 ms extra main raster time or 15% of the zero-light GPU
|
||
frame** at 32, 64, or 128 lights on the Linux physical reference GPU. The
|
||
[P3 lighting validation record](https://github.com/emil28092005/Faset_Engine/blob/main/docs/validation/p3-lighting-2026-09-24/README.md)
|
||
states the measured decision and scope. A software Vulkan run checks
|
||
functionality, not physical GPU performance.
|
||
|
||
For a direct comparison of the two algorithms on the same scene, invoke the
|
||
Release executable twice with `--lighting forward` and `--lighting tiled`,
|
||
using the same `--lights`, `--shadows`, `--visibility`, and output size. The
|
||
default `--light-layout dense` preserves the fixed benchmark scene;
|
||
`--light-layout localized` reduces point-light ranges to 1.75 units as a
|
||
separately labelled workload. Compare `gpu_build_plus_raster_ms`, which includes
|
||
`gpu_light_tiles_ms`, rather than raster time alone. One optional diagnostic
|
||
frame with `--tile-diagnostics on` reports candidate and overflow counts but
|
||
adds a GPU readback, so do not mix it into the timed runs. The
|
||
[Forward+ measurement](https://github.com/emil28092005/Faset_Engine/blob/main/docs/studies/23-p3-forward-plus-2026-09-24.md) retains
|
||
raw frames, shader hashes, and the decision.
|
||
|
||
## Current performance scope
|
||
|
||
The accepted MVP path uses direct draws and CPU culling; P2 adds optional GPU
|
||
visibility for opaque static meshes, with prepared LODs supplied by the project.
|
||
P3 adds local lights and bounded sun/local shadow atlases. The benchmark's
|
||
`lighting_path` and a Player profile's `effective_lighting_path` identify the
|
||
algorithm actually used. Both paths currently use one graphics queue and
|
||
synchronous full-image capture/readback. Use measurements to find the next
|
||
bottleneck before introducing parallel jobs or expanding GPU-driven rendering.
|
||
Neither an offscreen capture benchmark nor a tiny demo is a promise of a
|
||
production frame budget. Observed measurements and follow-up targets belong in
|
||
the implementation acceptance report with their source revision and method.
|