11 KiB
Profiling and measurements
Use a Release export to measure the shipping Player. Record the exact scene, hardware, driver, build configuration and resolution with the result. Small test scenes do not establish performance for a large game.
Capture a bounded Player profile
From a standalone generation directory:
./faset_player --headless --frames 240 --profile profile.json
--headless here means offscreen Vulkan rendering. A GPU/driver is still required.
The Editor's headless authoring mode is a separate feature. Omit this flag to measure
the windowed path. --profile requires an explicit --frames between 1 and 100000,
which bounds the stored samples.
The JSON contains raw completed-frame samples and nearest-rank p50/p95 summaries. No warm-up frames are silently removed. It records the presentation mode, device, resolution, validation activation, fixed ticks and timestep. A bounded run advances one synthetic fixed timestep per frame; it does not reproduce a real-time input session. Keep that distinction when comparing runs.
Startup starts at the Player application entry after platform argument normalization
and ends at the first completed frame. OS process loading and Windows wmain UTF-8
argument conversion are excluded. Frame wall times exclude writing the final profile and
capture files. Simulation and scene-snapshot times are separate from the renderer
call. Renderer CPU wall duration includes GPU waits and readback; it is not CPU
utilization. GPU timestamps measure the submitted graphics work and can be null
when timestamps are unsupported.
Resource counters report live renderer allocations and texture count. GPU allocation bytes include Vulkan allocation alignment and exclude driver-internal memory; they are not a whole-process VRAM meter. The fallback white texture is included.
Use --debug-physics or press F3 to show current physics box colliders. Debug
geometry increases draw count, so record whether it was enabled. The collider view
uses simulation poses; normal visuals can use interpolated poses.
Measure Editor and C++ workflows
From the engine repository:
python3 tools/measure_workflows.py \
--editor build/linux-debug/faset_editor \
--project examples/projects/collect-3d \
--output .cache/my-workflow-measurement
Use a new output directory. The tool copies the project, preserving your original, and records command startup, two-frame GUI startup/shutdown, first/cached Blender bundle import, initial/no-change/changed Debug builds and a subsequent Player frame. It checks that editing gameplay makes the schema stale and successful building clears that state. The initial build uses available dependency archives and OS caches; it is not a measurement of internet download speed.
On Linux, GNU time records peak RSS for each command and its waited-for children.
This is a maximum, not the sum of simultaneous compiler processes. Other platforms
report this field as null unless equivalent measurement support is added. The tool
keeps raw stdout/stderr, durations, hardware and revision information alongside its
report. A dirty source checkout is explicitly identified.
tools/verify_playable_exports.py separately verifies the two sample games in
relocated Release packages and records their Player profiles. Its assertions test
correct execution, not a frame-time threshold.
Compare P2 GPU visibility modes
The Editor diagnostics panel (F12) can switch its current viewport between Direct, GPU frustum, and GPU occlusion. Direct is the default reference. The selector is an Editor viewport setting; it does not change the saved scene or automatically change an exported Player. An exported Player can select a mode for a bounded run:
./faset_player --headless --frames 240 --profile gpu-frustum.json --visibility gpu-frustum
Accepted values are direct, gpu-frustum, and gpu-occlusion; Direct is the
default. The profile records the requested visibility_mode, the run's and each
frame's effective_visibility_mode, and each frame's gpu_visibility_active state.
Compare requested and effective modes before interpreting a GPU run: a missing GPU
profile falls back to Direct, while missing HZB can reduce GPU occlusion to GPU
frustum. The Editor shows the Effective path and any Fallback from line. See
Diagnostics for the counters and HZB preview.
For a repeatable offscreen comparison, build and run the P2 benchmark harness:
build/linux-debug/faset_render_gpu_acceptance_tests --benchmark /tmp/faset-p2.csv
It records 10 warm-up and 30 measured frames for Direct, GPU frustum and GPU occlusion in fixed frustum-heavy, open and occluded scenes. Run it three times and compare median/p95 by scene and mode. Keep the raw CSV, hardware/driver, resolution, shader bundle, validation state and source revision with any published result. The P2 acceptance protocol documents the scenes and CSV columns. The first measured report is a pre-optimization baseline: its Debug/validation profile found GPU MainCull substantially more expensive than direct GPU work. The optimized follow-up retains three additional raw runs and isolates the effects of device-local output buffers and bounded atomic append. MainCull p50 fell to 0.030–0.042 ms in those synthetic scenes. That comparison is useful for diagnosis, not a guarantee that GPU visibility speeds up a particular game or device.
The harness enables GPU visibility counters, so diagnostic readback is part of its
timings. In the Editor, opening diagnostics likewise enables these counters, and
Show HZB adds an on-demand image copy and preview upload. Close the panel and
disable the HZB preview for ordinary gameplay timing. GPU pass timestamps separate
MainCull, MainRaster, HZB, PostCull and PostRaster when supported; they are not a
measure of CPU extraction/upload. The renderer still waits for frame completion
and reads back the full image, so cpu_ms is wall time including waits, not CPU
utilization. An open scene can run slower with HZB; visibility correctness and
full-frame speed are separate findings.
Measure P3 lighting and shadows
A Player --profile sample includes effective_lighting_path, local lights
submitted/omitted, requested/effective sun cascades, requested/rasterized local
shadow faces, tile use, shadow drop reasons, caster draws, and explicit atlas
allocation bytes. gpu_main_raster_ms, gpu_sun_shadow_ms, and
gpu_local_shadow_ms are GPU timestamps or null when timestamps are
unavailable. A light can illuminate while its shadow faces are dropped. A
submitted-light count of zero is a different workload from 128 lights whose
shadows are disabled. See Lighting for the capacity policy and
Diagnostics for the Editor counters.
The same sample includes effective_lighting_path (forward or tiled),
gpu_light_tiles_ms, and light_tile_count. Stored candidate and overflow
counts are present only when visibility diagnostics readback was enabled;
light_tile_counts_valid: false means their null values are unavailable,
not zero. The normal Auto setting currently resolves to forward after the
fixed dense 1080p benchmark showed that tile construction cost outweighed its
raster savings. A C++ renderer integration can explicitly request Tiled for
a localized-light scene, then check the actual path before comparing timings.
The fixed-scene benchmark compares 0, 4, 16, 32, 64, and 128 local lights under Direct, GPU frustum, and GPU occlusion visibility, with shadows on and off. Its wrapper runs three independent 1920×1080 repetitions per configuration, each with ten warm-up and thirty recorded frames. First inspect the planned matrix:
python3 tools/benchmark_p3_lighting.py --list-runs
From the repository, after a Linux Release renderer build, run one shadow setting into a new output directory. Supply the actual device driver identity:
python3 tools/benchmark_p3_lighting.py --sweep \
--executable build/linux-release/faset_p3_lighting_benchmark \
--output .cache/p3-lighting-off \
--shadows off --driver 'REPLACE_WITH_ACTUAL_DRIVER' --validation off
The wrapper writes one raw CSV per run, merged.csv, and summary.json. Keep
all three with the exact source revision and device. It checks that every run
used its requested visibility mode and submitted every requested light. GPU
timestamps for the main raster isolate fragment-heavy lighting better than
renderer wall time, which includes GPU waits and synchronous readback. Shadow
time is split into sun and local GPU durations. The Forward+ decision compares
the median of three run medians against the matching zero-light configuration;
the threshold is 1.0 ms extra main raster time or 15% of the zero-light GPU
frame at 32, 64, or 128 lights on the Linux physical reference GPU. The
P3 lighting validation record
states the measured decision and scope. A software Vulkan run checks
functionality, not physical GPU performance.
For a direct comparison of the two algorithms on the same scene, invoke the
Release executable twice with --lighting forward and --lighting tiled,
using the same --lights, --shadows, --visibility, and output size. The
default --light-layout dense preserves the fixed benchmark scene;
--light-layout localized reduces point-light ranges to 1.75 units as a
separately labelled workload. Compare gpu_build_plus_raster_ms, which includes
gpu_light_tiles_ms, rather than raster time alone. One optional diagnostic
frame with --tile-diagnostics on reports candidate and overflow counts but
adds a GPU readback, so do not mix it into the timed runs. The
Forward+ measurement retains
raw frames, shader hashes, and the decision.
Current performance scope
The accepted MVP path uses direct draws and CPU culling; P2 adds optional GPU
visibility for opaque static meshes, with prepared LODs supplied by the project.
P3 adds local lights and bounded sun/local shadow atlases. The benchmark's
lighting_path and a Player profile's effective_lighting_path identify the
algorithm actually used. Both paths currently use one graphics queue and
synchronous full-image capture/readback. Use measurements to find the next
bottleneck before introducing parallel jobs or expanding GPU-driven rendering.
Neither an offscreen capture benchmark nor a tiny demo is a promise of a
production frame budget. Observed measurements and follow-up targets belong in
the implementation acceptance report with their source revision and method.