Xilly
PC Performance Testing Methodology
A before-and-after result is only useful when the workload, graphics settings, resolution, system state, and measurement method are comparable.
Results FPS optimization service Windows and DPC latency optimization Xilly optimization team
Who this page is for
For readers evaluating Xilly results and anyone who wants a repeatable way to test whether a PC change actually improved gaming performance.
- The methodology prioritizes repeatable workloads, multiple passes, frame-time behavior, thermals, and stability rather than one unusually high FPS screenshot.
Establish the baseline first
Record the hardware, Windows build, BIOS version, driver version, resolution, graphics preset, frame cap, upscaler, and test location before changing settings. If one of those variables changes, the result should be labeled instead of treated as a direct comparison.
- Same game version and repeatable scene where possible
- Same resolution, render scale, and graphics settings
- Same frame cap and latency-feature state
- Background downloads and avoidable overlays controlled
Measure more than average FPS
Average FPS describes throughput but can hide disruptive spikes. Where the game or capture tool permits, Xilly also considers 1% lows, frame-time plots, GPU utilization, CPU limits, clocks, temperatures, and power behavior.
- Average FPS for overall throughput
- 1% lows and frame times for consistency
- CPU and GPU utilization for bottleneck context
- Temperature and clock behavior for throttling
Repeat and reject bad comparisons
Run-to-run variance is normal. Repeat comparable passes, investigate large outliers, and avoid choosing only the best tuned run or the worst stock run. Online matches can provide useful real-world evidence, but they should not be presented as laboratory-identical workloads.
- Use multiple passes when the benchmark supports them
- Repeat unexpected outliers
- Separate synthetic benchmarks from real gameplay
- Label changes in maps, patches, or render settings
Performance is not stable until it survives use
A completed benchmark does not prove that an overclock or undervolt is stable. CPU, RAM, GPU, and curve changes should be tested with appropriate stress tools and then checked in multiple real games because different workloads expose different failures.
- Watch for crashes, driver resets, visual errors, and WHEA reports
- Test cold boots and restarts where memory training matters
- Check light-load behavior after aggressive CPU undervolting
- Use several games before calling a tune daily-stable
How Xilly describes results
Results belong to the tested PC and workload. They should not be converted into a guaranteed percentage for every customer. When original data is unavailable, Xilly distinguishes technical guidance from a measured Xilly benchmark.
- No universal FPS guarantee
- System specifications kept with the result
- Measured values separated from estimates
- Limitations stated when testing is incomplete
Primary measurement references
- Intel PresentMon - Official frame-time, percentile, GPU telemetry, and GPU Busy measurement documentation.
- NVIDIA FrameView - Official explanation of rendered FPS, displayed FPS, percentile performance, power, and render-present latency.
- UL 3DMark Steel Nomad - Official workload scope and CPU-limit caveats for the Steel Nomad graphics test.
- AMD Ryzen Master - Official Ryzen telemetry and tuning utility, supported-platform information, and overclocking cautions.
- Microsoft Windows Performance Toolkit - Official Windows Performance Recorder and Analyzer documentation for collecting and investigating system-level traces.
Frequently asked questions
Why are 1% lows important?
They summarize the slower frames that an average can hide. They are useful for consistency, but they still need a repeatable workload and should be read with the frame-time plot.
Does one benchmark prove an overclock is stable?
No. A benchmark is one workload. A daily-stable tune should also survive appropriate stress testing, restarts, light-load behavior, and several real games.
Can two online matches be compared perfectly?
Usually not. Player count, map activity, shaders, network state, and background work can differ. Online gameplay is useful evidence, but controlled built-in benchmarks are better for direct numerical comparisons.