Skip to content

Run comparison

On this page

Compare completed runs with the same workload version, scoring model and profile. The percentage change describes the difference between numbers, but by itself does not prove the effect of a tweak.

Before comparing numbers, check the compatibility of a pair of runs using the result files: each of them records the workload version — a permanent identifier of its composition — and the scoring model, as well as the statistics version. Compare these fields pairwise: if the workload or the scoring model do not match, the runs belong to different protocols and their scores are directly incomparable. How these identities are structured is described in the methodology.

  1. Record the CPU and GPU model, Windows and driver versions. Do not change the power mode or cooling between tests. Use the same set of background programs.
  2. First perform at least three tests without changing settings. This way you will see the normal spread of results. A conclusion about small differences may require more repetitions.
  3. Change one setting and repeat the tests. If a reboot is required, perform it before the series with the original settings as well. Before each run, wait the same amount of time.
  4. If possible, check the original settings, then the changed ones, and then the original ones again: A → B → A. Such a sequence helps separate the effect of the tweak from warm-up and background activity.
  5. Keep all results. Determine in advance which failures make a test unsuitable for comparison, for example an interrupted run. If you exclude a test, state the reason. A low score by itself is not a reason to delete it.

Keep the screen resolution and windowed or fullscreen mode the same. If you are testing the effect of resolution, repeat the test several times at each chosen resolution.

Δ score, % = 100 × (mean_after / mean_before − 1)
CV, % = 100 × sample_standard_deviation / mean

Compare the mean values and the spread between runs. Check not only the final score, but also individual scores, latencies in milliseconds and workload throughput. For latencies, a negative difference usually means an improvement; for scores, a positive difference means a score increase.

To check repeatability, count full runs. Thousands of frames within a single test do not replace repeated tests: all of them are obtained in one session and depend on shared conditions.

A small change relative to the normal spread requires additional repetitions. A universal threshold like “anything above 1% is significant” cannot be established from a single series. Even a statistically distinguishable difference may be too small for a practical effect.

Per-block diagnostics are already built into the results: for each workload, the spread of throughput across blocks, the “first-last” shift and the trend across steps are saved. This allows judging whether the difference between two runs exceeds the noise observed within them, even before new repetitions. The formulas by which blocks are aggregated into scores are disclosed in the score calculation.

In the series of eight runs, repetitions and resolutions were compared on an optimized system. It does not contain an unoptimized state and therefore does not prove a gain from BoosterX.

A synthetic test helps investigate changes in system behavior. To verify the benefit for a specific game, repeat the comparison in its reproducible game scenario.

Verified: 2026-09-20.