Skip to content

Benchmark repeatability check

On this page

In six runs at 1920×1080, the coefficient of variation of PC Score is 0.256%. Two runs at a lower resolution deviated from the mean of these six by +0.392% and −0.070%. On one PC the spread turned out to be small, but this data cannot be used to estimate noise on other configurations.

Here is the breakdown of all eight runs from 2026-09-04. Not a single run was dropped from the tables. The calculations were recomputed on 2026-09-07.

Parameter Series data
CPU AMD Ryzen 7 7800X3D, 8 physical cores / 16 logical processors
GPU NVIDIA GeForce RTX 5070
Windows Build 26200; version string in the results Windows 10.0.26200
Engine PCBenchmarkX 0.5.2
Profile Candidate, identical durations and test composition
Scoring model score-model-1.2-candidate-1
Internal resolution 1920×1080 in all eight runs
Output Full-screen borderless window
Environment state A detected hypervisor is noted in all results

According to the author, Windows with BoosterX optimizations was used. The full list of settings from the series has not been published. The archive also contains no comparison with a non-optimized Windows. The noted hypervisor does not prove that the run took place inside a guest VM.

The summary does not include the graphics driver version, and the archive does not allow determining component temperatures and the full list of settings during the tests.

The first three runs were performed within one Windows boot, the next three each after a separate reboot. This is visible from the system uptime: the calculated boot moment coincides for №1-3 and differs for №4, №5 and №6. The last two resolution tests were performed after №6 without a new reboot.

№ Windows boot System uptime before the test, min Output resolution PC Score
1 A 4.94 1920×1080 14 198.58
2 A 7.56 1920×1080 14 229.16
3 A 10.23 1920×1080 14 186.01
4 B 0.67 1920×1080 14 239.20
5 C 3.52 1920×1080 14 224.65
6 D 7.80 1920×1080 14 141.08
7 D 10.62 1280×768 14 258.79
8 D 13.48 800×600 14 193.24

Run №4 started less than a minute after boot. The wait before the tests differs, so it is impossible to isolate the “reboot-only effect” here: the behavior of the software after system startup is superimposed.

CV here is calculated between full runs, via the sample standard deviation with divisor n − 1:

mean = sum(x) / n
s = sqrt(sum((x − mean)^2) / (n − 1))
CV, % = 100 × s / mean
range, % = 100 × (max(x) − min(x)) / mean
PC Score series n Mean CV Range / mean
One boot, №1–3 3 14 204.58 0.156% 0.304%
Different boots, №4–6 3 14 201.65 0.373% 0.691%
All 1920×1080, №1–6 6 14 203.12 0.256% 0.691%

Comparing the first and second triple gives a difference of −0.021%. The spread between boots is higher than within one boot, but in this series it is still small. Three runs per group is not enough for a strict conclusion about the limiting error.

Component, №1–6 Mean CV between runs Range / mean
Performance 11 437.74 0.121% 0.330%
Core Latency Score 21 750.66 0.392% 1.129%
Consistency 12 878.65 0.832% 2.156%

Of the three components, Performance showed the greatest stability. The blocks that use the tail of the latency distribution fluctuated more noticeably. Therefore, final scores that look identical do not yet guarantee an identical picture in terms of frame distributions.

Resolution PC Score Δ PC to mean №1–6 Δ Performance Δ Core Latency Score Δ Consistency
1280×768, №7 14 258.79 +0.392% −0.054% +0.633% +1.149%
800×600, №8 14 193.24 −0.070% −0.089% −0.099% +0.022%

Lowering the output resolution did not lead to a systematic increase in Performance. This fits the logic of a fixed internal load of 1920×1080.

If we compare only with the nearest 1920×1080 run №6 in the same Windows boot, the PC Score changes are +0.832% and +0.369%. The difference depends on whether we compare with a single run or with the mean of several, so both variants are shown here.

At each low resolution there was only one run. The order was not alternated, and there was no repeat run of 1920×1080 after changing the resolution. The picture matches the hypothesis of a fixed internal load, but does not rule out the influence of resolution on individual components. Verification requires a repeat on several PCs with alternating order.

Verification of calculations and available data

Section titled “Verification of calculations and available data”

All eight runs completed: in the published CSV, all four scores and input aggregates are filled in for each run. Run suitability is assessed by the run quality model (Run Quality) described in the methodology: it takes into account background CPU load, hypervisor, battery power and a remote session. At the same time, external interference could still have affected the measurements: run suitability by this model does not prove its absence.

For all runs, Performance, Core Latency Score, Consistency and PC Score were recalculated from the saved throughput and percentiles. The maximum discrepancy with the publication is less than 0.000001 points. This is a verification of the formulas against aggregates; a full rebuild of percentiles from raw frames was not done in this analysis.

Save both files side by side and run:

Окно терминала
python analyze_repeatability.py

The CSV contains the scores and input aggregates, resolutions, system uptime and boot labels. The full source archive with identifiers and service data is not published.

To check repeatability on other computers, send your series to admin@boosterx.org with the subject “PCBenchmarkX: repeatability”.

Send several full results under identical conditions and specify CPU, GPU, RAM, Windows and driver version, profile, resolution, reboots and changes between tests. If you are checking the influence of resolution, alternate it between runs and do several runs at each value. It is better to send the whole series in full, including “bad” or rare cases.

Before sending, check the archive for personal data. State whether the results and computer specifications can be published without your personal data. After verification, we will be able to add the series here with conditions, calculations and limitations. Also send series with a large spread: they will help to understand under which conditions the results are less stable.

This breakdown was prepared by the BoosterX team based on the provided series. It is not an independent certification and does not prove a gain from BoosterX optimization. Independent reproductions will be marked separately.