Screen Test Pro

A thousand runs of benchmark 2.5: what the hardware actually scores

by Screen Test Propublished

Benchmark 2.5.0 went live on September 20. In the two weeks since, 1,065 people ran it and let the result be counted. That is the first time one version of this benchmark has stayed still long enough to collect four figures of runs, so this is the first time we can say what scores look like out there, and not only on the machines on our own desks.

Everything below comes from those runs, from September 20 to October 3. The charts are live: change the slice, change the grouping, switch any of them to a table. If you have run the benchmark in this browser, the first chart already knows your score.

Half of all machines score under 3,334

Of the 1,065 runs, 665 finished clean, and those are the ones the score charts use. Their median is 3,334. A tenth score under 636, a tenth over 13,583, and the best run of the fortnight was 31,468.

Complete runs by score, on a log scale: each bar is about a third wider in score than the one before it. Type a score to see where it lands in the slice you picked.

The tier names split that population more evenly than we expected when we cut them: 19% Entry, 28% Balanced, 32% Fast, 22% Extreme. Past 20,000 it gets lonely. Only 21 runs made it, about 3%.

The shape matters more than the median. There is a hump between 400 and 1,300, a dip just above it, and then a long plateau from about 1,800 to past 13,000. That is not one population of computers. It is two: machines running on integrated graphics, and everything else.

The GPU decides almost everything

Group the runs by GPU family and the two populations come apart.

The tick is the median, the bar is the middle half of runs, the thin line runs from the 10th to the 90th percentile. Groups under 15 runs are left out.
GPU family Complete runs Median Middle half
GeForce 207 9,204 6,116 – 13,156
Radeon 81 3,270 1,194 – 12,841
Apple silicon 134 3,128 2,132 – 4,257
Adreno 42 2,564 1,322 – 4,537
Intel integrated 130 973 632 – 1,796
Mali 31 785 533 – 1,307

A typical GeForce machine scores about nine and a half times a typical Intel integrated one. Nobody will be surprised by the direction. The size is worth knowing if you are deciding whether a laptop needs the discrete card.

Radeon is the odd row. Its middle half runs from 1,194 to 12,841, a tenfold spread, because the name covers both the graphics built into a Ryzen laptop chip and a desktop card that trades blows with GeForce. A Radeon median describes neither. Narrow the chart to Radeon and group by memory and the halves separate: 12,614 for machines with 17 GB or more, 2,496 for those with 9 to 16.

Apple silicon is the opposite: tight. The middle half of Apple runs sits between 2,132 and 4,257, phones, tablets and Macs together. Narrow to Apple and group by operating system and the order is what you would guess: Mac 3,961, iPad 3,451, iPhone 2,412. An iPhone at the median outscores the median Android phone by half again, and beats the median Intel laptop two and a half times over.

No machine wins every scene

The total hides the most interesting part. The benchmark is six scenes, each leaning on a different part of the machine, and the families do not keep their order from one scene to the next.

Each cell is that group's median in one scene, as a multiple of the median across all runs. Blue is above the field, amber below.

GeForce owns the scenes that are pure GPU. In Nebula Drift, a million particles updated on the graphics card, its median is five times the field and seven times Apple’s.

Then look at Starfall Cascade. That scene is physics computed on the CPU, in JavaScript, and Apple takes it: a median of 12,196 against GeForce’s 9,394. Split by operating system and it is starker. The median Mac scores 17,390 in Starfall and the median Windows PC 7,729. The median iPhone, at 11,593, beats the median Windows PC at CPU physics.

Canvas 2D goes the same way. Macs post a median of 10,085 on the plain circles scene and GeForce machines 4,448. Group by browser and Safari’s median there is two and a half times Chrome’s. That is not the chip alone. It is the chip plus a browser whose 2D canvas is very fast, which is what we found on one Mac with five browsers last month and is now visible across a few hundred strangers’ machines.

So a single score is a fair summary and a poor description. A Mac and a mid-range gaming PC can land on the same total having agreed on nothing along the way.

What a battery really costs

Take every desktop and laptop run and split it by power source. Plugged in, the median is 4,697. On battery, 1,700. That looks like unplugging costs you almost two thirds of your performance.

It does not. Narrow the chart above to GeForce and group by power source: on battery the median is 11,830, plugged in 9,207. Seventeen battery runs is a thin sample, so we would not claim battery is faster, only that there is no penalty to see. Intel integrated says the same with more runs behind it: 1,061 on battery, 964 plugged in.

The big gap is about who is unplugged. A machine running on battery is a laptop, and most laptops in this sample have integrated graphics. Compare like with like and the gap closes. If your own laptop scores far lower unplugged, that is its power plan, and it is worth a look, but it is not the general rule the headline number suggests.

The same trap is in the refresh-rate split. Runs on displays of 165 Hz and up have a median of 10,399, and runs on 60 Hz displays 1,971. A fast panel does not make the GPU faster. People who buy a 165 Hz monitor also buy the card to drive it. Within GeForce alone the gap shrinks to 10,506 against 6,609, and what is left is mostly newer cards sitting behind newer monitors.

One more of these, since it comes up. Among GeForce runs, Chrome’s median is 10,583 and Edge’s is 8,683. Both are the same engine. We read that as who uses which browser, and nothing more.

Which machines finish clean

A run is complete when every scene found a load the machine could hold and prove. It is throttled when the machine measured fine but slowed down as it warmed. It is degraded when at least one scene could not be measured properly. Across all 1,065 runs: 62% complete, 25% throttled, 13% degraded.

All submitted runs, not only the complete ones. The figure on the right is the share that finished complete.

GeForce machines almost never fail to measure: one degraded run out of 269. iPhones are the steadiest group of all, with 78% complete and a single degraded run in 64. Mali phones have the hardest time, with 39% degraded, and tablets as a class finish complete less than half the time. Thin, fanless and warm is a difficult place to hold a steady load for three minutes.

A quarter of runs throttled is more than we would like. Some of that is real heat, and reporting it is the point of the flag. Some of it is still the benchmark being too easily unsettled, and that is the next thing we are working on.

The screens

Sixty percent of runs came from 60 Hz displays. The next biggest group is not 120 or 144 Hz but 165 Hz and up, at 21%, more than 120 and 144 put together. The single most common screen is still 1920 × 1080 at 100% scaling, on 18% of runs.

How to read all this

These are runs from people who found a screen-testing site and chose to press the button. That is not a survey of the world’s computers. Gaming PCs are over-represented, and so are people who just bought something and want to see what it does.

The scores come from complete runs only, 665 of them. A group needs 15 runs before it appears in any chart, which is why Linux, 75 Hz displays and several phone families are missing rather than shown with a number we could not stand behind. Where a group is small, its run count is printed next to it. Treat a median from 20 runs as a hint.

What we store about a run is short: the score, the six scene results, the screen’s numbers, and coarse families for the GPU, operating system and browser. No IP address, no country, no identifier, nothing that says who ran it. The privacy page has the full list, and nothing on this page is finer-grained than that list.

Scores only compare within one benchmark version, so all of this describes 2.5.0. If you want your machine in the next round of numbers, run the benchmark. It takes about three minutes, and how the score is built is written down on the methodology page.

Share this page
DiscussionOne thread per language, shared across the whole site. Sign in with Google to post.
Discussion