Changelog
Newest first. Anything that changes a score gets a benchmark version bump and says so here. RSS
2.14.0
Five shaders join the gallery, and two of them are test patterns2026-09-10
- The effects gallery has a third tier. Five fullscreen fragment shaders now sit under the thirteen benchmark scenes on the effects page: Star Nest, Marble Drift, Synthwave Horizon, Neonwave Sunrise and Static Field. They are other people’s work, shipped under MIT or CC0, credited on the page and again in the fullscreen HUD — Shadertoy’s default terms are non-commercial, so a shader without an explicit permissive licence does not get added here however good it looks. Nothing about the benchmark moved: its workload has been frozen since 2.2.0 and no shader scene is ever scored. The homepage gallery still shows the thirteen scenes a score is made of.
- Two of them are diagnostics, not decoration. Static Field puts an uncorrelated random value in every pixel and replaces all of them every frame, which is the hardest thing a panel or a video pipeline can be handed. Run it fullscreen and slow pixel response shows as smearing, while a cast or a screen share falls apart outright. Synthwave Horizon at a high supersample count is a sustained GPU load, useful for hearing when the fans spin up or watching a laptop throttle on the live fps readout.
- The load knob means something else on a shader. A fragment shader has no particles, so there is nothing to count; it sets supersamples per pixel instead, 1 to 64. Sixteen samples is sixteen times the fill rate, and looks it. Speed and render scale work as they always have. Shader tiles also preview for real on hover instead of showing a painted approximation — there is no library behind the tier, so a shader compiles in the time three.js would still be downloading.
2.13.0
A score history you can read, and articles that fill their page2026-09-10
- The score history table stopped cutting your score in half. The popup took its column widths from the spec tables elsewhere on the site, which gave the timestamp nearly half the dialog and left the score free to break mid-number — 11,850 arrived as two lines. Each column now sizes itself: the date shrinks to what it needs, the run column takes the rest, and the score stays whole. On a phone the timestamp stacks the clock under the date instead of splitting the date at its hyphens, and nothing runs off the side.
- Articles and blog posts are the width of the text in them. The reading column has always been capped at about seventy characters; the page around it was 1,200 pixels, so some 500 pixels sat blank down the right of every post, with the byline rule and the share buttons running that far past the text they belong to. Both now end where the text ends. The words are untouched — same measure, same line breaks.
2.12.0
Controls that get out of the way, and a homepage every language can have2026-09-03
- The touch test’s controls move out from under your finger. The coverage strip and the close button now sit together on whichever edge you are not painting near — paint the bottom and they go up, paint the top and they come down — so every cell on the panel can be reached. The one-line instruction leaves at the first touch and comes back on reset, and on a phone the Reset button no longer names a key you don’t have. iPhone Safari gives pages no fullscreen API, so the browser’s own bars stay; the tool’s don’t.
- A play triangle that renders everywhere. Every start button ends in ▶ instead of ►, which drew as a squat, mismatched glyph on phones. Sixteen editions, the help cards, the burn-in resume button and the localized landers changed together.
- Less text above the first chip. The homepage headline is now “Screen test & benchmark”, the line under it is gone, and on narrow screens the panel hint drops the keyboard instructions — “Esc exits” means nothing on a phone. The topbar reads “Screen Test Pro” in every language; the domain form has left the share card and copied results as well.
- The one-click catalogue is ready to translate. Every chip name, group note, HUD line and readout label the English homepage shows now has a key in all fifteen dictionaries, and a language switches to the catalogue the moment its keys carry translations. Until then a translated homepage looks exactly as it did yesterday.
2.11.0
The full leaderboard, and the pattern tool learns to cycle2026-09-03
- A leaderboard page. The homepage board stops at ten. /leaderboard/ keeps going: up to 200 rows, one per machine, for the last 30 days or all time, and for every benchmark version we still hold samples of — pick 2.1.0 or 2.0.0 and see what the board looked like under that workload. Same coarse fields as the homepage rows, same one-row-per-machine rule; only the depth and the history are new. The privacy page says so in as many words.
- One cycle for all sixteen patterns. Yesterday’s pattern chips opened two different tools — the three broadcast cards in one, the thirteen drawn patterns in another — so “Test patterns ▶” and the arrow keys only walked part of the set. Now every chip and the group button open the same tool and → steps through all sixteen, cards included. The three-card tool remains on the TV lander and the localized pages.
- UFOs. The motion test’s moving object is now a UFO, as the genre expects, with six graphics to choose from in the bottom bar — and the striped block is still there for anyone who prefers edges to saucers. Icons from Flaticon, credited on the ghosting page.
- Smooth scrolling actually scrolls. The scrolling-text mode shipped with a wrap bug that bounced the text between the top and bottom of the screen instead of running through it. Fixed: one continuous ring of lines, bottom to top, no seam.
2.10.0
Every field, one click from the homepage2026-09-03
- Fifty-five checks on the front page. The homepage used to offer three tools and hide the fields inside them. Now the console shows the fields themselves: 16 solid colors, 13 gradients, 16 patterns, 4 motion tests and 6 measured instruments, each a chip that goes fullscreen in one click. The chips are drawn in the colors they’ll fill the screen with, so you can pick the field you want instead of stepping to it.
- More fields worth having. Solids gained orange, purple, four more grays and a custom hex picker. Gradients gained an RGB split, a hue sweep, the shadow and highlight quarters on their own, gamma steps and a 16-step staircase. A new pattern tool draws a fine grid, checkerboard, text at eleven sizes, one-pixel lines in three directions, dots, a one-pixel checker for native-resolution checks, near-black and near-white contrast steps, saturation, an overscan frame and a crack-tracing grid — all at device pixels, never scaled by the browser.
- Motion has four modes. The UFO lanes stay as they were. Joining them: a block chase that jumps one slot per frame (a dropped frame shows as a skip), a flashing block for flicker and refresh mismatches, and scrolling text for judder. Keys 1–4 switch modes inside the test.
- The board stays where it was, with a foot. The ten best runs and your own lane keep their place in the benchmark card. Underneath, the same hardware groups /benchmarks/ publishes show their medians once a group passes thirty samples — until then the line says how many samples exist, not a made-up number.
- The device readout fits again. With the board in the card, the fourteen readout rows now sit in two columns so a 1440×900 desktop sees the whole console without scrolling.
- English homepage only, for now. Locale homepages keep their three-row console until their strings get a proper human pass; every new tool still opens from the translated landers.
2.9.0
Throttled is a flag, not a disqualification2026-09-02
- Runs that throttle now count. Until today the top board only published runs where the device never slowed — and a phone a hundred seconds into a full-load GPU run always slows, because that’s what phones do. The result: no phone had ever reached the board on benchmark 2.2.0, and sixteen of its first twenty runs were silently invisible. Now a run where every scene was fully measured but the device warmed up gets a new status — throttled — and ranks on the board with that word right on the row. Nothing about the score changed: the same run produces the same number as yesterday, so the board keeps its history and benchmark 2.2.0 keeps its version.
- Degraded still means degraded. A run where a scene couldn’t be measured or its level couldn’t be confirmed stays off the ranking — that score is a floor, and ranking floors against peaks would mislead everyone. The results sheet now says this in plain words instead of leaving you to wonder whether the tool ate your score, and each scene’s flags say exactly what happened. Your own board lane treats throttled like complete, marks the flag at the front of the row (where a narrow phone screen can’t truncate it away), and degraded runs still hold the seat with an honest label.
- Advice before, diagnosis after. The start card now says the quiet part: a warm device throttles and scores lower, so give it a minute if you’ve just been pushing it — and if it throttles anyway, the run is flagged, not discarded. After a throttled run, the sheet suggests the obvious experiment: cool, plugged in, run again.
- The sample got smarter about heat. Shared samples (schema v5) now record why a scene was flagged and how much the device had slowed by scene’s end — a single ratio like 1.3. That’s the field data a future cool-down-aware benchmark revision needs, gathered the honest way. The privacy policy lists the new fields; as always, they describe the run, not you.
- Fresher board. A submitted score now shows up on the board without a five-minute wait — the page refreshes it after your run, and other devices see it within a minute.
2.8.0
A fair referee and a taller ladder2026-09-01
- Benchmark 2.2.0. Scores reset — here’s why it was worth it. A field trip put three Apple flagships side by side: an M2 Max laptop, an M4 Max Mac Studio, an M5 Pro laptop. The benchmark scored them within 9% of each other, strongest GPU underneath. Wrong answer, and traceable to three separate causes — a pass rule that favored laptop displays, one scene measuring the wrong thing, and a score ceiling real hardware had outgrown. All three are fixed below. Fixing them changes every number, so 2.2.0 starts a fresh board; how the score works has the full method, updated to match.
- The pass judge now understands displays. A fixed 60 Hz monitor can present a frame after 16.7 ms or 33.3 ms — nothing in between — while an adaptive laptop panel hands an 18 ms frame back as exactly that. The old tolerances quietly let adaptive panels cruise at 55 fps while desk monitors had to hold a true 60: same GPU, lower score at a desk. A ~600 ms pacing probe now reads how your display returns frames before scene one, and the pass rule — at most 5% of frames over the pacing-aware threshold, trimmed mean inside the 16.67 ms budget — makes “sustains 60 fps” the same claim on every panel. And if your display is stuck below 50 Hz (Low Power Mode, a 30 Hz cable), the run declines and says why, rather than grading a GPU through a straw.
- Liquid Chrome was measuring memory allocation, not raymarching. Changing its load resizes the render target, and the browser bills that as one long frame — which landed in the next measurement window and failed levels the GPU holds easily. An M4 Max “measured” at half an M2 Max because of it. Frames right after a load change are now discarded: the transition is not the workload. On our bench machine the fixed search settles 45% higher, at the level the silicon actually sustains.
- The score can’t hit its ceiling before the measurement does. The old flat 20× clamp froze two entire scenes at identical 20,000s on every fast Apple GPU — 38% of the score, constant across a chip generation. Ladder ceilings now sit 6–14× past the fastest machine measured to date (60.6 million nebula points, 32 million city instances, 180 raymarch samples per pixel), they cost nothing until a device earns them, and the score clamps exactly where the ladder ends, nowhere sooner.
- References re-centred on real, uncensored hardware. Scenes are normalized so the anchor machine — an M2 Max on a plugged-in bench — scores 13,000 with all six scenes agreeing within a few percent, instead of one machine reading 0.11× on one scene and 4.42× on another. Tiers are re-cut to match the new scale, validated by re-scoring every production sample: Entry under 500, Balanced to 1,499, Fast to 5,999, Extreme from 6,000.
- Shared samples now carry the raw numbers. Each scene’s actual maximum stable load — the figure the score clamp used to erase — plus a coarse display-pacing class (fixed / adaptive / unknown) now ride along with a shared result. That is what the next calibration will read, so it won’t depend on photographing a demo unit in an Apple Store. Nothing identifying was added; the privacy policy lists every field, as always.
2.7.1
One board seat per machine2026-09-01
- The top-scores board now seats one row per machine profile. The first month of benchmark 2.1.0 made the problem obvious: ten slots, and most of them read “GeForce · Windows · Edge” — several provably the same one or two computers, run again and again. Scores still carry no identifier of any kind, so the board now groups runs by the things a machine can’t change between attempts — GPU, OS, browser and version, CPU cores, memory, device class — and shows only the best run from each group. Running again can improve the row you already hold; it can’t buy a second one. Nothing new is collected, and the privacy policy didn’t have to grow to allow this — it got stricter.
- Board rows now name the browser version. “Edge 152” instead of “Edge”, on the global rows and on your own lane. The version was always part of a sample; showing it makes two different machines look like two different machines.
2.7.0
Gallery soundtracks, share buttons, fairer tiers2026-08-30
- The effects gallery plays music now. Thirty original tracks across the 13 scenes, each written for what’s on screen. Open a scene and tap ♪ — the choice and the volume are remembered on your device. Nothing loads until you turn it on, and the performance test stays silent so scores keep meaning the same thing everywhere. There’s more in the blog post.
- The tier labels got re-cut. The old cutoffs were set when scores stopped at about 4,300. Benchmark 2.1.0 raised the ceiling to 20,000, and half of all runs were suddenly landing in “Extreme” — a label that no longer meant anything. The new bands are simple multiples of the reference machine: below 1× is Entry, up to 3× is Balanced, up to 9× is Fast, past 9× is Extreme. Your score didn’t change; the word next to it might have. Details in how the score works.
- Every label on the result card is now a button. Click Extreme, degraded, drift, capped or failed and a short plain-language note appears right under it saying what that word means for your run — including the one people misread most: capped means the search ran out of time, so the true number is that or higher.
- Pages can be shared. A row of share buttons sits above the footer on tools, articles and posts — the big networks, email, a QR code for moving a link to your phone, and plain copy-link. The buttons load only when you scroll near them, so pages are exactly as fast as before.
- Language links moved to the bottom of the footer. Switching languages is a rare thing to do; the tests and pages you actually visit now come first.
2.6.0
The benchmark stops running out of ladder2026-08-29
- Benchmark 2.1.0: scores spread out again. Five of the six Standard scenes had been hitting the top of their own load ladder and reporting that limit instead of your hardware — Starfall Cascade in 98% of complete runs. With only one scene left free to move, unrelated machines landed on identical totals: an Android phone, a Windows desktop and a Mac all scored exactly 3,977. Every scene now has sixteen times its reference load to climb, and the search that finds the stopping point resolves to about 6% instead of 15%. Reference loads and weights are untouched, so 1,000 still means what it meant. Scores from 2.0.0 and 2.1.0 never compare, so the boards start empty and refill over the next few days.
- Two scenes stopped hoarding memory. Nebula Drift and Hypernova used to allocate their maximum particle count up front on every device, however few they ended up drawing — 58 MB between them, on phones and tablets that never got near it. They now repeat a smaller buffer instead, which costs half the memory and reaches a far higher ceiling.
- A run always ends with a result. On hardware that can’t hold the frame budget even at a scene’s lowest setting, the benchmark used to record nothing, close itself and drop you back on the page with no card and no explanation. It now shows the score it actually measured, says why it is what it is, and no longer stops the last scene from running at all.
- Slow devices are measured, not abandoned. Below about four frames per second every frame looked like a one-off stutter, so the engine threw the measurement away and sat at its starting load until the scene timed out. It now recognises sustained slowness for what it is and steps down to find a real answer.
- Your own run always has a seat. The you lane on the top-scores board showed the “run the benchmark” invitation even to people who had just run it, if the run came back degraded. It now shows your best result whatever its state, labelled honestly, at the bottom of the board.
- Same GPU, same pixels. The internal render size is finally capped the way the method page always said it was. A 4K display was rendering four times the fragments of a 1080p one at the same nominal load; those two scores were never comparable and now they are.
2.5.0
A top-scores board and a gallery with knobs2026-08-28
- Top scores. The ten best complete runs of the last 30 days, on the performance test page and the homepage bench card. Individual runs this time, not medians — hardware benchmarks still answers “what does this hardware typically score”; this board answers “what did the fast ones do lately.” Rows stay exactly as coarse as the consent card’s field list, and the privacy policy now says so in writing. The board also carries a you lane: your own best run, threaded into the ranking straight from this device’s local history — or a held seat until you take the test. Nothing about that lane is uploaded.
- Effects gallery is a page. All 13 scenes, each with three knobs the benchmark never offers: particle load on the scene’s own ladder, simulation speed down to 0.1×, render scale up to 2× native for a fill-rate stress. Free play stays unscored, and three.js still loads only when a Signature scene opens.
- Benchmarks page, drivable. A run button on the stats page itself — see a table, test your device against it on the spot. The tables also load into reserved space now instead of shoving the page around.
- Polish. The monitor test index got room to breathe; the homepage bench card centers on its run button, which now breathes a slow standby pulse (and holds still under reduced motion); the scoring link became an ⓘ that explains in place; the benchmark start card’s preview sits a little less flush.
2.4.0
Seven new pages and a real menu2026-08-26
- The test portfolio doubles. Touch screen test paints a multi-touch grid to expose dead zones and ghost touches. Burn-in test cycles uniform fields on a dwell slider — slow to inspect, fast to exercise retention. OLED screen test runs eight slides tuned to how OLED actually fails, true black through near-black to per-channel wear. Gradient / banding gives the existing ramps their own page.
- Device landers. TV screen test routes everything to a TV — built-in browser, cast tab or HDMI — with the broadcast patterns explained. Phone screen test packages the five-minute used-phone check.
- One map. Monitor test indexes all twelve instruments in the order a full check actually runs.
- A site menu. The topbar grew a menu button: every test and page, three dense columns, and mobile finally has navigation at all. The footer now links every lander from every English page.
- Classic mode, reachable. The eight v1 effects were fully wired as a scored Classic benchmark — engine, profile, translations — but nothing launched it. The performance test page now has the button. Classic runs score locally and never submit.
2.3.0
Four new instruments2026-08-26
- Every test gets its own page. Refresh rate measures the Hz your screen actually delivers, with a live frame-time strip. Resolution checker answers the reported-vs-physical-pixels question and computes PPI. Ghosting runs striped blocks at three speeds over five backgrounds. Monitor calibration walks black level, white level, gamma and gray balance in four fullscreen steps — and says plainly where only a colorimeter will do.
- Dead pixel test, deepened. A real FAQ (how many dead pixels are acceptable, do they spread) and links onward to the new tests.
- Two new guides. Calibrating by eye — and when you need a colorimeter and ghosting, response time and motion blur explained.
- Words un-glued. An Astro whitespace quirk was silently joining words across line breaks around links (“collected the policies inDead vs stuck…”). Fixed everywhere, with a regression test that scans every page.
2.2.1
The whole instrument, in 14 languages2026-08-26
- The benchmark speaks your language. Start card, HUD, results sheet, score history, and the tool help cards — 87 strings, human-translated into all 13 non-English locales. Scene names were already localized; now everything around them is too.
- Pattern notes localized. Each broadcast pattern’s “what to look for” line now shows in the page language.
- Related tools link home. The footer’s sister-site links now point at the matching language version of each site instead of English.
- Lighter locale pages. Locale homepages embed only the strings the tools actually use, not the full legacy payload.
- HUD progress fix. The progress bar no longer stalls during the bracket-search phase.
2.2.0
A lighter start card, livelier gallery, less data2026-08-23
- Less collected, not more. Samples no longer record country or site language — the field list shrank (schema v3), and the privacy policy shrank with it.
- The start card shows what you’re waiting for. A sample result preview — score, tier, per-scene bars — sits at the top of the benchmark card, and the technical notes folded into a “More information” toggle. One glance tells you what a run buys.
- Signature scenes animate on hover. All 13 gallery tiles now preview in motion — the five Signature tiles animate their Canvas 2D posters, so three.js still never loads until you actually open a scene.
- New header mark and a quieter topbar: the display readout lives only in the “This display” panel now.
- FAQ on the homepage. Eight short answers — dead pixels, banding, what the score means, what gets uploaded — each linking to the deeper read.
2.1.0
Hardware stats go live2026-08-23
- Hardware benchmarks page. Median scores by GPU family and Android device family, built from anonymous samples. Every row needs at least 30 samples before it publishes; medians, so one hot run can’t buy a rank.
- Sharing is on by default. The benchmark start card now carries a checked box instead of two buttons — with a live preview of the exact values your device would send. Uncheck it and nothing submits; the choice sticks. The privacy policy was updated in the same breath.
- Three new sample fields (schema v2): battery vs plugged in, form factor, and — Android only — a coarse device family like “Galaxy S24”. Rare models are floored to brand buckets at the point of collection; exact model strings never leave the device.
- A benchmark you can time. The HUD now shows scene 2/6, an overall progress bar, and a rough time-remaining estimate, with a scene-name interstitial between scenes.
- Start flows like the other tools. Confirming the start card drops straight into fullscreen and the first scene.
- The Signature scenes speak 14 languages. Nebula Drift, Neon Metropolis, Liquid Chrome, Starfall Cascade and Hypernova now have human-translated names and descriptions in every locale, and the gallery shows all 13 scenes everywhere.
- Terser homepage: less copy, bigger Run button.
2.0.0
The rebuild2026-08-23Screen Test Pro rebuilt from the ground up. Same URLs, same fourteen languages, new everything else.
- New benchmark. Five Signature scenes on WebGL2 (Nebula Drift, Neon Metropolis, Liquid Chrome, Starfall Cascade, Hypernova) plus the eight Classic effects from v1, ported faithfully and scored separately. Scores are versioned — this is benchmark
2.0.0 — and the methodology is public.
- Real measurement. Frame-time percentiles over statistics windows, an adaptive load ladder with hysteresis, thermal-drift detection, and hard timeouts on every state. The benchmark can be cancelled; it cannot hang.
- Anonymous hardware stats, consent-first. Finishing a scored run can submit one anonymous sample. The exact field list is shown before the run starts, skipping is one click, and the privacy policy was rewritten to match.
- New tool pages. Dead pixel test and performance test got proper homes with real guidance.
- New design. Dense, dark-first, instrument-style. The page shows your display’s live readout instead of a marketing hero.
- Faster. Static HTML, no framework runtime, no web fonts, scene code loads only when you run something.