PeekBench v1 · 77 exam tasks

Which AI gets your exam right — fastest?

We put AI models through the same screenshots and quiz pages Peeksolve solves every day. Every answer graded, every millisecond counted. Then run it yourself.

Tasks77 Subjects7 Models– Requestsreal Peeksolve
Top of the board
…Loading the leaderboard
PeekScore = 70 % accuracy + 30 % speedHow it works →
Leaderboard

Fast and right.

Official runs on PeekBench v1. Click a column to sort. Models marked “In Peeksolve” are the ones you can pick in the extension.

Benchmark your own →
#ModelPeekScore ↓ AccuracyFirst wordAuto-fill pageCost / 100 questionsBest at
Loading…
Speed race

Press ⌘⇧L. Who answers first?

Each car arrives at the moment its model's first word arrived, in real time — the median over 16 exam screenshots.

0.00 sfinish = first word on screen
Speed × accuracy · subjects

Where each model shines.

Speed vs. accuracy

Top left is where you want to be: right, and there in under a second. Point at a model to find it.

Right answers by subject

Share of tasks answered correctly in each subject, in percent.

What we test

Real exam work. Two ways in.

Every task is a screenshot or a quiz page exactly as the Peeksolve extension sends it — no clean API prompts, no tricks.

⌘⇧L

Screenshot answers · 16 tasks

An exam question on screen, answered from the picture alone. We time the first word and grade the answer.

⌘⇧E

Auto-fill · 61 fields on 5 pages

Whole quiz pages filled at once: radio buttons, dropdowns and number fields, the way the Agent does it. We time each page.

Try it yourself

Don't take our word for it.

Ask the models your own question, practise on exam pages with the extension, or check that every shortcut works.

Benchmark the AI you use.

Sign in with your Peeksolve account, pick DeepSeek, Gemini or Claude, and watch all 77 tasks get graded live. A run takes about half a minute and costs a few cents of your Pro credit.

Start a run →
Method

How the score is made.

01 · Same requests

Exactly what Peeksolve sends

Screenshots shrunk like the extension does, quiz pages captured by the real extension with its field list and key badges. Same prompts, same settings.

02 · Graded strictly

Right or wrong, no partial credit

Multiple choice counts by the option picked, numbers by value (“10'816”, “37,5 %” and “1/4” all count). An answer left out is wrong.

03 · Timed from Switzerland

Measured where Peeksolve runs

Every run starts on Peeksolve's own server, through the same gateway. Times are what your answers would take, not lab numbers.

PeekScore = 70 × accuracy + 30 × speed
speed = ½ · screenshot score (first word: 0.4 s → full, 4 s → none) + ½ · auto-fill score (page: 1 s → full, 10 s → none)
FAQ

Questions

Who can run the benchmark?

Anyone with Peeksolve Pro or a Day Pass. You sign in with your Peeksolve account; the run uses your own Pro credit, usually a few cents (DeepSeek well under one cent, Claude Sonnet more).

Which models can I run?

The three you can pick in Peeksolve: DeepSeek V4.1 Flash (the standard), Gemini 3 Flash and Claude Sonnet 5.5, each with or without “Think before answering”. The leaderboard also lists other models we test officially.

Why are my numbers different from the leaderboard?

Speed changes a little from minute to minute with the providers' load, and models are not perfectly deterministic. Run it again and you'll see the spread.

Is my data in the benchmark?

No. Every run uses the same fixed tasks. Your results page shows the model and the scores, never your account.

Can I share a result?

Yes. Every run gets its own link with the full result, every answer included.

Can I try things without Pro?

Yes. The practice sets and the shortcut check are free and work with any Peeksolve plan. Running the benchmark and asking the models side by side use Pro credit.