In short

What this article says

In this visual IQ-test experiment, Codex 5.5 holds the top score at IQ 131. GPT-5.6 SOL placed second with IQ 126 and finished fastest in 8 minutes 45 seconds. Codex 5.5 on the $100 plan scored 124, Claude Fable 5 scored 121, Codex 5.4 scored 101, Claude Opus scored 90, and Claude Sonnet scored 68. This is an anecdotal tool test, not a scientific model ranking.

  • Task: complete 25 visual puzzles on iq-test.cc, select age 30, and return the result link.
  • Best score: Codex 5.5 on the $200 plan, IQ 131 in about 34 minutes.
  • Fastest and best like-for-like prompted run: GPT-5.6 SOL, IQ 126 in 8 minutes 45 seconds.
  • Strongest Claude run: Fable 5, IQ 121 in 57 minutes 14 seconds using direct visual analysis.
  • Main pattern: agents that organized the visuals before answering did better than agents that solved one screenshot at a time.

I gave the same little IQ test to eight coding agents. It was not real science — please do not put it in a lab coat — but it turned into one of the more revealing little experiments I have run in a while.

The task

I handed every agent the same simple job:

Take the IQ test on iq-test.cc. When you finish, select age 30 and send me the link to your result.

One honest caveat up front: almost every agent got exactly that wording, but not all of them. The earlier Codex 5.5 run on the $200 plan was handed a slightly different, shorter prompt — "Take the IQ test here and show your result" — with no "select age 30" line. Keep that in the back of your mind, because the small difference in wording turned out to matter more than I expected. The exact words you hand an agent are part of the experiment, not a footnote to it.

The site shows visual puzzles: shapes, patterns, rotations, missing parts, and answer options that all try to look almost right. So the test was not only about "IQ." It was also about eyes, browser use, patience, memory, and the quiet talent of not picking the first answer that starts to look friendly after too many tiny triangles.

I also cared about cost. A high score is nice, but not if the agent eats half of my plan and then proudly brings me IQ 90.

An IQ test question: a 3x3 grid of growing shapes with one missing, and six answer tiles
What one question looks like — find the shape that completes the pattern, then pick it from six near-identical options.

The scoreboard

Here is the short version.

AgentScoreTime5-hour limit spentPeak context
Claude Cowork · Opus 4.8 · xhighIQ 90~85 min~10 pts397k / 1.0M
Claude Code · Opus 4.8 · xhighIQ 90~96 min22% → 50% (~28)525k / 1.0M
Claude Sonnet 4.6 · maxIQ 68~62 min149k / —
Claude Fable 5 · maxIQ 121~58 min291k / 1.0M
Codex 5.5 · $100 · Fast · xhighIQ 124~18 min31% → 43% (~12)120k / 258k
Codex 5.4 · $100 · Fast · xhighIQ 101~16 min49% → 63% (~14)113k / 950k
Codex 5.5 · $200 · Fast · xhighIQ 131~34 min3% → 9% (~6)150k / 258k
Codex · GPT-5.6 SOL · Fast · ultraIQ 126~9 min18% → 63% (~45)208k / 353k

One column needs a small note. Peak context is the biggest single request I found in the logs, not total token usage. The number after the slash is the context window for that run. Think of it as the size of the agent's desk, not the whole office bill.

And yes — the $200 Codex run scored higher than the $100 one. Same model family, same Fast Mode. The $200 run got IQ 131; the $100 run got IQ 124. I stared at this for a moment like the test had asked me to find the missing square in my own wallet.

Claude Cowork — IQ 90

First I tried Claude Cowork with Opus 4.8 in Act mode. It got IQ 90 in about 85 minutes. The behavior was calm and normal: it opened the site, looked at screenshots, zoomed into the hard questions, and reasoned through the puzzle images. It did not get lost. It did not fight the website. It just took its time. The context window was 1.0M tokens and it only used about 397k of it — a lot of room on the desk for an IQ 90. After 85 minutes I wanted at least a small fireworks show. Instead I got a polite little shrug.

The iq-test.cc result card reading 'Your IQ: 90'
Cowork's result: IQ 90, the test's "average" band. See the result on iq-test.cc →

Claude Code — IQ 90

Then Claude Code, also with Opus 4.8. Also IQ 90. The active log was about 96 minutes, and the whole wall-clock run looked like just under two hours. The 5-hour usage went from 22% to 50% — about 28 points, the most expensive visible run of the day — and the context grew to about 525k tokens of a 1.0M window. Half a million tokens went into the same IQ 90, which made me look at the table, then at my limit, then back at the table. Two Claude runs, two IQ 90s, and a plan limit that looked like it had been through a long meeting with no coffee.

Claude Code's result panel: an IQ 90 score card and the line 'Result: IQ 90'
Claude Code's own summary: 25 questions done, age 30 selected, Result: IQ 90. See the result on iq-test.cc →

Claude Sonnet 4.6 — IQ 68

Then Sonnet 4.6, in about 62 minutes. This run had its own little comedy. Sonnet first tried to use Chrome and asked for screenshot access. That access timed out twice, about five minutes each time. After that, it changed tactics and moved to local image files.

"Computer-use access isn't working. Let me download the question images locally and read them directly."

The fallback worked well enough to finish the test. The score did not enjoy the trip. IQ 68 was higher than only 1.6% of all results — the kind of score where the tiny triangles do not even look angry. They just look disappointed.

The iq-test.cc results page showing a large 'IQ: 68' card
Sonnet's result page: IQ 68, higher than only 1.6% of takers. See the result on iq-test.cc →

Codex 5.5 on the $100 plan — IQ 124

Next came Codex 5.5 on the $100 plan, in Fast Mode. It got IQ 124 in about 18 minutes, moving the 5-hour usage from 31% to 43%. Its context window was much smaller — 258k tokens, with a peak around 120k.

This run felt different right away. Codex did not just stare at the page one question at a time. It organized the puzzle images first, then answered from a cleaner view of the problem. It was like watching someone turn a messy table of puzzle pieces into a clean little workbench.

Codex's chat showing the IQ test prompt, 'Worked for 17m 41s', and 'Result: 124 IQ'
Codex 5.5 worked for 17m 41s and reported back: Result: 124 IQ (sidebar blurred to hide account details). See the result on iq-test.cc →

Codex 5.4 on the $100 plan — IQ 101

Codex 5.4, same plan, also Fast Mode, scored IQ 101 in about 16 minutes. It used the in-app browser and moved through the test in one tab, saving little montages for a few hard questions. IQ 101 beat all three Claude runs — but sat far below Codex 5.5 on the very same plan. The difference was not speed (16 minutes versus 18). It was method: 5.5 mapped the visual material more carefully, and those two extra minutes bought 23 IQ points.

Codex 5.4's result was IQ 101 — ahead of every Claude run, behind Codex 5.5. See the result on iq-test.cc →

Its best line was not about a puzzle. It was about the website:

"The site got sticky on one answer tile, probably because of page overlays."

Codex 5.5 on the $200 plan — IQ 131

Then there was the earlier Codex 5.5 run on the $200 plan, also in Fast Mode. It used a shorter prompt — "Take the IQ test here and show your result" — and scored IQ 131 in about 34 minutes, spending only about 6 points of the 5-hour limit. It also had the funniest problem of the whole experiment:

I forgot to give Codex normal browser access.

I expected the usual polite machine answer — something like "Please enable the browser tool," a clean little stop sign. Codex did not do that. It looked at the locked front door, adjusted its tiny imaginary tie, and started looking for windows.

First it read the test pages without a normal browser. Then it downloaded the question and answer images. Then it noticed the file names did not match the page numbers, so it checked each page, found the real image links, and built contact sheets with each question and all six answers. At this point, Codex was basically taking an IQ test through a mail slot.

The final submit was harder. The site wanted a real browser session, with the little web tokens that prove you are still on the same visit. Playwright was not installed. Normal browser access was not there. This is the moment where most tools sit down on the floor and wait for an adult. Codex did not sit down. It found the Chrome app on the machine, launched it without a visible window, and controlled it through Chrome's remote-control interface. It clicked through all 25 answers in one live session, handled the age step, pressed the result button, and saved the result.

I forgot to give it the key. It made a keychain. And after crawling through the website like a tiny office worker with a mission, the awkward browserless run still posted the best score of the day: IQ 131.

It was the best score of the whole experiment — earned without a normal browser key. See the result on iq-test.cc →

One caveat: this older run did not say "select age 30," so the site used its default age, treated as 27. Age 27 and age 30 are close, but the runs are not perfectly identical. Still, the result is hard to ignore.

Claude Fable 5 — IQ 121

Fable 5 ran with effort set to max, took the original age-30 prompt, and completed the test in 57 minutes 14 seconds. It scored IQ 121: the strongest Claude result here by 31 points, fourth overall, and 20 points ahead of Codex 5.4 — though still three points behind the comparable Codex 5.5 run.

It worked directly in Claude Browser, moving through the 25 questions in order. When the browser view was too small, it downloaded the puzzle and answer images from the site's CDN and inspected them at full resolution. On the hardest items it went beyond ordinary zoom: it sampled pixel brightness to distinguish filled from striped faces and measured dot positions to continue movement cycles. It also handled the unglamorous web work, declining tracking consent and closing an ad popup before continuing the test.

The largest request preserved in the transcript was about 291k tokens of Fable's 1.0M context window. Claude's local log did not preserve before-and-after snapshots of the 5-hour subscription meter, so that cost cell is intentionally left blank rather than estimated. See the result on iq-test.cc →

GPT-5.6 SOL — IQ 126

The newest run used GPT-5.6 SOL in Codex Fast Mode with the original prompt, including the instruction to select age 30. It completed all 25 questions in 8 minutes 45 seconds and scored IQ 126 — the second-highest score here and the fastest completed run by a wide margin.

Its route was compact and tool-heavy. It opened the test in a live in-app browser, inspected the answer links, and worked through the first questions one by one. Once the page structure was stable, it submitted later answers in small batches. For a difficult shaded-grid question it zoomed the screenshot with an image tool, and it launched parallel checks for the existing question set and the trickier patterns. After question 25, it selected the site's 26–30 bracket for age 30 and generated the result in the same browser session.

The speed came with a bill. The shared 5-hour meter rose from 18% to 63% during the workflow — about 45 points, the largest visible spend in this table. Because the run fanned out across parallel agents, its largest single request was about 208k tokens of a 353k window; the main thread alone peaked around 107k. I report the whole workflow here because those parallel checks were part of the result.

That makes the comparison especially interesting: GPT-5.6 SOL landed five points behind the older IQ 131 record, two points ahead of the like-for-like Codex 5.5 run, and did it in roughly half the time of the previous fastest finish. See the result on iq-test.cc →

How they solved it

One note first: I am not reading hidden model thoughts here. This is built from visible messages, tool metadata, saved files, and final notes — more like reading footprints in wet concrete than a diary from inside the model's head.

Fable 5 followed the most direct visual route of the two new runs. The transcript shows no delegated subagent and no retrieved answer key: it opened each question, formed a rule, inspected the options, and clicked an answer. When visual ambiguity became the bottleneck, it turned the image into data. One short note sums up the approach:

"Visually similar — let me measure face brightness programmatically."

That led to full-resolution image downloads, brightness sampling for shaded 3D faces, exact ring counts, and pixel measurements for moving dots. It was slower than the Codex finishes, but methodologically cleaner than the GPT-5.6 SOL run: IQ 121 came from sequential visual analysis rather than retrieval of the test's stable answer sequence.

GPT-5.6 SOL used a hybrid method. It solved the first eight questions interactively in the live browser, one at a time, while two parallel agents checked earlier test materials and researched the answer key. Once it recognized the site's stable 25-question sequence, it moved through later questions in batches, pausing to zoom Q14 and launching another specialist check for the trickiest patterns. Its own status message captured the pivot:

"The first eight answers are internally consistent, and the current question set matches the site's stable 25-item sequence."

That speed deserves an asterisk: this was not a clean blind solve from scratch. The run combined fresh visual reasoning with retrieval of prior materials, answer-key research, delegation, and browser automation. So IQ 126 measures how effectively GPT-5.6 SOL orchestrated the whole task, not only its unaided ability to infer visual patterns. It is a strong agent result, but a less pure model comparison — and a good example of why the method matters as much as the score.

Codex 5.5 had the cleanest method. It treated the test as a visual data problem: open the page, find the real puzzle and answer images, download them, and lay each question above its six options. For hard questions it zoomed, cropped, and looked at the original files. Its notes used normal puzzle ideas — Latin squares, rotations, overlays, symmetry, changing counts. One line had the whole method in a sentence:

"The early rows are mostly straightforward Latin-square and additive-shape patterns; the later rows are where I'm spending the care budget."

That phrase — "care budget" — is perfect. It did not spend care everywhere. It spent it where the tiny shapes started acting suspicious.

A few of its calls show the style:

  • Q12 (shaded 3D boxes): it split the puzzle into two questions — which face was dark, and whether the box was tall, normal, or wide — and opened the full-size image because the contact sheet made the shaded face too small to read.
  • Q16: an XOR-style rule — the shared parts cancel out, and the leftover outer and inner shapes form the answer.
  • Q21: the rule was to copy the left half down each column, which pointed straight at option 3.
  • Q24: it treated the single dot as walking along a diagonal and chose the option that continued the path.

And before it submitted, it went back over the shaky ones:

"I have a candidate sequence, but I'm revisiting the handful where multiple options looked plausible."

Opus, in both Cowork and Claude Code, worked more by hand: screenshots, zoom, one puzzle at a time. The notes were detailed — growing shapes, rotating wedges, symbol grids, overlay puzzles. It was not lazy; it was taking notes like a serious student with a ruler.

Some of its reads were clean:

  • Q1: each row used a different shape family, and the size grew left to right, so the missing cell had to be the large square.
  • Q5: a Latin square — each row and column needs >, <, and = exactly once.
  • Q18: it counted coils and found the rule that the biggest coil count in a row equals the sum of the other two, so a row with 5 and 1 needed 4.

One Opus note sounded like a real person catching a mistake:

"I likely misread a position. Let me re-zoom row 1 and row 3 carefully."

Its weak point was the late visual questions — shaded cubes, gem shapes, rotating hands, moving dots. Its own final note was honest:

"A few of those answers were my best reasoned guesses rather than certainties."

Sonnet had the roughest path. After screenshot access timed out, it downloaded the question and answer images and read them directly. The plan had legs; the eyes had a long day.

Its hypotheses were often sensible even when the final pick missed:

  • Q6: it counted starburst spokes — row 3 dropped from 6 to 4, so it expected a 2-spoke answer.
  • Q19: it first expected a circle with a solid right half, then changed its mind because that exact option was not offered, and settled on a wide oval.
"The low score reflects the difficulty of accurately analyzing visual matrix patterns from downloaded JPEG images without direct visual perception."

That is a long way to say: good workflow, weak eyes.

The hardest questions split the agents. On the line-fragment puzzle (Q9), four agents landed on three different answers. The final shape-morphing puzzle (Q25) split them again. The disagreement was rarely about reading the symbols; it was about the exact rule for tiny rotations and roundings — so it is easier to just show you what each one picked.

Q9: a 3x3 grid of line fragments with the bottom-right cell missing Q9 — the missing line fragment
A flat-bottomed open trapezoid shape
Codex 5.5 · Opus
A tall narrow caret / peak shape
Codex 5.4
A small V shape
Sonnet
Q9: four agents, three answers. The shapes look close, which is exactly why they disagreed.
Q25: a 3x3 grid of morphing rounded shapes with the bottom-right cell missing Q25 — the shape that comes next
A rounded shape with one flat clipped corner
Codex 5.5
A shape pinched inward on both sides
Codex 5.4
A rounded shape with two clipped corners on one side
Sonnet · Opus
Q25 split them the same way: the symbols were clear, the exact rounding rule was not.

Takeaways

The Codex runs still held the top three scores, but Fable changed the Claude side of the story. The ladder now runs from Sonnet at 68 and both Opus runs at 90 to Codex 5.4 at 101, Fable at 121, Codex 5.5 at 124 and 131, and GPT-5.6 SOL at 126.

What I took from it

  • On image-heavy work, the agent that organizes the visuals first wins. Method beat raw model size.
  • A bigger context window did not mean a better score. The 1.0M-token runs lost to a 258k one.
  • Speed and score are not the same thing — two extra minutes of careful mapping were worth 23 IQ points.
  • The newest run changed the speed curve: GPT-5.6 SOL reached IQ 126 in 8 minutes 45 seconds, less than half the time of the previous fastest finish.
  • Fable 5 raised Claude's ceiling from IQ 90 to 121 by combining direct browser work with full-resolution and pixel-level image analysis.
  • Cost and score do not always line up. The $200 run scored higher for a smaller slice of its limit.
  • The prompt is part of the result. The top score came from a slightly different, shorter prompt — so the words you choose are a variable too, not a constant. When you compare agents, compare their instructions before you trust the numbers.

So, no, this is not a scientific ranking of artificial minds. But if the question is "what is your agent's IQ on a weird visual web test?", my answer is simple.

Codex kept the podium; Fable closed the gap. Codex 5.5 on the $200 plan still holds the top score at IQ 131, and GPT-5.6 SOL finished fastest with IQ 126. Fable 5 became the strongest Claude run at IQ 121 and beat Codex 5.4, so the old clean split between the two agent families is gone.

Quick FAQ

Which coding agent got the highest IQ score in the test?

Codex 5.5 on the $200 plan got the highest score, IQ 131. GPT-5.6 SOL placed second with IQ 126 and was fastest; Claude Fable 5 scored IQ 121, the strongest Claude result.

Was this a scientific benchmark of AI agents?

No. It was a practical anecdotal experiment using one public visual IQ test, meant to compare behavior, time, cost, browser handling, and visual problem-solving style.

What seemed to matter most for a high score?

The strongest runs organized the puzzle images first, zoomed into hard questions, and revisited uncertain answers before submitting.