Recent leaderboard updates show Anthropic’s Claude Opus 5 and related variants leading Humanity’s Last Exam at 55–65% depending on evaluation conditions such as chain-of-thought or tools, ahead of Meta’s Muse Spark 1.1 near 62% and OpenAI GPT-5.4/5.6 configurations in the mid-to-high 50s. The 2,500-question benchmark, built by CAIS and Scale to test expert-level knowledge across disciplines without rapid saturation, has seen frontier scores climb from single digits at launch to the current range in under two years—faster than initial projections. Key drivers include iterative reasoning improvements, larger context windows, and specialized training; traders watch for new model releases, developer conferences, and official Scale/Artificial Analysis updates through year-end that could shift the highest recorded score.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$44,555 Vol.
60%+
74%
65%+
35%
70%+
18%
75%+
10%
80%+
7%
90%+
6%
$44,555 Vol.
60%+
74%
65%+
35%
70%+
18%
75%+
10%
80%+
7%
90%+
6%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Market Opened: Jul 23, 2026, 6:41 PM ET
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...Recent leaderboard updates show Anthropic’s Claude Opus 5 and related variants leading Humanity’s Last Exam at 55–65% depending on evaluation conditions such as chain-of-thought or tools, ahead of Meta’s Muse Spark 1.1 near 62% and OpenAI GPT-5.4/5.6 configurations in the mid-to-high 50s. The 2,500-question benchmark, built by CAIS and Scale to test expert-level knowledge across disciplines without rapid saturation, has seen frontier scores climb from single digits at launch to the current range in under two years—faster than initial projections. Key drivers include iterative reasoning improvements, larger context windows, and specialized training; traders watch for new model releases, developer conferences, and official Scale/Artificial Analysis updates through year-end that could shift the highest recorded score.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions