**Anthropic’s Claude Opus 5 and Mythos 5 currently lead HLE leaderboards at 64.7% and 64.5% (with tools), followed by Meta’s Muse Spark 1.1 at 62.1% and OpenAI’s GPT-5.4 Pro at 58.7%, according to August 2026 snapshots across BenchLM, Scale, and Artificial Analysis.** The 2,500-question expert-authored benchmark—spanning math, sciences, and humanities—remains far from saturated, with text-only frontier scores in the mid-40s% and tool-augmented results higher due to iterative reasoning, verification loops, and extended context. Rapid 2026 gains stem from specialized reasoning models and agentic scaffolding, outpacing earlier projections. Traders watch remaining 2026 releases from major labs, as further gains in chain-of-thought, tool integration, or new architectures could lift the year-end maximum before resolution on the official leaderboard.
Polymarket डेटा का संदर्भ देने वाला प्रयोगात्मक AI-जनरेटेड सारांश। यह ट्रेडिंग सलाह नहीं है और इस बाज़ार के समाधान में कोई भूमिका नहीं निभाता। · अपडेट किया गया$44,583 वॉल्यूम
60%+
74%
65%+
35%
70%+
18%
75%+
10%
80%+
7%
90%+
6%
$44,583 वॉल्यूम
60%+
74%
65%+
35%
70%+
18%
75%+
10%
80%+
7%
90%+
6%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
बाज़ार खुला: Jul 23, 2026, 6:41 PM ET
समाधान स्रोत
https://agi.safe.ai/रिज़ॉल्वर
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
समाधान स्रोत
https://agi.safe.ai/रिज़ॉल्वर
0x65070BE91...**Anthropic’s Claude Opus 5 and Mythos 5 currently lead HLE leaderboards at 64.7% and 64.5% (with tools), followed by Meta’s Muse Spark 1.1 at 62.1% and OpenAI’s GPT-5.4 Pro at 58.7%, according to August 2026 snapshots across BenchLM, Scale, and Artificial Analysis.** The 2,500-question expert-authored benchmark—spanning math, sciences, and humanities—remains far from saturated, with text-only frontier scores in the mid-40s% and tool-augmented results higher due to iterative reasoning, verification loops, and extended context. Rapid 2026 gains stem from specialized reasoning models and agentic scaffolding, outpacing earlier projections. Traders watch remaining 2026 releases from major labs, as further gains in chain-of-thought, tool integration, or new architectures could lift the year-end maximum before resolution on the official leaderboard.
Polymarket डेटा का संदर्भ देने वाला प्रयोगात्मक AI-जनरेटेड सारांश। यह ट्रेडिंग सलाह नहीं है और इस बाज़ार के समाधान में कोई भूमिका नहीं निभाता। · अपडेट किया गया



बाहरी लिंक से सावधान रहें।
बाहरी लिंक से सावधान रहें।
अक्सर पूछे जाने वाले प्रश्न