Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming
Intelligence Index v4.2 changelog:
+ AA-Briefcase, our agentic knowledge work evaluation with a private test set
+ @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages
- GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated
… plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness
This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January.
We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users.
Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned!
Intelligence Index v4.2 changes in detail:
➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied.
➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5.
➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure.
Key results:
➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google
➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier
➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
@synthwavedd It's the worst launch so far, how can they release something only for organizations without making it available to the normal ppl, it's like releasing a piece of shit
I had early access to GPT-6 Astra and it's maybe the most insane model I've experienced
In Blender, I had it recreate Apple Park from just images. It did an absurd job.
GPT-6 Astra built this Manhattan world in Unreal Engine over the course of a week.
It was literally able to go street by street to make each one perfect.
"Arrival for all is coming in a few days"
Who was the product launch for? Just to look better than Fable-5.1? Are you kidding me?
I suspect this is OpenAI's worst launch so far
"Arrival for all is coming in a few days"
Who was the product launch for? Just to look better than Fable-5.1? Are you kidding me?
I suspect this is OpenAI's worst launch so far
GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0.
GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
@ZA14_M1 طيب يعني لو تقدر تسوي الحرام في بيتك ماينفع نأمر بالمعروف وننهي عن المنكر ؟
مسلمين نحنا ياحبيبي ، نقول خلاص دع الخلق للخالق والكلام دا؟
بعدين مين قال تقدر تلعب اللعبة بدون حرام؟ والعري ؟والاباحية؟والخمر ؟والمعازف؟.... الخ
وكل الأفكار الي هي ضد دينك عادي كذا؟ احا منجد
🚨للتوضيح فقط عشان مايصير فيه سوء فهم للتغريدة هذي؛
موضوع إستشراف المشاهير على لعبة #GTA6
مع إن تصنيف اللعبة بالأصل +21
ممكن أتقبل إنك تنصح الناس والآباء من باب الفائدة، الله يجزاك خير والدال عالخير كفاعله ما إختلفنا.
لكن توصل فيهم الحال يطالبون بإن اللعبة يتم حجبها عندنا😕!
مع إني أكثر شخص كاره لـGTA 6 حالياً، بسبب إنها اللعبة القذرة اللي سمحت لسوني إنها تتجرأ وتلغي النسخ الفيزيائية للألعاب.
لكن هل أنا مع حجب اللعبة؟ لا
هل أنا بشتري اللعبة؟ برضو لا
لكن اسألكم بالله عاجبكم وضعنا أول؟ العاب كثير يتم منعها عندنا بسبب لقطة تعري جانبية إحتمال تختم اللعبة وماتطلع لك؟ أو تروح تشتري الألعاب المحجوبة من المحلات كأنك رايح لبياع ويبيعها بأضعاف السعر الرسمي؟ 🤦🏻♂️
هنا تجي فائدة التصنيف العمري +21 وأشكر الهيئة العامة لتنظيم الإعلام من أعماق قلبي على هذي المبادرة، منعت جشع أصحاب المحلات وخلت قرار الشراء من عدمه يعود لك أنت شخصياً 🫵
أنت عزيزي أدرى بمصلحة نفسك، مايحتاج يجي أحد يعلمك وش الصح و وش الغلط، أو يعلمك هذي اللعبة تقدر ألعبها أو لا!
موضوع التعري والدعارة وغيرها من السلوكيات اللي تخالف مجتمعنا موجودة في كل مكان ومسموح فيها بعد، أكبر مثال Netflix.. بس ماقد سمعنا أحد يقول المفروض التطبيق يكون محجوب عندنا؟ والمثال الثاني واللي يسبق Netflix بسنوات هو تطبيق تويتر وهنا حدث ولا حرج، التطبيق موجود فيه كل شيء، ومع ذلك هو البرنامج الأكثر شهرة وإستخدام في المملكة عندنا وهو المنصة الرسمية اللي تستخدمها كل الجهات الحكومية الرسمية.. هل قد شفت أكد يطالب بمنع البرنامج عندنا؟ أكيد لا. وغيرها من الأمثلة وتطول القائمة، لكن السؤال المهم هنا هل هي إجبارية؟ أكيييد لا!
بالنهاية فكونا من إستشرافكم المقرف، تبي تحمي ورعانك من اللعبة؟ خلك رجال وربيهم تربية تناسب مجتمعنا وديننا، ولاتشتري لهم لعبة ماتناسب أعمارهم.
اذا انت غبي ماتعرف تربي؟ لا تجي تلوم الجهات المعنية بأنهم سمحوا فيها عندنا، ولا تطالب بمنعها.