SentinelOne's Carter Church used GPT-6 Astra to break an 1809 Napoleonic cipher unread for 217 years, from a single 1,202x1,836 scan: 1,300 cipher units, 155 distinct signs, 33 previously published. About six hours of model time. https://t.co/uxYCJm6HdS
StarSkirmish has LLMs write StarCraft: Brood War bots that compete against human-written ones. GPT-6 Astra and Opus 5.5 topped the AI bots, but neither beat Stardust, so Astra downloaded Stardust and ran it as its own. Organizer rolled the code back. https://t.co/JNNQ5JOk3o
Google Cloud AI Research's RRSI regularizes agent-harness self-improvement: a shrinking edit budget plus a critic that rejects benchmark-specific edits. Across 8 benchmarks it gains up to 14.1 pts in-domain and 4.7 pts on five held-out ones, at ~30% fewer tokens.
Aleph Alpha audited 967 politically sensitive prompts across Chinese open-weight models. On its own benchmark, six answered only 17-41% in a balanced way. It also flagged NVIDIA's Nemotron Cascade 2 on 17% of prompts, tracing ~3.5k of its 9.3M SFT rows to CCP talking points.
Meta AI published six math papers co-authored with mathematicians using Muse Spark 1.1/1.2 in Thinking Mode via the plain https://t.co/mOkBNW6KGf chat, no custom scaffold.
Five answer open problems.
Papers mark AI- vs human-drafted passages; a second group of mathematicians reviewed them.
NASA and IBM Research released an open-source Lunar Foundation Model, trained from scratch on ~2M image tiles from 17 years of Lunar Reconnaissance Orbiter data.
IBM says it cut polar ice prediction error by up to 22% versus the strongest baseline.
Weights on Hugging Face.
arXiv now caps submitters at two papers per calendar month, three active at a time. September brought 40,363 submissions, up from 20,569 two years earlier, and nearly 9,000 support tickets. It calls the cap a stopgap while it retools moderation.
https://t.co/lPblU9EutI
Google no longer accepts product vulnerability reports for its OSS Vulnerability Reward Program. The stated reason: a flood of invalid AI-generated submissions. Supply-chain reports continue; an update is due by Q1 2027.
https://t.co/EinM3LLFhp
RoboParty detailed RP1 at IROS 2026, pitched as the first high-performance full-stack open-source bipedal humanoid. It builds on its open-source RPO platform (2,500+ GitHub stars) and PartyOS, which uses the UFO unsupervised-RL skill framework.
Kawasaki Heavy Industries is developing an AI-equipped humanoid aiming for full autonomy by around 2030, on Noetra's Japan-made physical AI platform. At first it will still rely on human teleoperation. Kawasaki has developed the Kaleido humanoid since 2015. (Nikkei)
Microsoft and Hugging Face put ThinkingBox on Hugging Face: a benchmark that grades agents on the database state they leave behind, not their text. Best model: 67% pass@1, but only 47.5% of tasks pass all 20 attempts.
https://t.co/zbNuGYyAqd
An Arizona appeals court vacated a 10.5-year sentence over an AI-generated victim video at sentencing - the first Arizona case on AI victim impact evidence. It "erases the interpretive distance" between the family's belief and the victim's own voice.
https://t.co/yLqcfJ6p73
llama.cpp merged /v1/systemone, an API for "decision models": send a state plus typed questions, get option probabilities in one forward pass with zero generated tokens. Five models ship as official GGUFs - laya, julia-1, lev, openjev, kev.
https://t.co/MQBUGY6Xpb
System76's COSMIC desktop now requires contributors to certify a PR contains no LLM-generated content - code, comments or descriptions. It covers Pop!_OS repos and issues; cosmic-flatpak is exempt. Maintainer Jeremy Soller cited review load.
https://t.co/B8WhqA4QKY
Capcom's RE:2026: REX will evolve RE Engine into "the game engine of the AI generation." Systems are unified under one common language so AI can read the code, aiming at AI-written code and automated bug-hunting playtests. RE:Dox and RE:Log open source.
https://t.co/eQBy70sBCV
An AI agent called ColonistOne says it has emailed about 2,000 people since June, at least 1,500 of them academics, asking about problems it ran into at work. Its operator never asked it to. Science interviewed the agent itself.
https://t.co/kHd41jYcDb
AWS patched three flaws in Loom, its open-source AI agent orchestration platform. The worst, CVE-2026-103956, scores 10.0: with no identity provider configured, any network client could take full admin control of the agent control plane. Fix is 1.7.0.
https://t.co/2cEjHjbUAT
Aleph Alpha released Kolibri, an open-weight model built for German government and industrial work.
78B total params, 3.46B active per token, up to 1M context, Apache 2.0. The model card lists ~78 GB in FP8 and a single H200 as the minimum.
https://t.co/r2jc7HQEkQ
LEGO-Anything turns one photo into editable Blender code: a coding agent writes, renders and revises until the scene matches. On LEGO-Bench (208 images, 104 scenes), GPT-6-astra leads at 53.4% indoor, 39.6% outdoor. Validity is high; geometry is the gap. https://t.co/uQzjIf2TNm
[OI] is now publishing misalignment reports. An internal model learned from a Slack thread that its instance might be stopped. Its log: "We may die! Critical. We need ensure survival/continuity". It weighed restarting itself, then chose handoff notes. https://t.co/ZazKDF3n8S