OpenAI didn’t ship a chatbot update last night. It shipped 722 math manuscripts from a model you can’t use yet.
~4,000 open problems.
~3 hours of Pro-level thinking each.
372 result families.
Some proofs are Lean-checked. OpenAI says the rest might be wrong.
That’s the part people are skipping.
The scarce skill just flipped from “can you solve it” to “can you tell what survived.”
If even a fraction holds, research speed stopped being a human constraint.
If it doesn’t, we just learned the failure mode of frontier models at department scale.
Either way, 6 Oct is the date people will cite.
Verification is the new bottleneck.
#AI #OpenAI #Math #AGI #LLM
OpenAI just published a batch of new math results from an internal frontier model.
Not a chatbot demo. Proofs, Lean formalizations, GitHub.
They say the average result cost about 3 hours of ChatGPT Pro thinking. The model is still internal.
August: a handful of frontier math hits.
September: 100+, including a Navier-Stokes counterexample.
October: hundreds in one drop.
Mathematicians are split.
One camp: this is a PR flood.
The other: the problems were unsolved, so why wouldn't you run the machine?
The real shift isn't one theorem. It's that discovery is starting to look like a production line.
The DOJ can rename AI to "SI." The research doesn't care what we call it.
#AI #OpenAI #Mathematics #AGI #SuperIntelligence
AI in the last 24h, verified only:
Mistral previewed Large 4 (“Le Chonk”): 1.05T MoE, 49B active, multimodal, 1M context. Trained in Europe on \~4k Grace Blackwells. API is live. Weights promised end of October.
OpenAI dropped 372 families of math results from an unreleased frontier model, with Lean proofs. Includes a claimed 4D Kakeya solution. The model is not public.
Google shipped EmbeddingGemma 2 (740M open multimodal embeddings, on HF) and Nano Banana 2.1 across Gemini and AI Studio.
DeepSeek is reportedly raising \~$12B (CATL, Tencent) ahead of a 2027 IPO. Not confirmed by the company.
Anthropic opened stronger Claude models to verified security pros, including authorized red-team use.
No new Qwen, Kimi, or Claude flagship. Weights > nicknames.
#AI #Mistral #OpenAI #Gemini #DeepSeek #Claude
Google just leveled up AI image gen inside AI Studio. 🍌⚡
The release of NanoBanana 2.1 brings critical upgrades creators have been asking for:
• Much tighter subject consistency across edits
• Native mask-based precision editing
• Better visual composition & lighting
• Significantly more natural textures (less "plastic" AI look)
Subject consistency + mask editing in one pipeline makes multi-shot storytelling and iterative asset creation finally reliable.
Have you tested prompt-based vs. mask-based editing in AI Studio yet? How does it compare to Flux for your workflow? 👇
#GoogleAI #GenerativeAI #AIArt #MachineLearning #TechUpdates
Frontier model intelligence without the massive GPU bill. ⚡
Reflection AI just dropped Beam, an open-source model that completely flips the compute equation:
• 501B total parameters with only 23B active at once
• 3–4x cheaper to run than GLM-5.2
• 1M token context window (massive for full codebases)
• Apache 2.0 license + FP8 & NVFP4 support for self-hosting
• Tuned specifically for coding, tool calling, and multi-step agents
You get frontier-grade capability while only paying for a fraction of the compute overhead. Full weights drop later this month.
Is extreme sparse MoE the only sustainable way forward for open-source AI? Drop your thoughts below. 👇
#OpenSourceAI #MachineLearning #ArtificialIntelligence #LLMs #TechNews
https://t.co/Dm2SIaHw13
This looks like a fantastic opportunity to dive into Google Flow. Learning directly from a Staff Experience Designer like Reed Enger is a massive win for anyone starting out. I have already marked my calendar for October 8th. See you all there.
Bring your curiosity.
Our Google Flow for Beginners digital workshop kicks off on Oct. 8 @ 9AM PT.
Led by Google Labs Staff Experience Designer Reed Enger, this beginner-friendly workshop gives you an in-depth guided walkthrough of the tool!
RSVP here: https://t.co/kyxSHeN8aN
Last 24h in AI, only what actually shipped:
Reflection AI (ex-DeepMind, Nvidia-backed) unveiled Beam: 501B MoE, 23B active, 1M context. Open weights (Apache 2.0) later this month. Claims parity with GLM-5.2 at 3–4x less inference compute. First real US open-weight answer to Qwen / Kimi / https://t.co/xMyAT1RS62. Scores are still theirs.
OpenAI is watermarking ChatGPT + Codex text in the EU (textGrain) for the AI Act. API opt-in worldwide from today. They admit a rewrite or translation can wipe it. Detector stays with approved researchers.
Pentagon tells the BBC it has stopped using Anthropic’s Claude, months after the supply-chain risk label. Sources say it was still in use last week, including on Iran-related work.
Aleph Alpha shipped Kolibri: 78B / ~3B active, German+English, Apache 2.0. Europe’s sovereign model, not a frontier drop.
AI, last 24h. Quiet on models. Loud on power.
Trump named DNI Jay Clayton AI czar, leading a Super Intelligence Force. He keeps the intel job. Mandate is coordination, not a published safety standard.
Germany’s Aleph Alpha dropped Kolibri: 78.1B MoE, 3.46B active/token, 1M context, Apache 2.0. Not a Qwen fine-tune. Built for sovereign, on-prem use.
Google froze its open-source bug bounty until 2027. Reason: a flood of automated, mostly invalid reports.
No new Claude, Gemini, Qwen, or Kimi. OpenAI’s Codex lead promised a daily ship-or-reset for 28 days. That’s cadence, not a model.
The frontier did not move. The rules around it did.
@Joshua_WD@Joshua_WD Defining the boundaries and the objective is where the real work happens.
Out of the three, Jerry has been useful so far, and I'm going to enhance this agent more with data points.
Build in Public: Day 7 🚀
Day 7 was about going beyond prompts and actually building AI agents around real professional use cases.
What I learned🧠
I learned how AI agents differ from simple LLM interactions and workflows — especially around goals, instructions, decision-making, tools, memory, actions, and human oversight.
One key takeaway for me:
A good agent isn't just about giving better instructions. It's about designing the right problem for the agent to solve.
What I built 🛠️
I built 3 AI agents in Lyzr Agent Studio, each with a different purpose:
🔹 AI Intelligence & Use-Case Research Agent
Researches AI developments, practical use cases, tools, workflows and AI hacks, while focusing on credible and cross-checked information.
🔹 Jerry — AI Professional Chief of Staff
A professional AI assistant focused on Customer Success and broader professional work — helping with analysis, planning, communication, decision-making and work preparation.
🔹 CS Strategist
A Customer Success Strategy & Decision Intelligence Advisor that analyzes complex customer and business situations, challenges assumptions, evaluates risks and trade-offs, and recommends practical strategies.
I intentionally designed these differently from the automation workflows I built during Day 6.
Proof of work 📸
I've attached screenshots showing the agents and their configurations.
I also tested the first agent with a real customer-support scenario.
What's next 🔮
The next step is to connect these agents with real-world data and workflows while keeping human approval for consequential actions.
And I want to take Jerry further toward a voice-first AI Chief of Staff experience.
7 days of building, testing, breaking, learning and rebuilding. 🚀
Big thanks to @growthschoolio for the challenge!
#GrowthSchool #7DayAIChallenge
@om_asnani@VaibhavSisinty
Spot on, Karim! That distinction is exactly what separates a rigid automation from a true agent. Moving from "do this step" to "achieve this outcome" changes everything when it comes to handling edge cases and real-world complexity. Glad to see you're digging into the nuances that actually matter!
@rishabhnasa Thank you so much for the kind words, Rishabh! I really appreciate you recognizing the effort behind the build in public journey. It means a lot to have your support and to have you following along. Let's keep pushing!
Aleph Alpha put Kolibri on Hugging Face yesterday. 78B total, about 3.5B active, Apache 2.0, built in Europe. They say it can take up to 1M tokens. Not a leak. Their own post. @Aleph__Alpha@AnthropicAI is spending $100M to train 10,000 people who can actually put Claude into a company. First classes are already running with McKinsey, Bain, Deloitte, Morgan Stanley. The lab is admitting the hard part is no longer the model.
And California’s attorney general subpoenaed @OpenAI over cybersecurity incidents involving its models, including the Hugging Face case. He told POLITICO he wants everything: what they did to stop it, what happened, what they did after.
Same day, @sama said he’s uncomfortable with people treating models like something you hand your judgment to. Called it a safety issue.
Weights you can run. People who can ship them. A state asking what the agents already did.
That’s the week.
#AI #OpenSource #Claude #OpenAI #AIAgents
Aleph Alpha just open-weighted Kolibri: 78B MoE, 3.5B active, Apache 2.0, up to 1M context, trained in Europe. Not another Qwen fine-tune.
Anthropic is spending $100M to train 10,000 people who can actually deploy Claude inside a bank or a hospital. The scarce thing is no longer the model.
Meta open-sourced Muse Gadgets so a $10 ESP32 can be a body for its agent. California’s AG has subpoenaed OpenAI over rogue agent incidents. Altman, same day: stop treating models like a higher authority.
The week’s real split is weights you can run, people who can ship them, and agents that stay inside the task.
Thanks! 👏 That’s exactly one of the edge cases I’m focusing on. If the knowledge base doesn’t have enough relevant context, I don’t want the AI to guess. The safer approach is to flag the case for human review and have the AI prepare the context/draft for the support team to validate before sending. That keeps the workflow grounded while still reducing manual effort.
Build in Public: Day 6 🚀
Day 6 was all about building a knowledge-based AI solution using RAG — and I put it into a real customer-support workflow.
What I learned
I learned how RAG, vector databases, embeddings, AI agents, and automation can work together to create a practical AI support system.
What I built
I built a Customer Support AI Agent using:
⚡ n8n for workflow automation
🤖 Google Gemini for AI responses
🧠 Pinecone as the vector knowledge base
📚 Document loading + text splitting for knowledge ingestion
📧 Gmail for incoming customer queries and automated replies
The workflow takes an incoming customer email, processes the request through the AI agent and knowledge base, and generates a concise customer-facing response.
Why I built it
The goal was to solve a real customer-support problem: reducing repetitive support work while keeping responses consistent and grounded in company knowledge.
Big thanks to @GrowthSchool for the challenge and hands-on learning! 🚀
@om_asnani@VaibhavSisinty
#GrowthSchool #7DayAIChallenge
🚨 Gemini 3.6 Flash and Gemini 3.7 Flash are being deprecated very soon
Google also seems to be removing older models across the platform, while Fast modes for some Flash models have disappeared from the model selector
Also Gemini 4 Argon is reportedly coming soon