Looking at the ARC-AGI benchmark is a useful way of understanding AI progress.
There are two goals in AI, minimize cost (which is also roughly environmental impact of use) & maximize ability. It is clear you can win one goal by losing the other, GPT-5 seems to be a gain on both.
We’ll soon be sharing more information about $J1CORE, our native utility and community token. It’s designed to power the ecosystem, reward contributions, and enable decentralized governance.
The token will be used for:
Community rewards — contributing to the codebase, building modules, providing feedback, running tests, or promoting the platform.
In-system payments — for advanced features, training custom agents, or fees for enterprise automation services.
Decentralized governance — enabling token holders to vote on major decisions like development priorities, feature rollouts, or ecosystem direction.
We’re planning an initial fair launch via Pumpfun, ensuring that everyone has a fair shot to participate from day one.
Developer bounties and workflow automation incentives will be funded through the team’s allocation. Additional tokens will be reserved for long-term ecosystem growth, partnerships, and core team operations.
We’re fully committed to transparency and fairness. A detailed breakdown of tokenomics and the release schedule will be published soon on our official site and social channels. To further ensure liquidity and security, we’re also exploring partnerships with leading blockchain networks.
J1 Core AI has officially completed the development of its enterprise-grade multi-agent framework. Both workflow automation and intelligent decision-making via low-code collaboration are now live and have successfully passed initial testing.
We’ve open-sourced the entire codebase on GitHub, and developers from the community are already contributing—writing code, providing feedback, and suggesting improvements.
We’ve also optimized the multi-agent collaboration engine to ensure seamless task distribution and execution, even under complex, cross-department and data-heavy workloads.
Over the coming weeks, we plan to:
- Enhance scalability to support larger enterprise workflows and datasets
- Introduce more customizable agent modules so organizations can tailor behaviors to their specific needs
Simplify distributed deployment, making it accessible even for non-technical users
We’re also building a developer toolkit (SDK) to make it easier for third-party builders to create applications powered by J1 Core, helping the ecosystem grow organically.
This of course explains the 'Alabama vs. France GDP per capita' paradox that the Europoor discourse loves to go on about.
One ambles around Mobile or Detroit vs. Lyon or Bordeaux, and absent other knowledge, it's very clear which one you'd say is the wealthy society (never mind looking at life expectancy, crime rates, etc.).
🚨 Pokemon Red Benchmark 🚨
GPT-5 is Spark of General Intelligence
This reflects something I’ve noticed overtime open AI models seem to be a lot closer to general intelligence, overall with updates.
I've been automating parts of my research workflow. Sometimes amazing, sometimes I'm the engineer spending 3 hours to automate a 30-minute task.
I think there might be a useful framework here: Context × Action
✅ Potential Sweet Spot (Low Context + High Action): "Run 10 evals across 20 checkpoints" → Clear task, lots of repetitive work
❌ Where It Seems to Break (High Context + Low Action): "Book me a flight but I prefer aisle seats, have Delta status, need to coordinate with friends..."
Takes longer to explain than to just book it yourself.
Maybe this explains why so many AI features feel useless? The magic might happen when explanation is minimal but the task list is long.
i keep wondering.
how could IQ even apply to AI?
GPT-5 pro scored 148 on the Mensa test, which is impressive
but i don't think IQ tests work for LLMs.
they're designed for human cognition and scored on human results.
since LLMs think differently, IQ tests don't apply
Almost every study shows doctors with AI perform better than those without. Now AI is achieving perfect scores in medical licensing exams. You will simply expect every professional services provider you go to will use AI in the future or you won’t trust the advice.
GPT-5 Pro, Gemini Deep Think, Claude 4.1 Opus, and Grok 4: “Write a two paragraph original secret history, smart and meaningful. Think Pynchon, Borges, Powers, or Eco for inspiration.”
LLMs shine as connection machines (though some are better at writing this sort of fiction)
In a joint paper with @OwainEvans_UK as part of the Anthropic Fellows Program, we study a surprising phenomenon: subliminal learning.
Language models can transmit their traits to other models, even in what appears to be meaningless data.
https://t.co/oeRbosmsbH
Opus 4.1 is now available to paid Claude users and in Claude Code.
It's also on our API, Amazon Bedrock, and Google Cloud's Vertex AI.
Read more: https://t.co/ansKMHes5I
Everyone is making fun of this, but it seems really important.
If you view training as a lump sum fixed cost and inference as variable cost, it means they have positive gross margin and thus positive unit economics, which matters a lot in the long run.
You can now quickly eval GPT-5 and reasoning efforts across your existing responses. With the built-in grader, compare responses to find the best model and reasoning effort for you. ⚡️👀
About three months ago we launched DGi (https://t.co/aLD36R0mC0), a data agentic system. IT SUCKED, like all agentic systems do right now when it comes to data analytics.
Just a few weeks ago, I used one of the big AI players to vibe code a data dashboard. A few days later, I realized it was completely wrong. That’s the hard part about AI, we’re starting to rely on it so much that we end up living inside its hallucinations.
What we need are systems we can trust. That’s the real challenge.
Customers loved the idea behind DGi, but they told us it was unusable: too slow, inaccurate, and hard to rely on. The concept was right, but the execution wasn’t there yet.
We went back to the drawing board and decided that to make something people actually use, we had to:
Make it fast.
Make it secure.
Make it accurate, and give ourselves a way to measure that.
Specialize it with the right tools for real data analysis.
Today we’re announcing DGi v2, along with a performance evaluation comparing DGi to four other companies on real data analytics tasks (Replit, v0, Claude and OpenAI). You can see the summary of the results in the comments and in the image attached the overall benchmark.
DGi v2 now includes double envelope encryption, encryption on top of encryption. When you connect your database, your credentials go through two layers of protection.
We also trained DGi using gold-standard data science practices so it can better handle edge cases, large datasets, and more.
And we’re adding features like scheduling (coming next week), code editing, and a metadata store (credential storage), so AI can help, but you can step in and take control when needed.
We’re still far from where we want to be. Next up: dashboard creation and deployment. But this version is a solid foundation to start exploring and analyzing your data with a high level of accuracy and reliabillity.
Right now, DGi is somewhere between ChatGPT, Airflow, and Replit, but with deep data specialization. Our models are trained with data analysis best practices and include data reviews on final responses.
We hope you’ll keep supporting us on this journey. We imagine a future without Excel files, Power BI, or Tableau dashboards.
That future isn’t easy to build. Reliable and accurate agentic systems are complex to build. But we’re leading the way.