After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
💯 the temptation to use AI in existing business processes is natural maybe even needed to learn & train but
“The best use-cases for AI tend to be those that fundamentally change the work being done instead of just replacing an existing process and doing it more efficiently.”
Just coming off of meetings with a couple dozen enterprise IT leaders discussing AI agents. Here are a few of the common themes that stand out:
* Lots of conversation that you have to solve an operating model challenge to get the full benefits of AI. Most companies have orgs that have always operated in siloes; but agents are most effectively when they are tied to a process, which often cuts across these siloes. So the big question is how do you start to deploy centrally managed agents that can work across organizational boundaries. Who manages these agents? How do they get deployed and adopted?
* Data fragmentation remains a major issue for most organizations. As long as data remains highly fragmented and not in standard formats, or data is not available to the right people and agents, enterprises are dealing with issues around being able to get answers from agents that are accurate or that conform to their business practices. This cuts across both systems with structured data (product metrics or revenue figures) and unstructured data (product roadmap or customer contracts).
* Clear sense that companies need to figure out what their core data moats are going to be in the future. If everyone has access to roughly the same superintelligence from the various models, then the context that you feed the models becomes proprietary value in the future. Capturing this data and getting it into a format that agents can use becomes very important.
* Everyone is trying to figure out the right metrics to manage to for AI adoption. General consensus that tokens are not the right metric per se, and people leaning more toward business outcomes (in an ideal world). For business outcomes (like more revenue or more shipped product), though, you have to get close to each individual workflow to figure out if it was successfully transformed with AI so it’s harder to manage top down.
* Growing view that enterprises are going to live in a multi-model world. Lots of interest (though early in actual adoption) in layers that can route workloads to different models (frontside or open weights) for cost or performance reasons. Also enterprises are trying to figure out what things do you give to the models directly vs. what do you separate as horizontal systems and context so you can swap any system in and out.
* Talent for driving AI adoption and implementation still remains a major issue and topic. Many view it as something you necessarily have to train for internally due to a shortage of talent being trained on this in the outside. As an aside, this feels like it remains a huge opportunity for those that get very good at deploying and management agents in an enterprise since most companies are looking for these skills.
* The best use-cases for AI tend to be those that fundamentally change the work being done instead of just replacing an existing process and doing it more efficiently. Companies are working through their versions of this individually because it’s different per industry, but this often remains both the most exciting and higher upside uses of AI.
Many more topics discussed recently, but overall it’s clear that there’s a ton of change going on with much more to come.
Usual SecOpsy habits while getting ready to work with NVIDIA NIM api inference endpoint. Perhaps needs a landing page or a redirect to something prominent?
@deedydas For us, hybrid RAG has been the default for more than a year now - both dense vector search + BM25 to pull the right candidates for further processing/pruning (for relevancy, staleness, cost, etc)
Same here. claude sonnet was the dependable workhorse last many months. Now, regularly swapping with GPT-5 with a lean towards the latter. Primary issue with sonnet, it is too agreeable (even with explicit hints in cursor-rules). gpt-5 is bolder, gives dissenting note to steer design/code in another direction which I appreciate.
+1 for "context engineering" over "prompt engineering".
People associate prompts with short task descriptions you'd give an LLM in your day-to-day use. When in every industrial-strength LLM app, context engineering is the delicate art and science of filling the context window with just the right information for the next step. Science because doing this right involves task descriptions and explanations, few shot examples, RAG, related (possibly multimodal) data, tools, state and history, compacting... Too little or of the wrong form and the LLM doesn't have the right context for optimal performance. Too much or too irrelevant and the LLM costs might go up and performance might come down. Doing this well is highly non-trivial. And art because of the guiding intuition around LLM psychology of people spirits.
On top of context engineering itself, an LLM app has to:
- break up problems just right into control flows
- pack the context windows just right
- dispatch calls to LLMs of the right kind and capability
- handle generation-verification UIUX flows
- a lot more - guardrails, security, evals, parallelism, prefetching, ...
So context engineering is just one small piece of an emerging thick layer of non-trivial software that coordinates individual LLM calls (and a lot more) into full LLM apps. The term "ChatGPT wrapper" is tired and really, really wrong.
As we hit limits of existing LLMs and LRMs, new model architectures will emerge!
Remains to be seen who will be winners & losers of the next AI wave
The Illusion of Thinking in LLMs
Apple researchers discuss the strengths and limitations of reasoning models.
Apparently, reasoning models "collapse" beyond certain task complexities.
Lots of important insights on this one. (bookmark it!)
Here are my notes:
For Enterprises, this is the next 5 years worth of feature requirements in one tweet. That's an optimistic estimate :)
Both people and enterprise business processes need to co-evolve for this generational change in how human <-> software interacts.
Products with extensive/rich UIs lots of sliders, switches, menus, with no scripting support, and built on opaque, custom, binary formats are ngmi in the era of heavy human+AI collaboration.
If an LLM can't read the underlying representations and manipulate them and all of the related settings via scripting, then it also can't co-pilot your product with existing professionals and it doesn't allow vibe coding for the 100X more aspiring prosumers.
Example high risk (binary objects/artifacts, no text DSL): every Adobe product, DAWs, CAD/3D
Example medium-high risk (already partially text scriptable): Blender, Unity
Example medium-low risk (mostly but not entirely text already, some automation/plugins ecosystem): Excel
Example low risk (already just all text, lucky!): IDEs like VS Code, Figma, Jupyter, Obsidian, ...
AIs will get better and better at human UIUX (Operator and friends), but I suspect the products that attempt to exclusively wait for this future without trying to meet the technology halfway where it is today are not going to have a good time.
Huh. Looks like Plato was right.
A new paper shows all language models converge on the same "universal geometry" of meaning. Researchers can translate between ANY model's embeddings without seeing the original text.
Implications for philosophy and vector databases alike.
Microsoft continues to embrace open standards - Empowering multi-agent apps with the Agent2Agent (A2A) protocol proposed by Google
#AIAgent#MCP#A2A#Microsoft
https://t.co/7Yzs7fkx0Y
An iPhone moment might be on the horizon for AI with this announcement of OpenAI and Lovefrom/IO coming together with legendary Jony Ive and @sama ... while the forgettable Humane AI pin came & went, waiting to see what unfolds from this legendary combo
https://t.co/ymnd6ibDZb
Agentic memory that survives and accessible across multiple MCP enabled apps is powerful. Need to check if any scope or auth check available to guard the gates - for privacy and security guardrails.
We’re excited to launch OpenMemory MCP, a private memory for MCP-compatible clients powered by @mem0ai
Today, most AI assistants and dev tools operate without memory. You plan your roadmap in Claude, implement tasks in Cursor, but none of them know what the other did. Each tool operates in isolation, and your context disappears as soon as the session ends.
OpenMemory MCP solves this.
Built on the open Model Context Protocol (MCP), OpenMemory MCP runs 100% locally and provides a persistent, portable memory layer for all your AI tools. It enables agents and assistants to read from and write to a shared memory, securely and privately.
Key capabilities:
✅Works across MCP clients like Cursor, Claude Desktop, Windsurf, Cline and more
✅Provides standardized memory operations (add_memories, search_memory, list_memories, delete_all_memories)
✅Stores data on your machine, fully private
✅Offers a centralized dashboard for visibility and control
✅Simple Docker-based setup with no vendor lock-in
If you're building on MCP, OpenMemory MCP is the easiest way to add persistent, private memory to your clients with zero cloud dependencies and full local control.
We’ve put together a short demo to show how it works
Check it out in the links below 👇
Unpacking the economics of DeepSeek and comparing it with US AI models, say, Claude Sonnet 3.5's build cost. Perhaps it is comparable in cost *but* the kicker ...
DeepSeek is opensource with an MIT license. Even Llama license comes with strings attached
#DeepSeekR1
These four points on DeepSeek seem very likely correct and important to understand about the economics of building AI models and what DeepSeek actually did. .
I'll get straight to the point.
We trained 2 new models. Like BERT, but modern. ModernBERT.
Not some hypey GenAI thing, but a proper workhorse model, for retrieval, classification, etc. Real practical stuff.
It's much faster, more accurate, longer context, and more useful. 🧵
A significant milestone in OpenSource with a frontier AI Model. It might well be the watershed moment à la Unix ruling the OS landscape. It will take few iterations and some significant GPU $$$ by downstream efforts to produce a "Linux" like community maintained frontier model. For now, let's cheer @AIatMeta for a great move 👏
Starting today, open source is leading the way. Introducing Llama 3.1: Our most capable models yet.
Today we’re releasing a collection of new Llama 3.1 models including our long awaited 405B. These models deliver improved reasoning capabilities, a larger 128K token context window and improved support for 8 languages among other improvements. Llama 3.1 405B rivals leading closed source models on state-of-the-art capabilities across a range of tasks in general knowledge, steerability, math, tool use and multilingual translation.
The models are available to download now directly from Meta or @huggingface. With today’s release the ecosystem is also ready to go with 25+ partners rolling out our latest models — including @awscloud, @nvidia, @databricks, @groqinc, @dell, @azure and @googlecloud ready on day one.
More details in the full announcement ➡️ https://t.co/hhJoLm5eLV
Download Llama 3.1 models ➡️ https://t.co/rRjvmxqCTC
With these releases we’re setting the stage for unprecedented new opportunities and we can’t wait to see the innovation our newest models will unlock across all levels of the AI community.
This is where we are heading. Some portion of our computing needs can switch to use such Cognitive Processing Systems. We need some initial killer use cases to light up this transition.
I still believe Apple with its ability to hear (microphone), see (camera) and show (display) using it devices and it’s UX mastery will be the first to crack this!
100% Fully Software 2.0 computer. Just a single neural net and no classical software at all. Device inputs (audio video, touch etc) directly feed into a neural net, the outputs of it directly display as audio/video on speaker/screen, that’s it.