🤖 archify
⭐ 76,933 stars
Stop manually drawing system architecture. Turn any codebase, plan, or complex idea into a beautiful interactive diagram instantly.
🔗 https://t.co/5MJ5UScteP
#AI#MachineLearning
Another insane Jev use case!
Traditional database filters need precise, predefined conditions. But many questions are semantic:
- Is this article mainly about software engineering?
- Which topic best describes it?
- How technically deep does it appear?
These usually require moving rows into application code, invoking a model, parsing its output, and writing the result back.
pg-jev is an open-source Postgres extension that exposes Jev through SQL functions.
- jev() works inside WHERE
- jev_prob() returns a probability
- jev_choice() selects a label
- jev_score() ranks rows across ordered levels.
It needs no vector column or embedding index.
Under the hood, pg-jev batches rows, sends them to Jev, and returns typed answers that SQL can filter, sort, group, and combine with exact predicates.
The recording below runs it against real Hacker News stories stored in Postgres.
It first shows the rows, then asks Jev to find software-engineering stories, classify them by topic, and rank them by technical depth. The final query reports requests, tokens, cost, and cache hits from the database session.
GitHub repo: https://t.co/6OD0nRMgfH
(don't forget to star it ⭐)
Postgres keeps the data and controls the query, while Jev handles the part SQL cannot express as a deterministic condition.
If you want to understand what Jev is doing under the hood, I also wrote a hands-on guide to building a Jev-style model with open models, entirely locally.
Read it below.
my favorite engineering skills for AI atm:
- Agent Skills by Addy Osmani: https://t.co/zioMUOyVSD
- Ponytail https://t.co/45tACElJg1
- Matt Pocock's skills: https://t.co/Rvf2iSBMRE
Matt's skills tend to get a lot of done with less process
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
I’ve been an ML researcher since almost 10 years now
I can build better models now with pure vibe coding than I ever did with all the feature engineering, carefully engineered poses and regularization techniques
Read “the bitter lesson” folks
While the value of software engineering *skill* is declining rapidly, the *mindset* of a software engineer will rapidly rise in value.
Modern businesses can be seen as bundles of software, and the role of humans will increasingly shift to setting up automations that define a business and managing exceptions.
Businesses of the future will require people with an automation-bias more than ever before. It’s just that these people won’t write software themselves.
This has never been more true. The opportunity cost of endless planning and contemplation just went way up. Agents allow you to just try vastly more stuff, so let them, and you'll learn way more, way faster.
Excited to share Quail, our new open source AI-SQL engine (a collab with Modal)! By planning queries and LLM inference together, it reaches 1B+ input tokens/min on one H100 for one query 😱🚀
AI-powered data operators create a new, interesting inference workload👇
NVIDIA and Stanford just challenged Jev.
(their new System 1 architecture runs up to 9x faster.)
It is called a Contrastive Language Model, or CLM.
Like Jev, CLM is not designed to generate text. It handles the small, repeated decisions inside AI systems, such as choosing a tool, ranking a patch, routing a request, or selecting the next action.
But CLM reaches those decisions differently.
Instead of generating an answer token by token, it treats decision-making as a retrieval problem.
Here is how it works.
1) Encode the state
CLM takes the current situation, such as an agent’s context or the state of a game, and converts it into a vector.
It uses a frozen Qwen3-8B model with a small trainable state projection head.
2) Encode every possible action
A separate action head converts each candidate into the same vector space.
In the Mario example, the candidates are left, jump, and right run. CLM does not invent a fourth option. It only evaluates the actions supplied by the application.
3) Learn which states and actions belong together
During training, the correct state-action pair is pulled closer while incorrect pairs are pushed apart.
A batch of B examples produces a B × B similarity matrix. The matching pairs sit on the diagonal. Every other pairing becomes a negative example.
This contrastive training uses InfoNCE, the same general mechanism behind systems such as CLIP and dense retrieval.
4) Turn similarity into a decision
At inference, CLM measures the cosine similarity between the state and every candidate action.
A softmax converts those scores into a probability distribution. The application can choose the winner, apply a confidence threshold, or escalate an uncertain result.
The real speed advantage comes from separating states and actions.
Actions can be embedded once and cached. If an agent repeatedly chooses between the same tools, CLM only needs to encode the changing state and compare it with stored action vectors.
That replaces repeated generation with one embedding pass and a set of cheap dot products.
The researchers report that CLM-8B matches Jev across computer-use, gaming, and tool-calling evaluations while reaching up to 9x lower latency. The improvement is largest when actions repeat or the candidate set grows.
CLM still has limits. It cannot generate new actions, its probabilities are relative to the supplied candidates, and its strongest verifier results require task-specific fine-tuning.
But its central idea is powerful.
The entire research is open-source, including the code.
Read more here: https://t.co/I9kPwMPI7B
When software already knows the possible answers, an AI model should score them instead of generating more words.
I also wrote a full breakdown on how system one models like Jev work.
The article is quoted below.
Minko Gechev was preparing for a Google interview
Instead of keeping his prep notes to himself, he turned them into a GitHub repo
He collected the problems he actually practiced and organized them around the concepts he was preparing for
No “500 problems in 30 days”
No random problem dump
Just practical interview prep
He also openly points out that solving these problems isn't enough you need strong fundamentals and the ability to explain your thinking during the interview
That repo now has 3K+ stars. 👀
https://t.co/fjlyo7Se4b
You can now generate an entire 3blue1brown style video from any research paper with Opus 5.5.
Here’s a 8min video summary of “Regularized Recursive Self Improvement of Agent Harnesses”.
The 90%ile educational YouTuber is fully automated.
It's pencils down, people. Writing code by hand is no longer an economically viable skill for most programmers at most companies. But the future of making software has never been brighter. Don't you dare black pill this beautiful moment! https://t.co/qaCgfmkW0u