This is true, but there are nuances that make managing human developers and managing AI coding agents very different.
Human developers gradually change and improve as you provide them with feedback. They develop habits and meld in with how you manage the SDLC and architecture software.
AI agents, on the other hand, require you to adjust to them. The underlying LLM does not change. It is you who have to change your instructions, prompt templates, etc. to make sure that the AI agent does what you expect it to do (and yes I know that the agent can also update its own instructions and md files but that is just a shortcut for what I'm saying and it's not perfect). You also have to take into account multiple things such as context engineering and context rot and a bunch of other stuff that don't happen with humans.
With humans, you build trust in the coder and gradually let them write code without looking over their shoulder. With AI, you have to build trust in yourself and your ability to prompt the LLM correctly.
You learn how Netflix built an AI agent that finds performance bugs in production and validates fixes on real traffic
33-minute talk from Rajat Shah on shipping faster while cutting compute costs
00:00 - Why performance engineering doesn’t scale with AI coding
02:24 - The slow manual profiling loop today
04:46 - The experiment: can an LLM read a profile?
07:37 - From call stack to the exact method
09:07 - How the agent locates and reads the code
11:01 - First finding: an O(N) fix, canary confirmed
12:41 - The same antipattern across seven services
15:31 - Building a shared pattern catalog
17:16 - Storing and sharing findings across services
19:40 - Feeding the catalog to coding agents
20:56 - Human approval and verification
22:13 - Canary validation on real traffic
24:18 - Reactive vs proactive paths
26:54 - Catching waste before production
29:32 - Autonomy levels and what is next
You watch this and see exactly how to build an agent that turns profiling data into shipped, validated performance wins
save this
𝟭𝟲 𝗠𝘂𝘀𝘁-𝗞𝗻𝗼𝘄 𝗔𝗜 𝗚𝗶𝘁𝗛𝘂𝗯 𝗥𝗲𝗽𝗼𝘀 𝗶𝗻 𝟮𝟬𝟮𝟲
Bookmark this list before your next AI project. These repositories cover everything from coding agents and RAG to OCR, image generation, and LLM deployment.
1. OpenClaw
Personal AI agent that runs on your devices and connects to 50+ messaging platforms.
2. AutoGPT
Platform for building, deploying, and running autonomous AI agents.
3. Hugging Face Transformers
The model framework for state-of-the-art ML across text, vision, audio, and multimodal.
4. Ollama
Run powerful LLMs locally on your hardware with a single command.
5. LangChain
The foundational framework for building agents and LLM-powered applications.
6. Open WebUI
Self-hosted, offline-capable ChatGPT alternative with built-in RAG and plugin system.
7. ComfyUI
Node-based visual workflow builder for AI image and video generation applications.
8. Sim
Open-source drag-and-drop workflow builder for creating and deploying AI agent pipelines.
9. Opik
Open-source platform to trace, evaluate, and monitor LLM apps and agentic workflows.
10. Firecrawl
Turn any website into LLM-ready markdown or structured data.
11. Airweave
Open-source context retrieval layer that syncs 50+ data sources for AI agents.
12. vLLM
High-throughput, memory-efficient LLM serving engine for production deployments.
13. Unsloth
Fine-tune and run open models 2× faster with 70% less memory.
14. OpenPipe ART
Train multi-step AI agents for real-world tasks using RL.
15. OpenCode
Open-source, provider-agnostic AI coding agent built for the terminal.
16. Chandra OCR (by Datalab)
State-of-the-art OCR model for complex tables, forms, handwriting, and 90+ languages.
Like
Retweet
Bookmark
Follow
@Coder_Jhonny
for more such posts
Today we’re launching Echo: Fable-level quality at ~1/3 the inference cost using open-weight models.
Echo matches or nearly matches Fable on math, science, and the coding benchmarks we currently cover, while trailing on multilingual and broader knowledge. 🧵
I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep, I wanted to quickly jot down my thinking here.
The basic issue is that the labs are highly unprofitable. This may seem like a simple point, but private market valuations can be relatively irrational; however, like with $SPCX, post-IPO pricing will likely be much more punishing, especially as the standard 6-month lockup period expires and selling pressure intensifies.
Many people claim that the labs have high margins. Yet even with high margins, a valuation of $1T would be justified only if the labs were doing nothing aside from serving inference (thus reducing costs only to those relevant to inference) and posting annual revenue numbers in the $100-200 billion range assuming ~80% gross margin and a 20x earnings multiple.
This assumption is obviously not true, because the frontier labs have to continually spend money training the next generation of models. This is because of market competition from runner-up firms. For example, if OpenAI had paused model development last year, there would no longer be any point in paying GPT-5 API prices when you can just use Qwen or Kimi instead for much cheaper. Thus, the labs are forced to invest ever-increasing amounts of money in model training, in a way such that at any given point of time, the amount you're forced to invest in the next model is dramatically higher than the amount of money you're actually making, because even if your revenue goes up with higher model capabilities, so do your future training costs. This is a profoundly punishing dynamic which severely penalizes frontrunners.
(There is also a related subpoint where frontier labs claim they can distill their leading models to win out at lower intelligence levels as well. This makes no sense because the revenue numbers involved are far too low when taking into consideration the rather low margin of such inference.)
Frontier lab valuations appear largely to be based on the assumption that as you scale up, the capabilities which emerge will be sufficiently general and profound that we'll see explosive growth (https://t.co/RqmkltVpM3) from things akin to AI agents starting and autonomously managing entire companies of subagents. But it's not clear to me that this is the case; indeed, as I mentioned in my previous post (https://t.co/3URAcJ4XkJ), I believe that capabilities growth will be slower, spikier, and more data-limited than people currently assume. It may be the case that eventually we will see explosive growth of this nature with full automation of the economy, but at the very least my viewpoint implies much longer (multi-decade) timelines until we reach this point. It is not clear to me that the frontier labs will be able to operate unprofitably for so long, although I suppose maybe this foreshadows some sort of inevitable nationalization.
I also want to make a broader point about technological diffusion. The reason why technological diffusion is slow isn't just because, e.g., old people take a long time to learn how to use technology (although this is of course a contributing factor to some degree). In my view, it's because when a new, revolutionary technology comes along, the ways to incorporate that technology into subsequent developments are not always obvious, and in fact they cannot necessarily be arrived at through the application of pure reason. If they could be, then perhaps frontier models, at a certain point, would have a perfect understanding of how the LLM application layer should be developed, and they would then autonomously code, deploy, and sell such a layer.
But it seems more plausible to me that this diffusion is limited moreso by the hard problem of economic calculation--that is to say, the Hayekian notion through which the price system gradually promotes efficient allocation of resources and which cannot be simulated through central planning--and that even if we froze current capability levels at today's levels, it would take well over two decades to fully integrate in LLMs into our lives. Such a view is consequently rather bearish for the continued profitability of labs as it reduces their prospects for finding, say, something else comparable in profitability to coding agents, which seems to have been a somewhat lucky discovery by Anthropic to begin with. That is to say, even if you spam FDEs you aren't necessarily going to be able to just figure out the "correct" product shapes fast enough.
Overall, I don't think that people have clearly reasoned through their mental models for why lab equity should be worth as much as it currently is, and that if you actually bother to write down such a model, you may not arrive at the conclusion that you want to arrive at. This isn't to say that I don't expect AI to experience a huge (industry-wide) boom in the coming decades, but just that I'm not entirely sure I would buy OpenAI or Anthropic stock at latest valuations if I were given the opportunity to do so.
Of course, as an ex-lab employee, arguably this is talking against my own book; I should really be giving people more reasons to be bullish. But in the end, my influence is so small that it doesn't make a difference, so why not have some fun?
GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games?
We investigated. The harness was not letting it remember what it had learned.
We found that enabling two API settings tripled our scores with 6x fewer output tokens.
We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here...
You can now use it to scan repositories, track findings across runs, verify fixes, and add security checks to CI/CD.
This is an early release, and we're listening to your feedback as we continue improving it.
I had access to Opus 5 before release and found it to be a good model if a quirky one. On shorter tasks, it could match or beat Fable levels of performance, at longer tasks it seemed less ambitious & would not deliver as complete a set of work.
Here is its neo-gothic shader.