Taking a short break from LLM research to go back to my roots.
Introducing... Concept Lenses! A training-free way to do similarity search along whatever concept you care about.
Typically, embedding-based similarity search is constrained by whatever notion of similarity the model learned in training. This means that if the model is primed for semantic similarity, but you actually care more about visual appearance, you'd need to switch models.
Under a Concept Lens projection, though, distance between embeddings means distance only along the concept you care about.
Blog post: https://t.co/aO36RaATpT
Demo: https://t.co/cM9ySdHvW4
Sharing my first of hopefully many research blog posts! https://t.co/qPvLWmsghr
This one is the kind of educational blog post I wish I'd had when I started with RL for LLMs.
I tried to make it as open as possible. Every rollout is browsable, the code is open source, and I walk through my entire thought process, from learning rate sweeps to reward shaping.
We gave our AI system a spec, and in under 2 weeks, it designed, verified, and deployed a chip that beats NVIDIA.
It’s built for low-power physical AI workloads. We’re running live inference on >B+ parameter models like Llama, Qwen, and Kimi, serving at 3.4x better perf/watt than NVIDIA Jetson.
From just a specification, our AI system autonomously generated all of the RTL Design, UVM verification, formal proofs, firmware, drivers, and kernels, co-designing the model, software and silicon as one optimization loop.
Better AI can now design better chips to run AI, leading to a loop of recursive-self improvement towards our path to abundant intelligence.
Code mode lost by 39 seconds.
It is supposed to be the faster way for agents to call MCP tools. Fewer round trips, smaller context.
We raced it against classic tool calling on the same Linear task and it lost. The reason isn't a code mode problem.
Most MCP servers ship input schemas and nothing else, so the agent knows exactly what to send and has to guess what comes back.
It guessed wrong three times, and every wrong guess is a full extra turn.
So we built a shape cache that remembers what each tool returns and carries it across sessions.
Rematch: 27% faster than classic, 47% cheaper, 42% less context.
@Liftyd on making [code]smith fast: https://t.co/J1XDbZhE2v
the funny thing to me about RLMs is the original paper SPECIFICALLY shows that they are "better" on long context tasks, but the abstraction is so nice and programmer-ey that everyone (including me) wants to believe its better for every other task too
@chipro isn't this close to what batch APIs end up solving? except instead of variable prices you have just two buckets of pricing depending on how slow you're willing to get your results. I'd imagine a more continuous pricing market has many operational and contractual complexities
@amystweets huh.. interesting, because my takeaway from all the robert caro biographies is that moses and johnson loved seizing power and ran things as dictatorially as they could get away with
@realmcore_ mostly in the context of building agents in product rather than a terminal harness, but you have to figure out
1. how your product integrates with the harness (async exec, lifecycle mgmt)
2. what tools, skills, clis, mcps work (harness design, evals)
... and it keeps going
Excited to share what I've been working on for the past few months! We have an amazing team and are doing really interesting work at the intersection of frontier AI and chip design.
We are proud to announce @architectlabs.
Our mission is to build an AI system that designs and provably verifies chips end-to-end. By doing so, we unlock a new era of purpose-built chips generated on demand, that powers the scale and distribution of intelligence impossible with current hardware paradigms.
Our founding team collectively has taped out 80+ production chips, led $10B datacenter product lines, been core contributors to Meta’s AI silicon, architected and designed the first neuromorphic chip out of Intel AI Lab, led research teams at Anthropic, xAI, and Google DeepMind, and contributed to fundamental AI research across nearly every frontier lab.
We’re fortunate to be backed by investors who share our vision, including @stevejang from @KindredVentures who led our $24M seed round as well as @TQVentures , @RaceCapital , @scaletogether , @ora , and Link Ventures.
We are also grateful for the support of our angels and advisors, including @snsf, @lukaszkaiser, @AravSrinivas, Kunle Olukotun, @tlbtlbtlb, @alexwg, Siddharth Nath, Thierry Tambe, @arashf, @ekaurghar, @CHHubbell, Selene Casabal, @semiDL and engineering leaders from OpenAI, NVIDIA, Google DeepMind, Intel and more.
We’ve come together to scale up and reinvent how chips are designed and provably verified end-to-end. If you want to work at the intersection of frontier AI, systems and silicon, consider joining us.
@mitchellh reminds me a lot of John Hopfields "Now what" that I read earlier in the week where he tries to distinguish between a problem and A PROBLEM to work on as a scientist
https://t.co/X6Ixdm3O9o
you can scrape any authenticated page with claude code in like 2 mins…
open the site in your browser, log in normally
open dev tools → network tab → refresh
find the first request that loads the content you want
right click → copy as curl
this copies your auth cookies with it. paste that curl in terminal and you'll get the same content outside the browser
now just give claude code that curl command and tell it to fetch from that endpoint. it'll handle the rest
works on basically anything - dashboards, gated content, private apis
the curl command is doing all the heavy lifting bc it contains your session cookies. you're basically cloning your authenticated state
only caveat is tokens expire so you might need to recopy if sessions are short
you are welcome