pivoting to hardware
jk, just a fun pet project where i try to build gravity-powered logic gates using only marbles, foam, Elmer's glue, and a kitchen knife. This is XOR. getting the (1, 1) -> 0 case to work reliably took way longer than I'd like to admit (in retrospect, a jam mechanism is not the way to go).
This is a bit of a strawman. Design quality in the typical sense is not a pure function of novelty/freshness/idiosyncratic style. It has an objective, evergreen component. This component has substantial mass. And LLMs are not even close to mastering it. That is where work is to be done. The existence of cyclical trends/fads don't have bearing on the value of this work.
@PhysInHistory Oddly enough, however, he was a middling chess player, despite his alien-like computational prowess. Whether that says more about von Neumann or the game of chess though is up to debate.
Doesn't matter if it's code or canvas, Codex/Claude output "slop" either way.
Design is a model problem, not an app layer problem. And for whatever reason, the frontier labs either aren't able (or just don't care enough) to solve it. Essentially zero progress on design ability has been made from GPT 5.1 through 5.5, and same with Opus 4.5 through 4.7.
1) There are more possible prompts just 60 tokens in length than atoms in the observable universe, after filtering for nonsensical, malformed constructions. And it's not even close.
2) LLMs can, to the best of our knowledge, provide plausible answers to most of them.
3) Clearly, strong generalization is happening. A few empirical examples of LLMs regurgitating doesn't imply memorization across the distribution in any meaningful sense. Even if what you say is not entirely baseless, it is phrased in a misleading way.
4) As an aside, your comment on statistical approximators being limited in their power I think underestimates the potency of emergence. My favorite intuition-building instance is that the definition of a prime number (and its supporting axioms) is simple enough for any 7yo to understand, but the pattern it forms (the distribution of primes) is so unfathomably complex that no mathematician has been able to precisely describe it.
disagree. there are structural headwinds inhibiting how good a general frontier model (i.e., "one set of weights") can get at everything. consider the prevalence of MoE, no-free-lunch theorem, etc.
additionally, even for the frontier models, most of the marginal improvement is now coming from post-training, not pre-training.
I think it follows that a robust post-train on top of a strong open-weight model for a narrow (but valuable) task set is definitely a moat. Cursor is already heading in this direction.
Excited to announce Gamut: a vision-based preference model for UI / graphic design.
Its telos is simple: given two variants, pick the 'better' one.
If you’ve ever tried reward modeling, though, you know the task's simplicity is an illusion; unverifiable reward functions such as these are notoriously hard to approximate. This is presumably why AI coding progress appears to lag substantially on frontend tasks compared to backend, where a set of unit tests is often enough signal.
Gamut scored near-ceiling (98%) on our internal, held-out evaluation set, where frontier LLMs (such as GPT-5.5 and Claude Opus 4.7) were unable to break 65%. To put that into perspective: 50% is your expected accuracy if you just guess randomly.
It's also extremely parameter-efficient (<90M total), meaning you can infer hundreds of pairwise comparisons in ~1s, contingent on your hardware.
Achieving these results required a lot of iteration on both model architecture and data. The training corpus was built and labeled entirely in-house!
Of course, don't just take my word: if you're a research lab or applied AI co. with a use case, I'd love to have you try it on your own eval set.
Learn more: https://t.co/r8ECtfDCkc
Try it on your own pairs: https://t.co/vs0PZ72yeK
Just gonna say it: dynamic UI is a solution looking for a problem.
I get the appeal from a builder's perspective. It challenges traditional assumptions and feels like an exciting new frontier to explore. But I promise no one will actually want to use your JiT email client over just regular old gmail, even if they think it's cool.
Even if the paper's stance is correct in principle, a sufficiently accurate simulation of emotional language is, in effect, emotion itself.
This is the core philosophical tension between the AI-is-a-bubble and AI-is-not-a-bubble factions. Even if AI is always just an imitation of some (supposedly) real, more pure notion, at some point, when the fidelity of that imitation crosses a certain threshold, you really have to ask: is the imitation vs real distinction even meaningful?
I don't want another GUI abstraction over html/css. I don't want to "co-design" with AI.
I want AI to extract the constraints from my brain and effortlessly one-shot perfection. I want it to collapse the hill-climb towards the right set of decisions -- which presently takes many hours, across many days -- to minutes. I want AI to design better than any human on the planet.
I don't care if the median shifts upward. I want the ceiling to expand.
meta's new ai is surprisingly good... at blatantly contradicting itself. not across a long, 1,000-tok response. literally in the next sentence. on simple discrete math.
It is entirely possible for a statement to be false and still function as good advice.
Good advice is good simply if believing it reliably pushes you towards a better state of being (whatever that means for you). That's it. No precise accordance with reality required.
Consider @pmarca 's takes on introspection. The correctness of his stance is debatable, to say the least. But for profiles that are highly susceptible to getting nerd-sniped / stuck in analysis-paralysis, it's pragmatically excellent advice.