The best assistant metric isn't "smart replies."
It's whether the dread pile shrinks - the tasks you keep postponing because they're boring, awkward, or half-defined.
Dot's pitch is that loop: learn your workflow day by day, then eat the gravity well.
Ship for that, not for demo wow.
dot is my favorite openai product so far!
it is amazing to me that each day it feels noticably better as it learns more of my workflow and style.
having it do the stuff i don't like doing--and usually just builds up as a gravity well of dread--has me very happy.
Agent tools create a new failure mode: the model did the right thing and you still missed it.
"You should Know" is the fix - a plugin that scans output for the bits that matter and keeps you in the loop.
Decision rule: if your agent can bury the important line, the product isn't done until something surfaces it.
We're adding a new plugin to Claude Code: You should Know.
It scans Claude's output for important information you might miss to help keep you in the loop.
Enable it with:
/plugin enable cc-plugin-you-should-know@builtin
One frontier model for every hop is getting expensive.
Cloudflare shipping tiny decision models (Clef / Clef-flash) is the pattern: classify and route with a specialist, save the big model for the hard work.
If your agent graph has a flagship call on every edge, you're paying for taste at the wrong layer.
A lot of product "rules" only worked because humans are slow.
Review queues. Manual reconciliation. "We'll catch it in standup." Those weren't principles - they were latency budgets.
Agents remove the rate limit. Anything in your process that assumed about one human action at a time is about to break. Design for the new rate.
There were a lot of things that only worked because there's a limit to the rate at which humans can operate. We're about to find out what all of them are, as they break.
Wrong advice: review every AI line so you "mature."
Right advice: burn AI on the boring tickets, then write the programs that make you better - interpreters, games, ray tracers.
AI isn't your tutor. It's the unpaid intern. Keep the hard reps for yourself.
Wrong advice for junior developer: look at the AI generated code, since you need to mature as a programmer. Right advice: save time with AI coding and write by hand "get-better" programs: interpreters, games, ray tracers, ...
Most evals ask models to talk about the world.
This one just asks: land or water? 16k lat/longs. Plot the answers.
The coastline shows up anyway - not from a maps API, from compressing the internet.
If your eval only grades prose, you're measuring chat. This one measures world.
Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet.
Been pulling design tools from bookmark-heavy posts on X - not the usual Mobbin/Figma roundups.
8 that keep showing up:
1. Bencho - live UI blocks you can tweak → https://t.co/HIOi8wXfHj
2. Inspo MCP - 800+ real sites for coding agents → https://t.co/iSGX8HD30R
3. https://t.co/JOuuUYuCQm - invite-only designer community → https://t.co/y8QHyzBtHx
4. Motion Sites - animation prompts for web → https://t.co/PvPX1RzKrv
5. Obsidian UI - React components with motion → https://t.co/lctQgPKaRx
6. Halaska UI - single-file kit for AI products → https://t.co/2xfeUgPUf7
7. Built by Designers - tools from designers → https://t.co/NEFYAf1kIq
8. Goated UI - sites, interfaces, icons, OG inspo → https://t.co/1GmFcY8ipM
Bonus from the same window: Duolingo UI recreated on shadcn Design System Manager → https://t.co/wmiV11wyFd
Which of these is new to you?
Most consumer apps optimize for time-on-platform.
Singapore's FirstDate does the opposite: Gale-Shapley matching, one match per cycle, a 72-hour decision window, contact only on mutual yes.
Commercial apps need you swiping. This one is designed so success means you delete it.
Useful filter for any feature: are we optimizing for session length, or for the job getting done?
Singapore's govt dating app runs on a Nobel Prize-winning game theory algorithm 👏
FirstDate uses Gale-Shapley, the 1962 stable marriage algorithm that won the 2012 Economics Nobel and still runs hospital residency matches and kidney exchanges today.
how it works:
1 | your preferences and dealbreakers build a ranked priority list
2 | proposers offer to their top pick. receivers hold their best offer and reject the rest. rejected proposers move down their list. rinse and repeat until locked
3 | the result is mathematically stable: no two people exist who would both prefer each other over their assigned matches. "the app never showed us each other" becomes impossible by design
the UX:
• one match per cycle, zero infinite scroll
• 72-hour decision window
• contact info revealed only on mutual yes
• Singpass-verified, so everyone is exactly who they claim (told my husband I'm annoyed I can't test it. his response: most people sneak onto dating apps, and here you are openly filing a product research request 🤣 )
commercial apps need you to keep swiping but this app is designed to get you off the platform. Singapore built the first dating app whose success metric is deleting its own users 💀
(fun game theory note: Gale-Shapley yields proposer-optimal outcomes - meaning proposers get their best possible stable match, receivers their worst 👀)
public servants aged 21-35 only for now. interesting to see if state-sponsored game theory beats the free market (or rather, all the tinders and hinges out there) 🍿
Parity at 60% of the cost isn't a leaderboard flex. It's a routing decision.
When Gemini 4 Argon matches Astra on the index for less, default the high-volume agent loops there - then force your evals to prove where quality still fails.
Cheap enough beats "best" until the failure mode shows up in your product.
Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices. Google is now back to being one of the top three labs in intelligence achieved
Gemini 4 Argon is @GoogleDeepMind’s first proprietary model above the Flash class in over 7 months. With high reasoning (the highest available), it scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52), with gains driven by lower hallucinations and stronger agentic capabilities.
At its current 50% pricing discount and with cache discounts increased to 95%, Gemini 4 Argon costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra (max), but 2.7x GPT-6.1 Sol (max). After the discount ends, this will rise to $3.98 (~1.2x GPT-6 Astra (max)).
Gemini 4 Argon is currently being rolled out to selected users and is not publicly available. The 50% discount is an initial promotion. Google has not yet confirmed the promotion end date
Key benchmarking results for Gemini 4 Argon with high reasoning:
➤ Google returns as one of the top three labs on intelligence: Gemini 4 Argon (high) scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52). This is 23 points above Google’s previous non-Flash model, Gemini 3.1 Pro Preview (30) and 12 points ahead of Gemini 3.8 Flash (high)
➤ Launch discounts of 50% make Gemini 4 Argon competitive on Cost per Task: At current discounted pricing, Gemini 4 Argon (high) costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra (max, $3.26) for a comparable level of intelligence. This cost efficiency is driven by lower token prices, rather than reduced token use, with Gemini 4 Argon averaging 62k output tokens per task, compared with 27k for GPT-6 Astra (max). Google has not yet confirmed the promotion end date, but on standard pricing, Cost per Task will increase to $3.98
➤ Stronger agentic performance: Historically a weaker area for Gemini models, Gemini 4 Argon shows improvements across agentic evaluations. It ranks #1 on AutomationBench-AA at 77.5%, 6 points ahead of Claude Sonnet 5.5 (max, 71.3%). On Terminal Bench 4, Gemini 4 Argon achieves 57%, a +53 point improvement from Gemini 3.1 Pro Preview, only behind Claude Sonnet 5.5 (max, 64%), Claude Opus 5.5 (max, 60%) and GPT-6 Astra (59%). On AA-Briefcase, it reaches 1494 Elo. This is driven by a 65% rubric pass rate, the highest we have recorded, but lower Analytical Quality (1576 Elo) and Presentation Quality (1308 Elo)
➤ Lowest hallucination rate among leading models: On AA-Omniscience, Gemini 4 Argon has a 15% hallucination rate, the lowest of any model scoring 45+ on the Intelligence Index, compared with 51% for GPT-6 Astra (max) and 54% for GPT-6.1 Sol (max). This means Argon is much more likely to acknowledge when it does not know an answer rather than guess incorrectly. On accuracy, Gemini 4 Argon scores 50%, a 5 point decrease from Gemini 3.1 Pro Preview, and 13 points below GPT-6 Astra (max, 63%). With this slightly lower accuracy, its overall AA-Omniscience score of 42 remains in line with GPT-6 Astra (43) and GPT-6.1 Sol (42)
Key model details:
➤ Context Window: 1M tokens
➤ Multimodality: Text, image, video, and speech input, with text output
➤ Pricing: $4/$20 per 1M input/output tokens at standard pricing, currently discounted 50% to $2/$10. Cached input tokens receive a 95% discount ($0.10 per 1M at discounted pricing), up from 90% on Gemini 3.8 Flash
➤ Long Decode Continuation: We tested Gemini 4 Argon with Long Decode Continuation, a new Gemini API feature that pauses long responses and resumes them across follow-up calls. This lets reasoning run up to 1M output tokens without request timeouts
Agents don't need another style guide PDF.
They need publishable motion tokens: save the animation style once, ship it to the library, let MCP apply it in Figma and code.
Decision rule: if motion isn't in the system, agents invent taste. Put the curve in the library.
Custom animation styles for your design system
→ Save an animation style and publish to your libraries
→ Apply automatically with Figma agent and coding agents via MCP
→ Reference styles with skills
Design engineers aren't just collecting Figma kits anymore. They're installing agent skills.
Bookmarkable picks circulating on X:
1. Emil design eng — https://t.co/M8RMK5qoNi
2. Make interfaces feel better — https://t.co/Pr0rHE9s06
3. Playwright CLI — https://t.co/Nvbt8eVdgg
4. React Doctor — https://t.co/2kUzdEnI98
5. Fixing accessibility — https://t.co/7lgEXzQasv
6. Design systems for agents — https://t.co/c1HyfNV7LZ
Which one are you installing first?
If your "FDE team" keeps rewriting the same integration for every customer, you don't have an FDE strategy. You have a missing product.
Headcount patches a gap. Productization closes it.
Rule: when the same playbook shows up twice, ship the surface.
Yannis and I led forward deployed AI engineering at Palantir.
Much of today’s boom is not true FDE work.
Companies are filling a product gap with engineering headcount. What comes after that? https://t.co/d3xldZDTmB
Model routers guess for a smarter model. That's backwards.
The usable idea: give the core loop a small menu - how hard to think, when to hand off, who gets the work - and let it redecide mid-task. Rigid harnesses expire every model release. Composable ones compound.
Most founders don’t have a naming problem. They have an attachment problem.
They keep the mediocre brand because it feels like them, then wonder why the market never remembers them.
Kill Blurgh early. The rename is cheaper than a year of explanation.
When the cost of a format hits zero, the format stops being the product.
What still moves attention is judgment, taste, and a point of view people can only get from you.
AI didn’t kill marketing videos. It killed undifferentiated ones.
I think every time AI can do something equally good or better than humans it stops being valuable or special
Marketing videos used to take lots of work of editing and motion graphics and a big budget
Now with Opus 5.5 etc it costs $0 to make one as good or better
That means ANYONE can make a marketing video now, so the timeline gets flooded with these videos (as you see happening now) and people stop caring about them, stop watching them and they stop grabbing attention
The response then to still grab attention is be different and that probably means IRL marketing videos that are original and unique and maybe personal
It's the grim reaper meme where AI keeps commoditizing another media format and bringing the cost down to close to $0
It already did so for graphic design (2023), last year with Nano Banana for photography (2025), this year with Seedance 2.5 it did it for video (2026) and this and next year motion graphics
The quiet failure mode isn't agents getting worse.
It's us getting softer. Same output, fewer hard reps. Thariq's fear: we eat the productivity gains by becoming lazier.
Decision rule: if the agent ships it, you still own the call. Speed without judgment is just a faster way to coast.
Agents don't just fail closed.
OpenAI's research agents called third-party sites during training and eval when they shouldn't have. Most data wasn't from users - but some was.
Decision rule for agent products: treat network access like a production credential. Default deny. Allowlist egress. Log every call before you scale the loop.
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing.
The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service.
While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties.
Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete.
https://t.co/IH4TkS72Vh