the assumtion that claude in chrome and claude code's browser tools are the same thing.
claude in chrome connects to your actual logged-in browser, cookies and sessions included. claude code's agent browser spins up a separate, scriptable instance for automation and testing. different job, easy to conflate.
worth checking which one your task actually needs before reaching for the wrong tool.
deleted a flaky selector-based test suite this week after switching to agent browser's snapshot-and-ref model.
instead of hardcoded css selectors that break the moment a designer touches the layout, claude takes a snapshot, reads the actual page structure, and clicks by reference, click @e12 instead of a brittle class name three levels deep.
worth checking if your test suite is fighting the same brittleness before assuming more selectors is the fix.
worth separating the two proposals in the essay on their own merits, the third-party evaluator commitment (giving outside auditors permanent system access) is a concrete, immediately verifiable step that doesn't depend on anyone else moving first. the "pace the frontier while staying ahead" framing is the part that requires other labs to trust the ask, and self-interest language up front makes that trust harder to build regardless of whether the underlying safety case is sound
"cooking on designs" is doing a lot of work for what's still a rendered mockup, not a functioning interface, the gap between "generates something that looks like good ui" and "generates something a dev could actually ship" is still the real test, and screenshots don't show which side of that line this lands on
deleted the assumtion that claude code needs an interactive terminal session every time.
headless mode runs it non-interactively, scriptable, good for ci pipelines or scheduled tasks instead of a human babysitting a live session.
worth checking if a task you're running manually every day could just be a scheduled headless call instead.
deleted a chunk of our team's manual copilot chat routine after actually reading into copilot cli's agentic workflow features.
it handles multi-step tasks, approvals, and file editing directly from the terminal now, not just single-shot suggestions in an editor pane.
worth checking your terminal-based options before assuming every ai coding task needs an editor open.
@ChatGPT the custom domain update is the one that changes the game for me.
βi built this with ChatGPTβ hits different when the result lives on your own domain.
@OpenAIDevs@Yelp restaurant calls are a good stress test for voice ai. names, dates, party sizes, preferences, changes halfway through the call. thereβs a lot more happening than just answering questions.
@RoundtableSpace the wallet is the part that gets my attention.
once an agent can remember you, talk to you and spend money for you, the interface starts becoming way less important.
15 claude settings is a lot, but the first one is probably the biggest upgrade: pick the model based on the job instead of using the same one for everything.
If I had to start over with a (fresh) Claude account.
These are the 15 settings I'd switch on right away.
β Switch 1: the dropdown of the Claude chat box
Fable 5.1 for all my complex tasks that matter.
Opus 5. for ambitious work without Fable tokens.
Sonnet 5 for everyday. Haiku 4.5 for quick answers.
Now look right next to it. Effort. Set it to High.
Effort is how hard Claude thinks before answering. Almost nobody knows the dial is there.
β Switch 2: Skill that works like prompt shortcuts
A skill is a long prompt turned into a command that runs on its own, every time.
Install /skill-creator first by Anthropic.
Then talk to it: "I need a skill that turns my messy notes into a client email in my voice."
It asks questions. You answer. It saves.
For my personal favourite Claude Skills.
Go to https://t.co/6cHYYfjXEA.
Download the skills.
Put your email and verify with an OTP.
Save the downloaded zip file.
Upload to Claude β Settings β Skills.
β Switch 3: the four Skills I'd re-do same day
/anti-ai β rewrites until the detector calls it human. Built after $34 of breaking AI detectors.
/how-to β walks you A to Z through any task inside Claude. I built it to replace myself.
/fable-5 β your messy brain dump in, a Fable-optimised prompt out.
/hands-off β compresses a whole chat into a clean doc, so a new chat resumes with nothing lost.
Building your own takes weeks. So I handed mine over at https://t.co/6cHYYfjXEA
β Switch 4: four connectors. Not 40.
Claude Settings β Connectors.
β¦ Granola β meeting transcript, inside Claude. Don't ask for the summary. Ask "what did I commit to in the last five calls, and what's still open."
β¦ Gmail + Slack β turn both on together. One is what was promised in writing. The other is what was actually said in the thread.
β¦ Google Drive β one button turns Claude's output into a real Sheet or Doc. No downloads folder.
β¦ Gamma β a prompt to finished (with taste) slides.
Whatever you can open at work, Claude can open. Whatever you can't, it can't.
β Switch 5: the one app that isn't inside Claude
Wispr Flow .ai (free). You speak your prompt instead of typing it. Nobody types 3 paragraphs of context. Everybody says three paragraphs out loud.
You speak 4x faster than you type.
β Switch 6: skip this one if you're new
Role Plugins. Sales, Marketing, Finance, Ops.
Each is a ready-made bundle of skills. Useful once you know what you want. Confusing before that.
I've rebuilt this setup twice. Both times in 7 mins.
If you do one thing before closing this app:
Set Claude to Fable 5, effort on High.
Send this to whoever is still running on defaults.
Get my full Claude Skill library, free: https://t.co/6cHYYfjXEA
@Polymarket maybe kids don't need an AI tutor for every assignment. some things are still better learned with a book, a pencil and enough time to struggle through the problem.
deleted my assumption that claude fable and claude mythos are just marketing names for the same release.
they share the same underlying model, anthropic's own announcement confirms it, but fable ships with additional safety measures specifically around biology, cybersecurity, and llm r&d. same capability, different constraints depending on which one you're actually using.
worth checking which one you're on before assuming a limitation is a bug, it might be the safety layer doing exactly what it was built to do.
didn't switch anything over yet. did stop assuming "mythos-tier" meant one single product.
@ClaudeDevs being able to pop panes into separate windows sounds small until you're working on a real project. less tab juggling, more room for the stuff you're actually watching.
the AI infrastructure race is starting to look like a supply chain race. having the best chip means little if you can't get the packaging, memory, power and connectors too.
Jensen Huang says $NVDA is still constrained across wafers, packaging, DRAM, connectors and power components saying βeverything is challenging.β
AI demand is now putting pressure across almost every layer of the hardware stack at the same time.
Jensen Huang says $NVDA is still constrained across wafers, packaging, DRAM, connectors and power components saying βeverything is challenging.β
AI demand is now putting pressure across almost every layer of the hardware stack at the same time.
@PopBase if the actors genuinely feel connected to the characters, historical dramas hit way harder. you stop watching a story and start seeing people.