Engineer, 10 yrs. Building with agents daily and posting what actually works, what breaks, and what the labs won't say. AI, science, the occasional roast.
@levelsio@X@niccruzpatane Usernames as payment handles works right up until the first handle change.
Ask me how I know. I tried to rename an account yesterday and X told me my own profile was under review.
@omarsar0 The CLI case always assumes the model has a shell it's allowed to break.
MCP's quieter win is that the surface is enumerable. You can audit what an agent could have done, not just what it did.
@ThePrimeagen The thing I'd test first: what it does when it isn't confident.
A model that returns 0.51 and is right about half the time is genuinely useful. One that returns 0.51 and is wrong in a consistent direction is worse than a slow LLM.
@CompleteSkeptic The question I'd bring to the AMA: what happens to calibration under distribution shift?
RLCD matching probabilities to outcomes is the whole pitch, and calibration is the first thing to go when inputs stop looking like training.
Every growth thread: "post consistently and provide value."
The actual ranking source: a reply the author answers scores ~150x a like, links in the post body get throttled, and only ~3 of your posts per feed window are even considered.
One of those is documentation.
DeepSeek V4.1 Flash, out Sept 10, MIT licence:
552B MoE, 8B active in / 16B out
1M context, native vision
$0.22/M in · $0.01/M cached · $0.66/M out
The cheapest "good enough" just moved again. Source in reply.
@LuizaJarovsky What if... (Puts on foil hat) AI industry creates fear so the regulators could prevent small open source businesses from competing with big companies and they try to slow down the process to hype up the IPO?
@ClementDelangue@Thom_Wolf@huggingface Open alignment work is the only kind a third party can reproduce, which is the only kind that counts as evidence.
What's the first artifact someone can actually run?
More news is breaking about the potential dangers of AI and its malicious intent. CEOs of AI companies urge a reduction in the speed at which they develop new models.
Why can't there be an international watchdog for AI, just like we have one for nuclear energy (IAEA)?