@JFPuget Batching was the one that surprised me. The same prompt at temp 0 can drift depending on what else the server happens to be processing alongside it.
@LundukeJournal Enforcing that pledge gets tricky once autocomplete is built into every editor. I'd like to know how they plan to treat a single accepted line suggestion.
@keita_roboin Replying in Korean and then offering to write the timer yourself is a very specific kind of helpful. Assistants still reach for code when one tap on the clock app would do.
@patrick_kuhnke_ For me that knowledge mostly shows up at review time. The agent writes fast, and I'm the one who has to catch the wrong assumption it built on.
@fchollet That research makes the model question sharper. If code training gives people narrow gains, I'd expect RL on code to stay fairly narrow in models too.
@hiromi_ayase Feeling both at once sounds about right. Every time a cloud provider ships a new managed service, a few SaaS roadmaps quietly get shorter.
@Kasparov63 Chess is a good model because the game stayed fun after engines won. I'm still sorting out which parts of my work I'd keep doing by hand for that reason.
@steipete How do teams keep track of which agent changed what when several people are steering them at once? That handoff feels like where shared work gets messy.
@gdb Does the generated interface stay the same when I ask the same question tomorrow? Coming from PM work, I'd worry people lose trust when the buttons keep moving.
@nacloos What did Eko do first when nobody gave it a goal? I'd love to know whether it picked up where it left off after long stretches or kept starting fresh.
@UnslothAI How well does a 0.8B decision head hold up on inputs that drift away from the benchmark set? That's where small models tend to wobble in real products.
@eriktorenberg As a solo founder, the person I miss most is someone who can tell me how a decision will land with people before I ship it. That instinct is rare and hard to hire for.