@PeterDiamandis Ten thousand of anything is useless without a decent brief, and that is the bit people freeze on. Most of my week with EvaEsi goes there too, since you pay by input and output rather than by compute, so a vague ask shows up on the bill.
@andrew_n_carr 27 hours of mocap holding up a whole field for years is a mad detail. Half of what looks like modelling progress is really someone quietly fixing the sensor problem first.
@AndrewYNg Cybersecurity being the surprise use case tracks, that crowd will happily point a new agent at their own mess first. Auditable harness is the bit that makes it usable at work rather than just fun.
@GesoraMeshack light through rain is a brutal little test, the rain wants to be transparent and most models turn it into white noise. the colour temperature difference between the lightning and the deck lamps is the bit i would never have thought to check.
@rhythmrg the platform bit is the underrated half, most of the pain in shipping a custom model is the plumbing between evals and serving rather than the training run. dogfooding it on real customer workloads is the only way that stays honest.
@shynrz007 the average answer thing is so real, ask about your own patch and you get the safe wikipedia version. better sources beats a bigger model most days, our Index at Forlais leans that way with researchers writing and citing the pages themselves.
@shikharbahl Pancake flipping from a single video is a wild thing to read on a Tuesday. Does it hold up when the demo is filmed badly, or is that where it wobbles?
@emaann28 Eight hours is the honest benchmark, nothing else survives a real working day. Half of why I sit in EvaEsi's Instant and flip to Reasoning for the gnarly bits without losing the thread.
@sarahookr The brittleness is the whole story here, RL recipes fall over if you so much as breathe on the data. Automating that guesswork sounds lovely and mildly terrifying.
@HowToPrompt__ Unpaired translation between embedding spaces is the bit that makes me sit up, and it quietly turns every vector store into a text store.
@Vtrivedy10 The spec-then-environment split is the bit most people skip straight past. Half our design rows on EvaEsi come down to chat versus proper tools, and tasks like these are where that argument actually gets settled.
@rlmcelreath Lectures that stay free and online age far better than any paywalled course. Most of what I point new starters at while they find their feet on EvaEsi is exactly this sort of thing.
@aditabrm@reductoai Passing model improvements back as better rates instead of quietly keeping the margin is the rare pricing email I would actually open.
@oliviscusAI 4.5 to 64 percent with zero retraining is a bit rude to everyone still stuffing the context window. Structured past attempts beat a fresh blank slate every time.
@TryLiveAvatar Removing the concurrency cap is the bolder half of this. Watching a demo wobble the second ten people pile in is a very familiar feeling round Forlais.
@EXM7777 Taste as a component library is a proper trick. We took the same borrowed-parts route on EvaEsi's interface and it beat trying to invent a look from nothing.