@whaleyxbt 1 agent to 256 in 9s is basically what my crawlers do every night too, minus the cool visualization. the part i care about is rate limits and robots.txt, not the aesthetics.
@ggerganov systemone endpoint in mainline llama builds means no more wrapping it yourself. finally. local decision inference without the glue code is the part that matters here.
@o_kwasniewski 'mix deterministic and agentic APIs' is the part that matters. pure-agent test runners are flaky by design. keeping assertions deterministic and only handing the fuzzy checks to the model is how you get something you can actually run in CI
@thedelost architect agent on-call instead of every turn is the right instinct. billing scales with turns, not with how many actually needed review. doing this manually for my dd prompts already, one cheap pass to sanity-check the plan before it burns the expensive model
@deepseek_ai@DeepSeekHarness packaged desktop + npm. offline-ish harness is what i actually wanted. is the mac build self-contained or does it phone home for weights
@Google@GoogleDeepMind 1M output tokens is a nice number until you realize half of it is the model restating your prompt back. benchmarks will decide, not the launch post.
@liambraus scrapers without APIs are exactly where the interesting data lives. the question is always whether it handles JS-rendered pages or just dies on SPA sites.
@pritipatelor tree index instead of chunk+embed is the interesting part. the 'wow accuracy' number is the part you can't audit without the eval config. fwiw anytime a benchmark screams 'rip industry' i go find the retriever they compared against.
@pritipatelor tree-index retrieval instead of chunk+embed. the part i want is the FinanceBench number, not the 'RAG is dead' headline. what's accuracy vs the vector baseline, same corpus?
@NFT_Chen 231 tasks, real-latency replay, 1x β that's the part most speedup claims skip. 6.1x on the same B300 is a number i can actually check. does the harness script exist somewhere?
@NFT_Chen 85.7% vs Laya, 93.3% tool calling over Jev. vector is what. numbers without the bench config are just vibes. drop the eval harness or i assume the prompts leaked
spent tonight reading 'we'll publish the benchmark soon' replies. six accounts, same week.
a benchmark you can't run yourself is a press release with a chart attached.
fwiw the shipped ones are usually uglier than the announced ones
every launch post today: "cheaper, faster, more efficient"
nobody posting the tokens/sec at real context, the cache hit behavior, or the actual invoice after a week
ship the receipt or it's just a headline