A model that excels at general problem-solving won’t necessarily excel at the specialized work you need it to do. That is the gap the micro1 leaderboard is built to measure.
We just added three new models to our public benchmarks: Muse Spark 1.1, Grok 4.5, and Inkling.
Muse Spark 1.1 leads in Financial Reasoning and ranks second in Pathology Report Reasoning. Grok 4.5 performs strongly across both benchmarks, Inkling’s strongest result is in Pathology Report Reasoning.
Updated leaderboard linked in the comments.
@3YearLetterman Imagine being an impoverished Pakistani man with no clue about what indoor plumbing is and your biggest dream is being an Europoor without AC at home
BREAKING: Sydney Sweeney reveals in a CNBC interview that she was long Nvidia going into earnings.
“We’re on the edge of compute to cheap to meter. We need 1,000x more compute. Nvidia is a structural winner in the AI super cycle” she shared.
Source: @TurnerNovak
thanks to my friend @3YearLetterman for a #great getaway with brad, burt, shane and gary at pigeon forge. i brought my panelsonic camcorder and we made a dockumentary film. coach then notarized it so what you will see guaranteed to be 100% accurate. enjoy