Check to try: compare review counts and visibility separately. Record a baseline, choose one change and measure again. Treat any improvement as a result to investigate.
Measured 2026-08-10. https://t.co/I7xvIb8Mxy
Use the fillable worksheet on the research page to choose one action, record a baseline and set a review date. Want to understand your own business? Explore its free Growth Brief there.
Check to try: write down your actual diagnostic fee, warranty and service hours, including conditions. Can a customer find and understand each before calling?
Measured 2026-08-10. https://t.co/TW6mMaE5pu
Use the fillable worksheet on the research page to choose one action, record a baseline and set a review date. Want to understand your own business? Explore its free Growth Brief there.
For @AnthropicAI to really fumble this bad with all the negative news before their upcoming IPO means either they are incredibly mismanaging everything... or they don't care because they have something up their sleeves anyways... so I really hope it's the latter
This may reveal a gap between benchmark performance and production reliability.
A model can be excellent at synthesizing data while still being too confident about what the evidence proves.
Has anyone compared Gemini 3.7 Flash and GPT-5.6 Luna on similarly evidence-heavy work?
Gemini 3.7 Flash didnโt outperform GPT-5.6 Luna in our evidence-heavy data analysis testsโeven though public benchmarks suggest it should.
Has anyone else observed this?
We sent both models the exact same prompt and data. Hereโs what happened ๐งต
So my conclusion isnโt that Luna is universally better at data analysis.
Itโs that Luna was more reliable for our specific task: analyzing large dossiers where provenance, uncertainty, and defensible recommendations matter more than confident presentation.