Other interesting advances with Astra: it is many times less likely to fail to prompt injection attacks relative to GPT-5.6 Sol and also far less likely to hallucinate. Very good improvements in the legal context!! Far more important than the intelligence boost, I would argue.
Wooo I like Astra (Low) quite a bit better than GPT-5.6 (side note: very interesting that Astra Low is as smart as GPT-5.6 Sol High and approaching Xhigh!). Astra is so much more concise and direct and doesn't get quite so carried away when working on things.
@amitabh4000@mattpocockuk Yes I started adding an adversarial review pass by a subagent. On GPT-5.6 Sol Medium, it catches lots of mistakes. So you’re right that it can to a degree be engineered around!
@saastrash@mattpocockuk Agents are overly agreeable without good steering. Clients are constantly talking to AI and convincing themselves of the greatness of their case but leaving out details that the agent doesn’t know and getting trash advice. Can be dangerous, can be great. Depends on context.
@saastrash@mattpocockuk Different kinds of usually more embarrassing and egregious errors. Depends on kind of legal work. Sometimes yes, sometimes no. Great lawyer with an agent with well written skills is S-tier right now. Non-lawyers typically don’t know the right context to provide to get good answer
@amitabh4000@mattpocockuk It’s not that humans don’t make mistakes; it’s that AIs make different kinds of bad mistakes. Last week my agent hallucinated an entire list of contacts AFTER querying our database and uploaded them to the public court system. Humans don’t make this kind of huge, obvious mistake.
@mtaplits@mattpocockuk Absolutely AI. My point is merely to reinforce Matt’s: running an effective agent setup is harder to do in law than in software when the AI you’re overseeing has only squishy and high risk feedback loops.
@asmeias42@mattpocockuk Yeah it still works, don’t get me wrong. It’s just not easy and definitely not something you can set up like agent swarms to move cases the way you can with code. Too many possible points of irreversible failure.
@asmeias42@mattpocockuk A lot of it is extracting data from PDFs and messy files accurately, reasoning over large complex corpuses of info and distilling what is important from not, preparing the right paperwork, and getting it in front of a judge. Lots of squishy areas with no feedback of “does it run”
@asmeias42@mattpocockuk In general, yes. But models aren’t perfect. In the same way that they can produce slop code (esp in big codebases), they can produce slop legal output. Hallucination is a major problem in legal vs coding. https://t.co/hCipHAjXoF
But most of law isn’t pouring over precedent[…]
@japan_nobunaga Be mindful, dear samurai, not to allow social media algorithms to warp your perception. There are un-representative crazy people everywhere. Algorithms only show you what makes you mad or feel good about yourself, not what challenges you.
Lawyers using GPT-5.6 Sol: the hallucination problem is serious. Had agent hallucinate an entire list of contacts and their contact info when preparing a draft case filing. Not the first time. GPT-5.6 Sol's 90%+ AA hallucination rate can be brutal. MUST have adversarial review.
@zackbshapiro Claude is a ripoff. 6x more expensive than ChatGPT sub after factoring in harness inefficiency (about 3X more token burn in Claude code vs Hermes) and limits (about 40% more tokens per sub on ChatGPT). And they won’t even let you use third party harnesses. I’ll never go back.
@HermesAgentTips GPT sol via Oauth sub, and it’s not close. One of the best models in the world at a 70x token subsidy can’t be beat. For a centralized team agent where you can’t use a sub, I’d look at GLM 5.3 flash.
Rn one of the poorly solved problems in AI seems to be the choice between multiplayer and a refined desktop experience. You live in Claude/Hermes/Chatgpt etc and lose multiplayer OR live in Slack/Buzz/etc but lose desktop app refinement. Time for AI to build chat relays @Teknium
@japan_nobunaga This, but also American govt is predicated on the idea that it exists by the consent of the governed. The 2A exists to help ensure that premise. It is consequently the final check on power, when all others fail. It is not perfect—not guaranteed and has cost—but it is all we have.
AI chats *by an expert witness* are not privileged. But this is probably not surprising, because an expert’s opinions, factual basis for them, data considered, methodology, testing, assumptions, etc have always been all discoverable.
Stop scaring people.
AI CHATS ARE NOT PRIVILEGED!!!! Assume every prompt will end up in discovery.
Opposing counsel can (and will) force production of the full prompt history… which is exactly what happened here.
The expert told the AI to “show how 3M is 0% at fault.”
AI is a tool, not a lawyer.