Ling-3.0-flash-Sante launched this week. A medical AI model built for diagnosis and clinical reasoning, marketed on evidence-based retrieval and safety-focused evaluation.
Evidence-based retrieval gets treated like it settles the accountability question. It does not.
Every HIPAA-adjacent workflow I have built treats a citation as the start of a review, not the review itself.
A model that shows its sources still needs a name attached before that answer reaches a patient chart.
The detail that stands out is Qualcomm using Bedrock to speed up its own chip design and EDA work. Once an AI model influences a decision that turns into physical silicon, you cannot patch it later the way you fix software. Whoever signs off on that step needs a real record of what the model recommended and why it was accepted.
@prathoshap Same problem outside AAAI. A model flags something in a healthcare record or a legal filing, and someone downstream still reads every flag to find the real ones. No first level filter on the model's own confidence, and the reviewer becomes the gate.
An uncensored model doing what it's told is the whole point of running one locally, no surprise there. What stands out in the replies is people wiring these into Sonarr, Radarr, download tools, real system access, not just chat output. Once that's true, the thing worth controlling is what the model can reach, since the model itself will refuse nothing.
@unusual_whales This measures how employees feel about AI at their company, which is a sentiment question. Whether anyone can trace what the AI actually did in a specific case is a separate question, and that's the one that matters when something goes wrong.
@paulg Retention numbers this high for legal AI say something else too. A firm keeps renewing because every answer traces back to the specific clause or filing it pulled from. That trail is what a partner actually checks before the work reaches a client.
Subagent spawning is the part worth watching here. Each subagent it kicks off is another actor taking action without a human in the loop. In a regulated build we need a record afterward of exactly which subagent did what and why it got spawned. Proactive delegation without that record is an audit gap waiting to surface.
AI finding the bug in under two weeks is the headline. The month between reporting it to Tencent and confirming the patch was live everywhere is what actually protected anyone. Same discipline any AI accelerated finding needs, give the fix time to land before the finding goes public.
@GulatiYajat Clicking through Tally like a person is the fun part. What matters is these entries eventually feed a GST return. If a regulator asks who booked a transaction and why, the agent needs to leave the same trail an accountant would, timestamp included.
Mine break at the handoff between steps rather than inside a single call. One step passes a plausible looking output forward and the next trusts it without checking. In a regulated workflow that handoff needs its own record, because nobody can point to exactly where a confident wrong answer started once it reaches a customer.
The microVM isolation per machine is the right instinct. In audited environments the harder part comes after that. Once an agent can read and write to any connected machine, you need a log of exactly which command touched which device and when. Isolation stops it from doing damage. A name attached to the log is what proves it didn't.
Running inside the customer's perimeter is the right pitch. The part that actually gets tested later is whether that boundary is provable after the fact. A log showing which model touched which data on which machine is what a regulator asks for months down the line. The marketing diagram does not answer that.
The mounted directory abstraction is what makes this usable, every machine looks the same to the agent. Once one of those machines is running proprietary measurement software though, the record still needs to say which physical machine actually executed the command. The abstraction that helps the agent work is exactly what a reviewer needs undone later.
The AI angle is almost beside the point here. One phone call, a Dutch speaker posing as an IT colleague, no synthetic voice involved, and that is how ShinyHunters got data on 6 million people. Voice based trust for support access was already this fragile. Add cheap voice cloning next and this failure mode only gets more common. Worth building callback verification and out of band confirmation into any process where a single phone call can trigger account access, before this becomes an AI story too.
The jump from incredible to stupid without warning is the harder problem for anyone trying to deploy this in a regulated flow. A model that's consistently mediocre is easy to write a review process around. One that swings between genius and broken on the same task type means you cannot decide in advance how much a human needs to check, only after the fact whether they should have.
@sqs Steering instead of queueing makes sense for speed. The harder part in a regulated build is the log keeping pace with the agent's course corrections. If it changes direction three times before landing on an answer, the audit trail needs all three moves, not just the final one.
@IntCyberDigest Police had to specify the voice wasn't AI, which says something about how normal a cloned voice call already sounds. Either way the real failure is the same: no single support login should reach 6 million customer records, human caller or AI voice.
Feeding public government documents into an external AI model already means giving up control of where that data lands. A training run does not get walked back once it happens.
The phishing page getting the attention here is a fake CAPTCHA that talked someone into pasting a command straight into PowerShell.
Same root gap either way. Nobody signed off on what happens once something leaves the building, a terminal session or a training pipeline.
An AI system doing speech recognition and case management for 6000+ courtrooms is publicly claiming data stays in India and stays encrypted. That should get independently verified before deployment, as part of vendor review. Instead it took an outside researcher testing it after launch to find neither claim was true.
Models improving themselves without human input is the headline. The real gap is that the response to it is voluntary self-restraint. A safety standard nobody outside the lab checks is a promise, not a control. Internal policy alone never satisfies a regulator. Someone outside has to confirm it held.
Serving generative AI to 50 million citizens directly from government data centers is a real scale test for the model. The harder part starts after launch. When someone acts on a wrong answer about benefits or legal rights, there needs to be a record of exactly what the model told that specific person and when.