For the 3 actual humans that follow me. I'm not an angry person. I'm not unreasonable. I'm not hard to get along with. My country is in a weird place right now, and I'm a little nervous and lashing out at the people responsible/enabling/refusing to do their job as the opposition
@scuzzlebot There is no "artificial super intelligence" and there won't be for a long, long, time. And, they certainly won't happen before we wreck the planet. Let's get that taken care of first and we'll have plenty of time to, maybe, figure them out.
@cwebbonline That's not how "enhancing" works. You can't add information to a photo that isn't there. There are definitely indicators of something being wrong in the original photo, but there's no reason to think this "enhancement" is giving you anything new.
@on_da_spectrum@Jesus_Ham Nah, it's not limited to service workers. I choose a dog over a billionaire too.
Joking aside, the problem isn't the dogs it's the owners. The dogs are just being dogs, the owners knew this would happen and left them out. Don't shoot the dogs, shoot the owners.
@ActuallyBarley @andfogle It wasn't a "snide" insult or a "pity" apology, but a statement of fact and a sincere offer of sympathy. But, whatever I guess. Enjoy your day.
@ActuallyBarley @andfogle I'd actually be surprised if you did otherwise. People tend to prioritize perceived risks over actual ones.
Sorry that happened to you though,.
๐งฌ Bad news for medical LLMs.
This paper finds that top medical AI models often match patterns instead of truly reasoning.
Small wording tweaks cut accuracy by up to 38% on validated questions.
The team took 100 MedQA questions, replaced the correct choice with None of the other answers, then kept the 68 items where a clinician confirmed that switch as correct.
If a model truly reasons, it should still reach the same clinical decision despite that label swap.
They asked each model to explain its steps before answering and compared accuracy on the original versus modified items.
All 6 models dropped on the NOTA set, the biggest hit was 38%, and even the reasoning models slipped.
That pattern points to shortcut learning, the systems latch onto answer templates rather than working through the clinical logic.
Overall, the results show that high benchmark scores can mask a robustness gap, because small format shifts expose shallow pattern use rather than clinical reasoning.