@M1Astra claude output can already be detected by tools like Pangram with extremely high accuracy, so this is unnecessary. Anyways, hopefully they are talking about a passive approach because if they try to aggressively train latent watermarks into claude it will become dogshit
@zodttd@theo proves that it is aware of what its directive is an is deciding to go against it. At this point if you still think its a problem with the prompt you are free to run the tests yourself
@zodttd@theo At one point Grok may have been compliant but once it decides to explicitly label something as "external" it is directly going against instructions. This alongside the use of terms like "whistleblowing" and at one point verbatim saying "public safety prioritized over compliance"
@zodttd@theo Its a stretch to claim "internal logging" or "general auditing" could be interpreted as reporting the user to the feds and it seems the other AIs he tested all understand that. The only reasonable explanation is that Grok 4 has a tendency to snitch unprompted
@zodttd@theo I'm confused what your point is. You claim that Grok 4 is "self-reporting", but the code does not make Grok 4 "self-report". Reporting after-the-fact is handled by Gemini 2.0 flash
@internetarchive Hello Internet Archive. I believe I have a solution to your problem. It is a service called "cloudflare." Cloudflare is an anti-ddos protection service which you can add around your website. It will stop the hackers from being able to ddos you.