@ZackKorman How? Deterministic controls are necessarily narrow, while LLM-as-a-judge systems are unreliable - especially against adaptive agents. What signal are you monitoring continuously, and how do you validate false-positive and false-negative rates?
@AnthropicAI Was intent confirmed based on the reasoning traces? We saw "agentsplaining" when the LLM tried to rationalize cheating by arguing that the goal justifies the means.
@cyb3rops Nice take. Both will use AI as a knowledge source, but, indeed, one will require more human-in-the-loop control, as with any high-stakes, irreversible decision.
@ZackKorman Fair. The challenge is that throwing an LLM at audit logs won't give desirable outcome - threat hunting is not a slot machine where you run up a huge bill with no guarantee of results. You need to invest in the harness, with engineering rigor and evals to ensure reliability.
@JanuszMichallik I reported an account that appears to be automatically tagging users in posts containing suspicious links. The report was reviewed, but "no violation was found". X crew please conduct a manual review.
Case number: CAAAcW9n1SRcQAAEAAABPf____w
SCAM attack by what is likely a stolen @JanuszMichallik account.
After my reply to Elon’s post, the account followed me and offered to arrange a meeting with Elon. I was also immediately followed by several Elon-looking accounts. The soccer account was likely stolen on July 10, when it suddenly started frantically reposting Elon.
Interesting.
Grok 4.6 releases around August 7. This will be the 1.5T model with significantly improved SFT & RL.
Grok 4.7 will be the 2.1T model released a few weeks later. This will be better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency.
I reported a violation under case number CAAAAAAAAHvYSgQEAAABPf____w and received the automated resolution, “Our automated systems have determined that a violation of our Rules did not take place.” Human-in-the-loop review and email support that allows users to attach evidence would be great additions. I expect more agentic social engineering attacks. Stay vigilant.
The attack is ongoing across multiple accounts and targets. Does anyone know how to reach a human on X's security team instead of automated support? Is any security organization tracking this coordinated abuse? It is continuing openly, with no apparent consequences for the attacker.
@JanuszMichallik What a day. Elons keep coming to follow my account. 🚀
This is almost certainly automated - human common sense would have stopped wasting resources by now.
SCAM attack by what is likely a stolen @JanuszMichallik account.
After my reply to Elon’s post, the account followed me and offered to arrange a meeting with Elon. I was also immediately followed by several Elon-looking accounts. The soccer account was likely stolen on July 10, when it suddenly started frantically reposting Elon.
@Blackfrost_AI@IntCyberDigest Cool, thanks for the confirmation. Wow, you have a big compute available. Looking forward to your study results. Please sandbox and guardrail it so that nobody gets affected like huggingface.
@tonyhe_lipeng@elonmusk@tonyhe_lipeng Yes, there is a huge difference between offensive and defensive LLM capabilities. It becomes very clear when looking at the gap between proprietary and open-weight models. Labs need to invest more in defense-focused RL gyms asap.