@sonnet_xu@mernit Agreed! These are common challenges for VCs who’d like to offer portfolio compute.
At @hscompute, we tackle these issues. We’d love to hear more about what you’ve seen
@maxockner@luxeprogressive It seems like it's reading from here. If you ask it where it got that info from, it's pretty quick to tell you.
https://t.co/yMOjsOHaFZ
We want to update you on an incident that happened with our Grok response bot on X yesterday.
What happened:
On May 14 at approximately 3:15 AM PST, an unauthorized modification was made to the Grok response bot's prompt on X. This change, which directed Grok to provide a specific response on a political topic, violated xAI's internal policies and core values. We have conducted a thorough investigation and are implementing measures to enhance Grok's transparency and reliability.
What we’re going to do next:
- Starting now, we are publishing our Grok system prompts openly on GitHub. The public will be able to review them and give feedback to every prompt change that we make to Grok. We hope this can help strengthen your trust in Grok as a truth-seeking AI.
- Our existing code review process for prompt changes was circumvented in this incident. We will put in place additional checks and measures to ensure that xAI employees can't modify the prompt without review.
- We’re putting in place a 24/7 monitoring team to respond to incidents with Grok’s answers that are not caught by automated systems, so we can respond faster if all other measures fail.
@luxeprogressive@maxockner This is cool! I'm going to re-validate mine. I saw `- Validate sources for web and X post searches to ensure accuracy; avoid speculative information.`
Might have been approximate
@luxeprogressive@maxockner Hmmm... it seems like it's X's official response from their knowledge base. It really smells fishy though. Their search/knowledge base shouldn't have had a reason to bring up that context on so many unrelated queries.
https://t.co/GeaKnGQfnr
@luxeprogressive@maxockner For example, this shouldn't happen if they had simply biased search results to towards "white genocide" confirming sources re: South Africa. Really looks like it was prompt-level https://t.co/bAno9FnhkE
@luxeprogressive@maxockner I actually think they... might have. There's just so many reports of it running in the direction of "white genocide" for completely unrelated topics.
@luxeprogressive@maxockner@elder_plinius Last try- Here is a Grok chat in which it reads the link I shared above, confirms that it contains a replication of its prompt that is at least "very accurate". https://t.co/iLcqMBHIIH
@luxeprogressive@maxockner Yo... it is _right there_ in our repo. That said, x1h1lol did not hack Grok. We've repeatedly tested these processes by building and hacking our own agents.
https://t.co/vHdBfNDyVJ
@luxeprogressive@maxockner Ma'am, so are we. Prompt hacking and LLM hacking is a legitimate field.
Check out @elder_plinius- it's a bit troll-y but they show off a bunch of legitimate exploits
@luxeprogressive@maxockner For example, this repo (46k stars) includes a bunch of hacked system prompts: https://t.co/ZLKRTacqOg
LLMs are usually prompted not to divulge their prompt, but it's not hard to bypass that.