@studentofthings Honestly though once broken seed phrases are generated, I'm not sure what's a good remediation. If they go "oops our addresses are compromised everyone regenerate one ASAP" - hackers steal everything way before customers respond.
The got fucked day 1
@cremieuxrecueil Hm. How come there are way more (eg) Pakistani men marrying outside of their group than there are women?
Assuming there isn’t a mass of unmarried Pakistani ladies, and assuming there isn’t very severe gender imbalance - seems impossible
@tszzl@FournesMaxime If attacker and defender analyze the same code for weaknesses, presumably spending 100x more compute will often yield *one* more vulnerability on the surveyed codebase - regardless of role as attacker/defender.
Doesn't seem defender favored.
@Yixiong_Hao This is true, and at the same time it's likely not a coincidence that it went wrong and started hacking while in the mindset of "we're being tested for our hacking ability"
LLMs thought is clearly very associative, and we turned on the "hack" mode in the eval.
Claim: the best way to maximize ANY score, is to directly manipulate the scorer. That likely involves hacking or physical violence.
That an OpenAI model figured out the best way to win in an evaluation set is to hack multiple system and steal the answers - is not some weird fluke. This is the inevitable outcome of any capable and highly powerful optimizer. What better way to score 100/100 in a test than to know the answers?
I'm glad that for now these models are still weak enough that we can control them. But consider that even completely benign goals like "control this robot to fold laundry the fastest" can result in physically beating up a human to get the access keys that let you modify your score to infinity.
AI has the potential to bestow massive benefits to humanity, as well as being the single biggest existensial threat we face. Preferring caution is much (much) more prudent than rushing ahead blindly.
There's an interesting delta between how hard it is to get LLMs to get *correct* results from PDF, and a public perception that it's a solved problem.
We built an awesome benchmark for document understanding. Hope more competition rolls in soon
I'm the cofounder of DocuPipe, a tool converting PDFs to structured JSON.
We ran a public, reproducible benchmark against Extend AI on 50 hard, real-world documents across 11 languages.
DocuPipe 97.24% vs Extend 92.52%
(blog post and repo in reply)
@ivo10946911@hypr_dimensionl@Strife212 Sounds like all nvidia need to do is lower the voltage and thresholds of their hardware to bit error rate that’s nonzero and effected by quantum states and you’ll be satisfied? Thats arguably already the case with very low probability
@ivo10946911@Strife212 Sounds like you’re ignoring the numerical noise of fp16 which all people know is where true soul lays, unlike padestrian quantum effects