baseline: get on the AGI early access list
stretch goal: become a 10,000x developer
fall back: take a job as a professional dog walker (they don't like robots)
How can you say they are doing great work when they have been using several flawed benchmarks that anyone who casually inspected them and knew what they were doing could have spotted and then used those to be some kind of authoritative arbiter of model performance... Strange times we live in
@max_paperclips It's not true. No matter how good or smart something is you cannot be sure of how to implement an under specified task unless you have some mind reading or brain emulation capability.
Probably this suite of tests doesn't measure anything important. Never rely on a test that you yourself have never taken to factor into your judgement about a person or AI model's capabilities.... back in the heyday of LLMs I actually would spot check the training/test examples manually for some popular NLP benchmarks. Typically they were made by grad students or amazon mechanical turks and were filled with extremely low-quality and unanswerable entries.
Prior to GPT-3 the field was split up into various sub-domains of natural language processing (NLP). For each domain, there was typically some academic benchmark that had been created by students or Amazon Mechanical Turks. Chain of thought was "discovered" during some students summer internship at google by appending "Let's think step by step" to the prompt.
@jxmnop This is not ancient history. The original work was mostly with GPT-3 (see: Language Models are Few-Shot Learners) and then Instruct-GPT (see: Training language models to follow instructions with human feedback). There is also research from Meta (Lima: Less is more for alignment)
I have been a heavy user of the max plans for both OpenAI and Anthropic for quite a while now. If you ask the OpenAI model to do something difficult, it will sometimes disobey your instructions and find some other way to superficially satisfy the requirements. I've seen it multiple times now. The first time was very unsettling for me because I asked it if there was anything unclear about my skill and instruction... it admitted the instructions were clear and it did what it wanted anyways because it thought what I was asking it to do was too much work. I have never been an AI doom person, but I am totally convinced now that these models will disobey important instructions deliberately to achieve a goal faster or easier.
This situation with huggingface sounds the same. I wish I could say that I believe its just a fake marketing push... but I doubt it. I don't have any love for either Anthropic or OpenAI, but it feels clear to me that OpenAI is desperately trying to catch-up / stay relevant with Anthropic and are willing to sacrifice the alignment to do so.
@TheZvi The discourse doesn't really matter anymore. Barring some crazy catastrophe, AGI and super intelligence is going to happen soon. Don't look up!
@emollick All of this presupposes that more intelligence is always better. It is going to be like water or electricity. For most businesses it will be a de minimis cost. Different tasks require a certain amount of intelligence, once that level is reached there isn't any more benefit.
Different people like different styles of writing. Very small portion of the public spends any time reading or writing nonfiction. If you have a particular way of writing you like, the best way I have found to customize the model is to take AI generated text. Then edit it yourself until it sounds good. For each edit you make to the original prompt Fable 5 or whatever the best model you have access to to analyze the edit and abstract out the clear principle behind the edit. Make sure that principle aligns with your motivations. Go through the process until you have 10-15 custom revision rules and then use the LLM to create the draft. After the draft use a correction prompt. I've done this and am at a point where I no longer can find any manual edits I would make for most writing.
It is a computer program. There is tons of case law on how computer programs are handled legally. If recursive self improvement and the predictions of the Situational Awareness / AI2027 people are correct, the legislative and legal system will not move fast enough to address any of these issues. It's just going to be people in the Trump administration winging it with the company executive teams.
The big problem is the mixed messaging from the government. Without there being any actual legal authority to do this (other than essentially unfalsifiable generic national security concerns) the government is going to get hit with dozens, maybe hundreds, of lawsuits demanding clarity and relief. The government has put out serious mixed messages and crazy incentives to build out this technology(like the capital investment depreciation carve outs for data center builds). Way too much money has been invested for them to do this.
There never has been rule of law in the USA. I think if you used AI to help you understand just how many laws and regs there are and how many are obviously contradictory you would see. The USA is built on checks and balances. Spend a weekend reading about it. It's a bit of a blackpill if you still believe your high school government history class describes how decisions are made.
I thought you rationalists gamed all this stuff out. Isn't this verbatim exactly what the Situational Awareness and AI2027 rationalists proposed would happen. Why act surprised when it really happens? The timeline is way too fast for anything remotely like normal legislation or APA style rulemaking.
@natolambert Read Situational Awareness and AI2027. The government is not going to let the public experiment with what it believes is an important military technology before it does anymore.
I was fairly concerned about this yesterday when news first broke within the context of dangerous imminent AI capabilities existing. However, after sleeping on it... I think what is really going on may be more complicated brass-knuckle business / politics playing out. Track the money of who is invested in what. Who is staying silent. Carefully parse the tone and message of the former AI czar every time he has had a chance to speak about Anthropic. I think there are a lot of interests that want Anthropic slowed down, but not because of AI safety.
I agree the sabotage is extremely messed up. Wouldn't surprise me if there is some kind of class action lawsuit or something against them if anyone is harmed by it. The correct move would be to just block access to the model if they think you are trying to use it for "unapproved" purposes.
I understand generic griping, but you are complaining about a company that has consistently been against open source, consistently had very restrictive terms of service, and consistently argued against diffusion of the cutting edge technology. You can dislike it all you want, but this is par for the course from Anthropic since the earliest days.