can AI do research-level mathematics? make conjectures? prove theorems?
there’s a moving frontier between what can and cannot be done with LLMs.
that boundary just shifted a little. this is my experience with AI proving a new theorem.
1/
Are we able to agree on what we mean by "AGI". I've been using this definition from OpenAI which I thought was relatively standard and ok:
https://t.co/lMHgmMaLgG
AGI: "a highly autonomous system that outperforms humans at most economically valuable work"
For "most economically valuable work" I like to reference the index of all occupations from U.S. Bureau of Labor Statistics:
https://t.co/tGVU57v7WG
Two common caveats:
1) In practice most people currently deviate from the above definition to only mean digital work (a relatively major concession looking at the list).
2) The definition above only considers the *existence* of such a system not its full deployment across all of the industry.
Some people say GPT-4 is already AGI, which per above definition would be clearly not true. LLMs are useful tools for most of these jobs but you clearly couldn't hire them to autonomously perform them in full and autonomously at human+ capability.
Last note some people say the goalposts keep moving, which I mostly disagree with. I think the definition above makes sense, it has been stable, and has clearly not been reached.
Prompt engineering (or rather "Flow engineering") intensifies for code generation. Great reading and a reminder of how much alpha there is (pass@5 19% to 44%) in moving from a naive prompt:answer paradigm to a "flow" paradigm, where the answer is constructed iteratively.
example
https://t.co/G0bd6j0PD1
U.S. ranks 33rd worldwide ... nearly double that of the next-highest developed country and almost three times that of Canada. ... Denmark, Iceland, and Switzerland, for example, have all reported zero police killings.
On e/ia (intelligence amplification), I'm continually impressed with how Web-connected LLMs can help me clearly explore ideas within a news article.
Q: "Search the web for information on the topic in this article, include trends in other nations where possible."
"The reality of goodness, not the perception of it ... Let’s see how Earth responds to that ... by that time we'll have digital gods ... we live in the most interesting of times ... even if I knew that annihilation was certain ... it was the most interesting thing ... "
@sama, congratulatios to the #OpenAIAssistant team!!
Looking forward to their upcoming smarter assistant-creating assistants that help us mold what great responses look like. Exciting times!
@sama , amazing innovation by the #OpenAIAssistants team!
Here's an upvote on features that help us express "answer quality".
(plus, one-day, personalized likes maximization).
How did our medical #LLM, Med-PaLM 2, become the first to perform at “expert” level on U.S. Medical Licensing Exam-style questions? Check out our new paper: https://t.co/dyMoJVSJyE
"... ChatGPT’s math where it makes nonsensical errors such as adding two numbers in a subtraction problem or dividing numbers incorrectly. ..."
If "math" is in the way, then my money is on v5. ;-) https://t.co/1AWYGvwUBC
Oh oh, @MonsterEnergy uses #erythritol . It's their third ingredient, after water and citric acid, so there could be a lot per can. Looking forward to them publishing a response.
#Metabolomics analyses reported an increased risk of #Cardiovascular Disease associated with the #ArtificialSweetener#Erythritol, supported by mechanistic studies showing that high levels of erythritol enhanced platelet reactivity & #Thrombosis formation.
https://t.co/FWQE2k3hWx
12/12 Everything has a price, and now, in the spring of 2022, we must pay this price. There's no one to do it for us. Let's not "be against the war." Let's fight against the war.
@UnitedAirlines_ I see that you've responded to others today that your online services are flailing. Do you have a webpage that tells me the latest status? Hopefully not random checks on Twitter! (not even your own feed currently mentions the problem). Thanks.