Come join the CBI (cognition, behavior, and information) initiative at @asu_ssbs!
Two positions available, one potentially at the associate level. Deadline Dec. 1st.
https://t.co/WbCD4oK2L9
@DavidMSidhu I assigned writing magic school bus fanfiction for an intro cognitive science class - pick an experiment and a neuroimaging method - explain the physics of the img method and the experiment. Was fun, would do again. Good science communication and independent research practice.
Before everyone leaves twitter pls share with any promising UG students that my lab and the Cognition Behavior, & Information group at ASU are
recruiting (funded) MS students for 2023!
My lab: https://t.co/xSPA6wMJMD
CBI: https://t.co/t956qptckc
@talyarkoni@Ted_Underwood Agree. Text hasn't had its stablediffusion moment yet, but all the xformer and quantization work is shrinking hardware requirements for the biggest LLMs rapidly. Once it runs on consumer hardware it's a different game.
@_akpiper You're right those steps take time. In consultations about corpus methods the first q I ask is "is there more text than you can read or do you have a specific hypothesis to test?" If the a is no and no, then just read it! But often the data > what is readable in a human lifetime.
Join a high-energy team in a growing School at ASU to help us transform science. @asu_ssbs is hiring TWO tenure-track / tenured faculty in Psychology or proximate field as part of a continuing expansion of our Cognition Behavior & Information initiative.
https://t.co/T6Xr3Eei3n
Do causal inference folks have anywhere they keep track of causal inference related job openings? Because I've got two in psychology at assistant/associate level:
https://t.co/WbCD4oK2L9
@yudapearl@rlmcelreath
Heather Wild @gottabewild’s talk “Valence effects in the wild: Analysing word learning in language learning apps” was awarded “Best Talk” at #MentalLexicon2022 🎉 Congratulations Heather!
We need to keep asking under what conditions can cross validated error be taken as a valid indicator of generalization error? And do these conditions hold for where LLMs are getting deployed? When validation sets and deployment sets aren't from the same pop, CV is optimistic.
This overlooks an important aspect of the initial comment though, which is that LLM eval tasks (including probably more than half of BIG Bench) mis-represent what they are measuring; scaling ofc improves perf on those tasks but both human & ML metrics on such tasks are misleading