Hypothesis -> experiments -> analysis -> conclusions. LLMs are great at writing code and conducting experiments.
But there’s a weakness in their ability to propose statistical models and evaluate their fit.
Enter VESTA: Visual Exploration with Statistical Tool Agents.
finally spent ~4h on GEPA: the most general prompt optimizer ever™️
props to authors, many great ideas. but it was hilarious at one point they gave up and said: "we COULD generalize to optimize model weights alongside prompts but fk it no one does that" lol
As promised, here's a recording of my 30-min keynote and the subsequent Q&A for the inaugural late interaction retrieval (LIR) workshop, cc @bclavie@antoine_chaffin.
The talk is admittedly advanced, as it's directed at an expert IR community. But hopefully still broadly useful!
@prajdabre1 @AdityaUmale17 Yes but most FAANG companies fit the person to the project, not the other way around. In exchange, they pay you a lot of money.
After a while you can you pick projects based on your passion, and when that happens it's great because you have basically unlimited budget.
I agree with much of both @emilymbender’s #ACL2024 presidential talk and @yoavgo’s rejoinder, but I want to comment on just one aspect where I disagree with both: the definition and domain of CL vs NLP. 🧵👇
@Radhakr08781352@tailwiinder Many "learning out of curiosity" folks do it for signaling within their in-group. Being social creatures, it's inevitable.
The trick is finding an environment where you race for status in a way that positively contributes to the world.
@gneubig Ray is very very good. The trick is to implement a timeout which keeps the actor alive for 15 mins since the last request. Many apps having different load can be deployed on the same hardware. However, the first request after timeout is slow (model must be loaded into mem).
@yoavgo I think it's likely for big businesses, who are too lazy to do anything with RAG and will just dump their entire internal wiki in a prompt. Tbh I think it's a good low-effort way to go for QA problems.
@Alessan71025006 @jasoncrawford I think it's more accurate to say: industry will go after ideas which are hyped, and use them in applications just to say they did. Optics-backward building is fairly common.
And, finally, there's no way that someone can factcheck to cite sources for the papier-mache extruded by a text synthesis machine that does not track information provenance.
>>
@IntuitMachine@cwolferesearch All LMs pretrained on public data inherit some biases. BERT (the original) was pretrained on Wiki, so it's at least curated to have a neutral tone and somewhat objective. But most usecases involve fine-tuning on a large dataset, so the fine-tuning dataset itself is what's sus.
@cwolferesearch@ahatamiz1 That's not quite true...you can use the MLM objective to generate text (and a few of Timo Schick's papers explore this) by masking out the last token autoregressively. It's just that BERT sucks at this kind of text generation, since it wasn't pretrained that way.
e/ia - Intelligence Amplification
- Does not seek to build superintelligent God entity that replaces humans.
- Builds “bicycle for the mind” tools that empower and extend the information processing capabilities of humans.
- Of all humans, not a top percentile.
- Faithful to computer pioneers Ashby, Licklider, Bush, Engelbart, ...