Can we use LLMs to evaluate open-ended instruction following generations? Introducing LLMBar, a benchmark for evaluating LLM evaluators
🧐LLMBar is manually curated, objective, and adversarial😈
🤯Most LLM evaluators cannot beat random guess!
📜https://t.co/w0ta1dgA4T
[1/n]
Can we use LLMs to evaluate open-ended instruction following generations? Introducing LLMBar, a benchmark for evaluating LLM evaluators
🧐LLMBar is manually curated, objective, and adversarial😈
🤯Most LLM evaluators cannot beat random guess!
📜https://t.co/w0ta1dgA4T
[1/n]
Zeno Build is a tool we're building for analysis of large language models. We have an exciting new feature thanks to @LiteLLM: you can now compare 100s of different models with the same code! Just select "litellm" as the provider when generating outputs: https://t.co/pA2IJOyum2
📢 Thrilled to announce our latest paper!
"Merge Conflicts!" Exploring the Impacts of External Distractors to Parametric Knowledge Graphs:
https://t.co/knZReXIibj
🤖 Experience the CLASH of external knowledge and LLM's parametric knowledge through its inner mechanism!
(1/n)
We’ve added support for a number of LLMs for Prompt2Model: Cohere, Anthropic Claude, PaLM, etc! Just set the `api_tools.default_api_agent` parameter in the Prompt2Model notebook/script and run everything the usual way.
https://t.co/bdohmjpZMV
Whereas people in HCI and CS would be more interested in how to build new start-ups with agents rather than spending a fortune on policy simulations... Well, there aren't really any idealists like that, after all.
Today, I came across a talk given by Prof. Michael Bernstein. I just feel that I'm more interested in Michael Bernstein's ideas on agents than what the charlatans are doing—simulating how the human world reacts to complex environmental changes.
For example, the California government could create a GTA-sized simulator and then see how their policies affect the city of LA.
Somehow, this is related to sociology, and the funny thing is that sociologists probably don't have such strong engineering skills.
2. Data is more flexible than a model: with these data, users can further train models with RLHF, PEFT, or whatever. We can integrate these training methods into Prompt2Model, but the center is still high-quality personalized data.
After chatting with a start-up company, I found an interesting topic:
The previous focus of Prompt2Model is small models, but as we realized and wrote in our paper, the center is data. Our next concentration is probably Prompt2Data rather than Prompt2Model.
Prompt2Model: Generating Deployable Models from Natural Language Instructions
https://t.co/1Z48ST624l
Not just "write prompt get model", but get finetuned smol model >20% better than GPT3.5 while being 700x smaller.
Read thru @gneubig and @vijaytarian's great paper this weekend. If ML had AutoML, this would be "AutoLLM" - the natural next step in Automatic Prompt Engineering would be:
- get prompt for a model to create
- use HyDE to match against @huggingface dataset (with human in the loop to choose) and model (e.g. Flan T5)
- generate extra data from GPT3.5 to supplement (there's an ablation to show having generated data + offtheshelf data does better than either alone, and matches expensive custom labeled data)
- finetune the retrieved HF model, eval it, and serve it
The Prompt2Model repo includes a published and maintained python framework for you to use it immediately (very interesting trend in academia to ship maintained software as artefacts). Lots and lots to learn and pick up from, while I also share their concerns about the generalizability of this technique.
Also, I think Prompt2Data is more generalized than Prompt2model in:
1. Data is the center of the current paradigm: training a small model is easier than getting personalized data.
There are two main concerns on Prompt2Data:
1. How to create it? We are focusing on distillation and its tricks.
2. How to evaluate it? We are finetuning the model with these data to reflect its quality. Is there a better way?
So I finally did what I promised 3mo ago – We wrote a reflection paper on using LLMs to replicate crowdsourcing pipelines, based on an assignment in our Human-Centered NLP course! https://t.co/VwHBdZiNZg
20+ students implemented 7 crowd. pipelines from prior research & find…
1/
Our paper (w/ @dapatil211, @apsarathchandar & @strubell) “An Empirical Investigation of the Role of Pre-training in Lifelong Learning” is now officially published in #JMLR (will be presented at #NeurIPS2023 Journal-to-Conference Track)!
Paper 👉 https://t.co/CdtvMyQSfS
🧵👇 (1/n)
Without seeing this sheet in details, I can already figure out wether some of Tsinghua Professors are on the list. @thudcst@thukeg.
You guess what? I am 100 percent right!
Hot take: what if Google Scholar reported two new metrics: (1) median citations per paper and (2) *percent* of papers with 100+ citations?
I computed these metrics for some ~200 senior AI researchers: see https://t.co/IhYTlUS9ll. The top researchers by median citations per paper are Kaiming He (699), Alec Radford (440), and Alex Krizhevsky (265). (If you want your name removed, DM me).
The reasoning is that I'm sure I'm not the only one frustrated these days by (1) the large volume of papers, and (2) the low quality/rigor of papers. And in terms of incentives, we don't really discourage this. For example, Google Scholar reports total citations, h-index, and i10-index (# papers with >10 citations). None of these penalize low-quality work---even a sloppy, incremental paper that you self-cite later can boost total citations and i10-index.
Median citations per paper and percent of papers with 100+ citations would encourage fewer, high-quality papers. Paying attention to these metrics could be quite revealing:
- Will people accept invitations to be on random papers? That could decrease median citations per paper.
- Will people spend time trying to publish incremental results? That might drag down percent of papers with 100+ citations.
- Some researchers might have hundreds but only 5% of papers with 100+ citations, or very low median citation count. Is this the type of work you want to be doing? (Not a rhetorical question, but a serious one. Some people prefer to work on narrow topics, which will by nature get less citations. That is perfectly fine.)
Incentives matter, and we all should think carefully about how the incentives we (sometimes silently) set drive the community that develops. And yes, these metrics also aren't perfect. To name some caveats:
1. These metrics shouldn't be used in isolation (someone might only have 1 paper, so # total citations is also important). But I definitely think they're interesting and at least better than i10-index.
2. Some citations should be merged on Google Scholar, since duplicated citations penalizes these scores. In the spreadsheet, I removed patents and papers with 0 citations since those were likely to be either very new or duplicates, but I didn't do any deduplication.
3. Similar to total citations and h-index, these metrics will differ per field. E.g., percent of papers with 100+ citations is probably on average higher for AI than for theoretical math, and it's important to note that.
4. These metrics also differ between companies versus universities. At companies, researchers are expected to have experience and each paper is written with impact in mind, so these metrics should be higher. At universities, professors are expected to train students---writing a paper is a good intellectual exercise for the student even without the impact, so it is natural to have lower metrics compared to researchers at companies.
(Of course people are more than their Google Scholar profile, their SAT scores, or how much money they make. I'm not trying to make things more metrics-focused. Google Scholar already exists and surfaces metrics, BTW. I'm just advocating for different metrics).