We assembled prominent benchmarking environments covering:
→ Software engineering
→ Customer support
→ Personal assistance
→ And more
All peer-reviewed by the research community. No vibes. Just benchmarks. And yes, more benchmarks are needed. This is just the start.
Working on LLM-Based Agents? Want to understand how to evaluate them effectively?
Check out our latest survey on LLM-based agent evaluation https://t.co/MOwQVb8v7h
Survey on Evaluation of LLM-based Agents 🤖
Our paper is the first to provide a comprehensive overview of LLM-based agent evaluation 📜
Paper: https://t.co/43ByXGXLkQ
@omarsar0 great initiatives! We are running this year a new shared task of automatic generation of long summaries for NLP/AI papers as part of SDP workshop@EMNLP 2020 https://t.co/n0NcccoBwT, and would be great if we can utilize (and credit) your summaries as part of training data. @SDProc