๐ค How do frontier models perform on core investment banking tasks?
Introducing Playgent-IB-Bench, an investment banking benchmark we built with former M&A bankers from bulge-bracket and elite-boutique firms.
Writeup: https://t.co/hQGONrOhht
We then built an RL harness on @PrimeIntellect that mimics the human setup. Agents work in a filesystem with 58 tools across:
- working with spreadsheets
- creating presentations
- reading PDFs
- editing Word documents
The environment enables learning a policy for long-horizon financial tasks using the same tools as human bankers.