🎉 We are thrilled to share our open-source work "MirrorBench", an extensible framework to evaluate user-proxy agents for human-likeness.
GitHub: https://t.co/OMtdKgSu7U
arXiv: https://t.co/cwFf4z4KCu
#ai#AgenticAi
@alexinexxx I made an animated video explaining this paper few years ago.
How TensorFlow Works? | Tensorflow Architecture | #tensorflow
https://t.co/vafpcftTZ7
📍Our Contribution
✅ Finetuned LLMs on DiaFORGE generated multi-turn dialogue data perform more accurately than top-tier closed-source models in the real production-grade assistant.
✅ We introduce & provide insights into the dynamic evaluation protocol for evaluating LLMs.
🚀 New paper alert!
We’re thrilled to share our paper, “Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky.”
https://t.co/EqVMXHZDbN
📍 What we’re releasing?
✅ End-to-end data-gen + finetune pipeline methodology
✅ An open set of 5000 enterprise APIs + corresponding DiaFORGE-generated & validated dialogues