Thrilled to have attended my first NeurIPS in December and present our work on constructing complete datasets to address real-world data scarcity for fairness testing! 🍁🌟
The CDT is proud to sponsor and support our PhD students as they present their research at #neurips2024 showcasing their hard work and academic excellence on an international stage! Read about their experience here https://t.co/H5HRuabaNT
LLM evaluation is a confusing mess. Standardized eval frameworks have helped, but there's a long way to go.
I think we should promote evaluation as a specialized activity carried out by impartial third parties who are not also developers. Of course developers will evaluate their own models and agents, but third-party evaluations should be considered more credible. For example I think HELM results should be much more prominent in model comparisons. (I have no affiliation with HELM.)
Journalists reporting on AI releases should prioritize third party evals over the developer's claims when possible. Or they should explicitly note that third party evals are not yet available.
Separating testing & evaluation from building is the norm in most engineering fields. It's time for AI to grow up.
We start with a busy morning of talks in AI bias, AI and warfare and AI regulation. These topics always lead to some insightful discussions over break! @EMAI_eu
Thrilled to see our work on Gaussian Splatting gaining some attention! 🤩 A huge thank you to @janusch_patas for the shout-out.
The Gaussian Pancakes paper is available on https://t.co/qFrJQziocK
As AI systems get deployed broadly, we are going to need a greater set of organizations to test out their capabilities and potential misuses or safety issues. Here’s a @AnthropicAI post that discusses how we see third-party testing relating to AI policy https://t.co/aGcTfYKzJu
Launching our new #steeringAI podcast!🎙️
Join Reuben Adams for a deep dive into the impact and risks of AI with leading experts Prof. Marc Deisenroth, Prof. Lewis Griffin and @RoyalHolloway's Prof. Chris Watkins . A must-listen for any #AI enthusiast.
https://t.co/q1t0POQduQ
Latent Attention (Latte) is a new linear time transformer model. Unlike sparse methods, all tokens play a role in the attention calculation and Latte can be used in both causal and non-causal settings. We welcome comments!
https://t.co/GXJILQ7geQ
It has been an incredible two years of Postdoc experience at @ucl, working together with the members of @ucl_wi_group. Many thanks to Emine for her continuous support. Also excited to announce that, from March 1st, I will join @SheffieldNLP as a Lecturer. The next chapter.🥳
COVID-ICU was tough, but one thing that helped me was working on a project with the amazing Emily Holmes & colleagues to design/deliver a @wellcometrust funded clinical trial of an intervention that might help ICU staff.
https://t.co/a270hafhrr
MY FIRST PAPER 🎉 An exciting clinical trial to reduce intrusive memories of trauma 🧠 We used sequential Bayesian analysis to optimise a brief digital intervention, and showed positive results for NHS ICU Staff who experienced work-related trauma👩⚕️
(https://t.co/W2jS1CIVCB)
As I was privately babbling out my own diatribe on the FLI letter yesterday, I was honored to be pinged by @emilymbender and @timnitGebru about joining forces to put out coherent thoughts. Excited to be able to share it today.
https://t.co/WGVGf2S0EH