New preprint w/ (co-first author) @royjamesadams and @suchisaria: "Evaluating Models Robustness Under Dataset Shift" https://t.co/ykOu7G9EQE
How can we evaluate *ahead of time* whether or not a model's performance will generalize from training to deployment? 1/
Our work is finally published in @npjDigitalMed! 🚀 In https://t.co/QAB9D3UJoc , we show that AI is capable of scaling up analyses of medical device regulatory documents, and has a lot to offer to patients, regulators, researchers, and vendors
@roydanroy Wow, huge win for DeepMind! My first ML conference was UAI 2017 in Sydney. I’ll never forget your talk on the “nonvacuous generalization bounds” paper. Congrats!
What if a platform existed to help doctors make faster more accurate diagnoses?
INDEX intends to create a high-quality medical imaging platform that does not currently exist to benefit the entire imaging ecosystem. https://t.co/IVfovONQiA
@beenwrekt @jjcherian @iswgibbs If you define a set of groups with respect to covariate values (e.g., demographic info), then if you shift the prevalence of the groups you will see a corresponding covariate shift . Not sure if the reverse direction works the same
It turns out monitoring the performance of ML models is hard, especially in settings like healthcare where model predictions affect user decisions
If you're at NeurIPS check out Jean's talk and we'd love to hear your thoughts re monitoring
Our framework brings together ideas from causal inference and statistical process control. Please drop by if you're interested! Preprint here: https://t.co/ROvBdBudOy It was wonderful working on this with coauthors from UCSF and FDA! @pirracchior @_asubbaswamy@harvineet_singh
Our framework brings together ideas from causal inference and statistical process control. Please drop by if you're interested! Preprint here: https://t.co/ROvBdBudOy It was wonderful working on this with coauthors from UCSF and FDA! @pirracchior @_asubbaswamy@harvineet_singh
I'm talking at today's NeurIPS #RegulatableML workshop about why designing monitoring systems for ML-based medical devices is complicated, why we need a systematic framework for comparing all the different monitoring strategies, and our first steps to developing such a framework.
How do people manage to keep track of ML papers? This is not a request for support in my current state of bewilderment - I'm genuinely asking what strategies seem to work to read (or "read") what appear to be 100s of papers per day.
@vitaliikl@St_Megas@rlmcelreath Assuming faithfulness, the collider is distinguishable from the pipe and fork. As the image states, in the collider case X and Y are marginally independent while for the other two X and Y are marginally associated. See, e.g., https://t.co/DaAaUNJRaN
@roydanroy I'm unfortunately having to do this... I'm about to watch the best paper in my batch get rejected because of a spiteful reviewer and uninformed AC, and the worst paper get accepted because of uninformed reviewers and uninvolved AC
Why do people care about neurips papers again?
@EugeneVinitsky @ccanonne_ @shortstein imo this only really happens when papers are shepherded and there are opportunities for revisions. I’ve had much better experiences as an author with journal submissions because of this
Our #UAI2023 paper is on detecting novel subgroups (aka classes/categories) when distribution of non-novel data shifts.
If this sounds interesting, or you’re wondering what it has to do with unicorns, check out our poster/paper/drop a line.
w/ @suchisaria
https://t.co/7nrMfDblfx
If you’re at #ICML2023 join me here: https://t.co/Kzca5T5dCf
Giving an invited talk later today on AI Safety.
Below are some of the papers our lab has written on the topic in last 5 yrs; this talk will give a brief overview and focus on learning/evaluation.