Excited to share our work on improving safety evaluation of Clinical Voice AI transcription, using @DSPyOSS and GEPA to build a clinician-level LLM Judge!
🚨 ASR errors in clinical dialogue can be dangerous, and WER doesn’t know it. Today we release “WER is Unaware”.
Using DSPy + GEPA, we optimise an LLM Judge that reaches clinician-level performance at detecting safety risks.
🔗 https://t.co/RZOKEa4JWq
📄 https://t.co/h5YxKghv3W