We will continue to keep the model and dataset open so that the community can evaluate, use, and contribute to them.
Demo: https://t.co/cjpftudm7l
Model: https://t.co/B1H5R4Dahp
Dataset: https://t.co/BPVHAdwDEB
After a period of improvement and reflecting on the limitations of our first version, we are excited to release Meddies PII v2, a new version of our multilingual medical de-identification model.
For this version, we changed the model architecture and training approach to achieve three major improvements.
- Approximately 5x faster inference on CPU compared with v1.
- 2x less compute required during training.
- Significantly improved model quality. Meddies PII v2 achieves a macro F1 score of 0.833. Under the same evaluation setup, Meddies PII v2 also outperforms the models we compared against, including OpenAI's model.
However, we also want to be clear that Meddies PII v2 is not perfect.
Medical de-identification is a difficult problem. The goal is not simply to remove as much information as possible, but to minimize the risk of missing information that could identify a patient while preserving the information needed for downstream medical analysis and reasoning.
Therefore, when used in real-world environments, Meddies PII v2 should not be treated as the only layer of security. Additional verification and safeguards are still necessary to reduce the risk of personal information being missed.