OpenMed Arabic PII
Detecting and protecting personally identifiable information in Arabic clinical text.
in progress
Problem
Arabic medical text is full of personal data, and you cannot safely use it for downstream AI without first finding and protecting that PII. Arabic morphology makes this harder than the English-centric tools assume.
Approach
- Identifies PII entities in Arabic clinical text - names, identifiers and other sensitive fields.
- Accounts for Arabic-specific morphology and orthography that trip up off-the-shelf detectors.
- Aims to make Arabic medical data safe to use in downstream AI workflows.
Tech
Arabic NLPPII detectionHealthcarePython
Status
In progress. Described honestly as ongoing work.
Links
No public code link - described honestly below.