- Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information and design transformations based on the data type and downstream use case
- Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods
- Build production pipelines that anonymize raw data before it enters downstream processing, training, evaluation, or synthetic data generation workflows
- Create evaluation frameworks that measure privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts
- Design systems that remain robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields
- Work with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards