Clera
Senior Research Engineer, Privacy and Anonymization
San Francisco · Posted 3h ago
Job Description
About the Role
Build privacy and anonymization systems that help make sensitive real-world data safe and useful for AI training. You will develop end-to-end methods to protect sensitive information while preserving the structure and signal needed for downstream training, evaluation, and synthetic data workflows.
What You'll Do
Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, and tailor transformations to data types and use cases.
Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
Create production pipelines that anonymize data before it enters processing, training, evaluation, or synthetic data workflows.
Develop evaluation frameworks for privacy risk and retained utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
Design robust systems that handle new sources, schema drift, unusual formats, and sensitive information in unexpected fields.
Partner with engineering, research, operations, and customers to turn privacy requirements into practical safeguards.
What We're Looking For
At least 2 years of experience building production data or ML systems in Python, with strong proficiency in the language.
Hands-on experience with PII detection, removal, or anonymization, including transformations that preserve useful data characteristics while hiding underlying information.
Experience with information extraction, named-entity recognition, classification, or related methods for finding rare or sensitive content.
Ability to build end-to-end data pipelines and compare approaches across recall, precision, latency, cost, and downstream utility.
Understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation.
Experience handling schema drift and edge cases; work with sensitive data or privacy-enhancing techniques such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is valuable.
Experience with low-latency or high-throughput ML inference and data processing is beneficial.
Compensation & Benefits
Salary range: $130,000 to $225,000 annually. Visa sponsorship is available.
Location
On-site in San Francisco, California, United States.