Overview
ABOUT THE ROLE
Full job description
ABOUT THE ROLE Build privacy and anonymization systems that help make sensitive real-world data safe and useful for AI training. You will develop end-to-end methods to protect sensitive information while preserving the structure and signal needed for downstream training, evaluation, and synthetic data workflows. WHAT YOU'LL DO
- Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, and tailor transformations to data types and use cases.
- Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
- Create production pipelines that anonymize data before it enters processing, training, evaluation, or synthetic data workflows.
- Develop evaluation frameworks for privacy risk and retained utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
- Design robust systems that handle new sources, schema drift, unusual formats, and sensitive information in unexpected fields.
- Partner with engineering, research, operations, and customers to turn privacy requirements into practical safeguards. WHAT WE'RE LOOKING FOR
- At least 2 years of experience building production data or ML systems in Python, with strong proficiency in the language.
- Hands-on experience with PII detection, removal, or anonymization, including transformations that preserve useful data characteristics while hiding underlying information.
- Experience with information extraction, named-entity recognition, classification, or related methods for finding rare or sensitive content.
- Ability to build end-to-end data pipelines and compare approaches across recall, precision, latency, cost, and downstream utility.
- Understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation.
- Experience handling schema drift and edge cases; work with sensitive data or privacy-enhancing techniques such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is valuable.
- Experience with low-latency or high-throughput ML inference and data processing is beneficial. COMPENSATION & BENEFITS Salary range: $130,000 to $225,000 annually. Visa sponsorship is available. LOCATION On-site in San Francisco, California, United States.
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →