Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
hud (YC W25)

Research Engineer, Privacy and Anonymization

hud (YC W25)

. Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information and design transformations based on data type and downstream use case .

Posted 9/23/2026full-timeRemote • California • United States, SingaporeMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building systems for detecting and anonymizing sensitive information, with a strong focus on privacy-enhancing technologies and data processing pipelines. Proficient in Python and experienced in translating privacy requirements into effective technical solutions.

Highest-signal resume keywords
Python ProficiencyData Processing PipelinesPrivacy-Enhancing TechnologiesInformation ExtractionExperimental Analysis

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Information ExtractionNamed-Entity RecognitionClassificationRedactionMaskingPseudonymizationAnonymizationSynthetic DataMachine Learning SystemsData Utility Metrics
Soft Skills
Attention to DetailCollaborationCommunication
Tools & Technologies
Differential PrivacyK-AnonymitySecure AggregationFormat-Preserving Encryption
Industry Keywords
HealthcareFinanceSecurityAdversarial Re-IdentificationSchema Drift

Tech Stack

Tools & technologies
Python

About the role

Key responsibilities & impact
  • Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information and design transformations based on data type and downstream use case
  • Develop and benchmark detection approaches combining rules, statistical models, classifiers, and LLM-based methods
  • Build production pipelines that anonymize raw data before downstream processing, training, evaluation, or synthetic data generation
  • Create evaluation frameworks measuring privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts
  • Design systems robust to new data sources, schema drift, unusual formats, and sensitive information embedded in unexpected fields
  • Work with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards
  • Own the full pipeline for protecting privacy without destroying the structure and signal valuable for AI-agent training
  • Participate in a 2–3 day work trial during the application process

Requirements

What you’ll need
  • Strong proficiency in Python and experience building reliable production data or ML systems
  • Experience with information extraction, named-entity recognition, classification, or related methods for detecting rare or sensitive content
  • Strong experimental instincts and ability to compare approaches across recall, precision, latency, cost, and downstream data utility
  • Understanding of redaction, masking, pseudonymization, anonymization, and synthetic data, and when each is appropriate
  • High attention to detail and ability to reason about subtle leakage paths, edge cases, and adversarial failure modes
  • Experience building data processing pipelines end-to-end without a fully prescribed roadmap
  • Hands-on experience with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption (strong candidates may also have)
  • Experience with sensitive data in healthcare, finance, or security (strong candidates may also have)
  • Experience building low-latency or high-throughput ML inference and data-processing systems (strong candidates may also have)
  • Experience working in unstructured problem spaces and taking ownership from early research through production deployment (strong candidates may also have)
  • Early-stage startup experience and strong communication skills for collaboration across teams and time zones (strong candidates may also have)
  • Ability to work hours that 70–80% overlap with either San Francisco or Singapore time zones
  • Full-time availability

Benefits

Comp & perks
  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
  • Lunch and dinner when you’re in the office (in-office employees)
  • Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
  • Unlimited access to tokens for ChatGPT, Claude Code, Cursor, etc.
  • Equinox membership (US employees)
  • 401k (US employees)
  • Commuter benefits (US employees)
  • Support for relocation and visas for strong full-time candidates to the US or Singapore