FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Expertise in developing and deploying NLP systems for entity resolution and data quality analysis, with a strong focus on AML risk patterns and compliance. Proficient in Python and major ML libraries, capable of translating complex model performance into business language for diverse audiences.
Highest-signal resume keywords
NLP Pipeline DevelopmentEntity ResolutionAML Risk AnalysisPython ProficiencySQL Proficiency
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
NLPEntity ExtractionNamed Entity RecognitionRecord LinkageMachine LearningData Quality AnalysisFeature EngineeringGraph-Based MethodsMultilingual NLPStatistical Analysis
Soft Skills
Excellent CommunicationMentoringCollaboration
Tools & Technologies
PyTorchSpaCyHuggingFace TransformersLangChainLangGraph
Certifications & Qualifications
Master's or PhD in Computer ScienceComputational LinguisticsStatisticsApplied Mathematics
Industry Keywords
AMLSanctions ScreeningAdverse MediaFinancial Crime DetectionData Science
Tech Stack
Tools & technologiesPythonPyTorchSQL
About the role
Key responsibilities & impact- Improve the quality, coverage, and freshness of Watchlist data through next-generation ingestion pipelines
- Design and execute data quality analysis pipelines to identify anomalies, evaluate dataset health, and ensure high-fidelity model-training inputs
- Apply NLP and AI to classify and enrich raw source data into normalized schemas
- Extract structured entity attributes from unstructured sanctions, PEP, adverse media, and enforcement sources
- Expand multilingual capabilities across Latin and non-Latin scripts
- Build NLP systems for entity resolution, deduplication, alias resolution, and canonical profiles
- Develop approaches for evolving entity profiles, names, aliases, and sanctions status
- Measure and benchmark entity resolution quality and improve coverage and accuracy
- Design and scale real-time name-matching and identity-classification models
- Build multi-signal risk scoring using name similarity, entity type, geography, list type, and other attributes
- Maintain benchmarking frameworks, golden datasets, and regression tests
- Build analytics for customer screening-threshold tuning
- Develop backtesting and counterfactual analysis capabilities
- Design evaluation frameworks for AI-powered autonomous decision systems and monitor production drift
- Develop mathematical analysis and feature engineering for AML risk patterns in transaction data and payment message fields
- Maintain AML taxonomy and risk signal library
- Apply graph-based methods to identify indirect sanctions-risk exposure
- Lead technical initiatives across Watchlist Data Science
- Collaborate with Product and Engineering to bring research into production systems
- Prototype and deploy advances in NLP, large language models, and entity resolution
- Mentor peers and contribute to technical rigor and continuous improvement
Requirements
What you’ll need- Master's or PhD in Computer Science, Computational Linguistics, Statistics, Applied Mathematics, or a related field; or equivalent professional experience
- 7+ years of experience in data science or machine learning, with meaningful work in NLP, entity resolution, or information extraction
- Experience in AML, sanctions screening, adverse media, or financial crime detection is strongly preferred
- Hands-on experience building and deploying NLP pipelines for entity extraction, named entity recognition, and record linkage at production scale
- Familiarity with multilingual NLP and non-Latin script processing is a strong plus
- Experience with LLMs and agentic AI frameworks such as LangChain/LangGraph is a plus
- Strong proficiency in Python and major ML libraries including PyTorch, spaCy, and HuggingFace Transformers
- Strong SQL proficiency and experience with large-scale data pipelines and production ML systems
- Excellent communication skills, with ability to translate model performance tradeoffs into compliance and business language for non-technical audiences
- No visa sponsorship available
- Must be located in one of the talent hubs: New York, San Francisco, Seattle, or Miami
- Must be eligible to work in the United States indefinitely without visa sponsorship
- Must reside within 45 miles of a Socure talent hub
- Cannot reside in Delaware, Hawaii, Iowa, Kentucky, Mississippi, Nebraska, New Mexico, South Dakota, Vermont, West Virginia, or Wyoming
Benefits
Comp & perks- Equity
- Annual bonus or commission plan
- Accommodation support during the application or hiring process
- Equal employment opportunity protections
