FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Data Scientist
Oxylabs.io. Develop and maintain entity resolution techniques to match and link records across sources with inconsistent identifiers .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in developing and maintaining entity resolution techniques, building and evaluating ML and NLP models, and automating data acquisition workflows. Proficient in ensuring data accuracy and reliability while collaborating with stakeholders to drive insights and improvements.
Highest-signal resume keywords
Data ModelingMachine LearningNatural Language ProcessingPython ProgrammingSQL Proficiency
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Entity ResolutionML Model DevelopmentData ExtractionData ValidationData Enrichment
Soft Skills
Clear CommunicationAttention to DetailCollaborative MindsetSelf-Driven
Tools & Technologies
SparkDagsterDbtSupersetTrino
Industry Keywords
Data ScienceData AnalysisData EngineeringConcept DriftRevenue Insights
Tech Stack
Tools & technologiesPythonSparkSQL
About the role
Key responsibilities & impact- Develop and maintain entity resolution techniques to match and link records across sources with inconsistent identifiers
- Prototype, build, and evaluate LLM and ML-based extraction and inference models that turn unstructured or free-text data into structured outputs
- Automate scraping logic and source queue generation using data-driven models to streamline and scale source acquisition workflows
- Build, iterate on, and monitor applied ML models to proactively identify data or concept drift
- Enrich and enhance datasets with product-ready calculated fields and derived metrics
- Assist in building automated or ad-hoc QA validation processes to validate model output accuracy, reliability, and consistency
- Take ownership of the applied ML backlog, including entity resolution, LLM-based structured extraction, and topic modeling
- Turn unstructured data into structured, product-ready outputs and ensure data accuracy and reliability at scale
Requirements
What you’ll need- 4+ years of previous experience as a Data Scientist, Data Analyst, or Data Engineer, with a strong background in data modeling, ML, and NLP
- Excellent programming skills in Python
- Proficiency in SQL
- Hands-on experience with Spark
- Deep, structural thinking with the ability to communicate clearly and align with various stakeholders
- Collaborative, self-driven mindset with keen attention to detail and a propensity to dig into deeper layers to inspire improvements
- Experience with Dagster, dbt, Superset, or Trino (nice to have)
- Previous experience working closely in a team with data engineers (nice to have)
- Strong business acumen and understanding of how insights convert to value and unlock new revenue (nice to have)
- Excellent written and spoken English (nice to have)
Benefits
Comp & perks- 40+ internal learning options
- External conferences
- Mentorship
- Year-round knowledge-sharing
- Private health insurance
- Psychotherapy
- On-site well-being consultants
- 24/7 gym access
- Wellness app
- Team events
- Overseas workation
- Quarterly team-building budgets
- Milestone celebrations