Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Oxylabs.io

Data Scientist

Oxylabs.io

. Develop and maintain entity resolution techniques to match and link records across sources with inconsistent identifiers .

Posted 9/15/2026full-timeVilnius • LithuaniaMid-LevelSenior💰 €3,500 - €6,800 per monthWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing and maintaining entity resolution techniques, building and evaluating ML and NLP models, and automating data acquisition workflows. Proficient in ensuring data accuracy and reliability while collaborating with stakeholders to drive insights and improvements.

Highest-signal resume keywords
Data ModelingMachine LearningNatural Language ProcessingPython ProgrammingSQL Proficiency

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Entity ResolutionML Model DevelopmentData ExtractionData ValidationData Enrichment
Soft Skills
Clear CommunicationAttention to DetailCollaborative MindsetSelf-Driven
Tools & Technologies
SparkDagsterDbtSupersetTrino
Industry Keywords
Data ScienceData AnalysisData EngineeringConcept DriftRevenue Insights

Tech Stack

Tools & technologies
PythonSparkSQL

About the role

Key responsibilities & impact
  • Develop and maintain entity resolution techniques to match and link records across sources with inconsistent identifiers
  • Prototype, build, and evaluate LLM and ML-based extraction and inference models that turn unstructured or free-text data into structured outputs
  • Automate scraping logic and source queue generation using data-driven models to streamline and scale source acquisition workflows
  • Build, iterate on, and monitor applied ML models to proactively identify data or concept drift
  • Enrich and enhance datasets with product-ready calculated fields and derived metrics
  • Assist in building automated or ad-hoc QA validation processes to validate model output accuracy, reliability, and consistency
  • Take ownership of the applied ML backlog, including entity resolution, LLM-based structured extraction, and topic modeling
  • Turn unstructured data into structured, product-ready outputs and ensure data accuracy and reliability at scale

Requirements

What you’ll need
  • 4+ years of previous experience as a Data Scientist, Data Analyst, or Data Engineer, with a strong background in data modeling, ML, and NLP
  • Excellent programming skills in Python
  • Proficiency in SQL
  • Hands-on experience with Spark
  • Deep, structural thinking with the ability to communicate clearly and align with various stakeholders
  • Collaborative, self-driven mindset with keen attention to detail and a propensity to dig into deeper layers to inspire improvements
  • Experience with Dagster, dbt, Superset, or Trino (nice to have)
  • Previous experience working closely in a team with data engineers (nice to have)
  • Strong business acumen and understanding of how insights convert to value and unlock new revenue (nice to have)
  • Excellent written and spoken English (nice to have)

Benefits

Comp & perks
  • 40+ internal learning options
  • External conferences
  • Mentorship
  • Year-round knowledge-sharing
  • Private health insurance
  • Psychotherapy
  • On-site well-being consultants
  • 24/7 gym access
  • Wellness app
  • Team events
  • Overseas workation
  • Quarterly team-building budgets
  • Milestone celebrations