Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Firmable

Lead Data Engineer

Firmable

. Architect and own the end-to-end extraction and ETL pipeline transforming unstructured web data into a B2B dataset across 13 markets .

Posted 9/23/2026full-timeRemote • IndiaSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in architecting and implementing end-to-end extraction and ETL pipelines, with a strong focus on production LLM infrastructure and data quality frameworks. Proficient in handling complex data challenges, including anti-bot defenses and schema drift, while ensuring compliance with data privacy regulations.

Highest-signal resume keywords
Expert PythonAdvanced SQLExtensive Airflow ExperienceProduction LLMs in Extraction PipelinesCloud Data Platforms

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
ETL Pipeline DevelopmentWeb Data ExtractionData NormalisationData DeduplicationData ValidationIncremental ProcessingAgentic WorkflowsAutomated TestingAnomaly DetectionComplex Transformations
Soft Skills
Strong JudgmentProduct MindsetCollaboration
Tools & Technologies
AirflowAWS LambdaAWS S3AWS ECSAWS GlueSnowflakeRedshiftClaude CodeCursorBraintrust
Industry Keywords
B2B DataData PrivacyGDPRCCPAData Quality Frameworks

Tech Stack

Tools & technologies
AirflowAmazon RedshiftAWSCloudETLJavaScriptPythonSQL

About the role

Key responsibilities & impact
  • Architect and own the end-to-end extraction and ETL pipeline transforming unstructured web data into a B2B dataset across 13 markets
  • Set architectural standards for extractor patterns, proxy strategy, LLM infrastructure, agentic escalation workflows, coverage, schema, and accuracy
  • Design extraction, normalisation, deduplication, validation, and load processes
  • Own cost, performance, reliability, scheduling, incremental processing, and recovery design
  • Build extractor frameworks and ship difficult extractors handling anti-bot defences, JS-heavy rendering, schema drift, and low-quality structure
  • Build agentic extraction pipelines with rule-based triage, LLM escalation, structured-output validation, retries, and human-review queues
  • Establish production LLM infrastructure with versioned prompts, labelled evaluation sets, precision/recall measurement, rollback, traces, and drift detection
  • Develop evaluation and observability scaffolding and model-choice playbooks
  • Create versioned SKILL.md specifications and orchestration patterns using Airflow or equivalent
  • Provide technical input to data platform, product, and analytics teams
  • Deliver approximately 80% hands-on architecture/reference implementation and 20% cross-functional technical input

Requirements

What you’ll need
  • 6–10 years shipping production extraction, ETL, or data pipeline systems in business-critical environments
  • Deep web extraction at scale, including anti-bot defences, proxy architecture, JS rendering, schema drift, and recovery
  • Production LLMs in extraction pipelines, including structured outputs, versioned prompts, labelled eval sets, logged traces, precision/recall judges, and vendor-model drift detection
  • Production agent workflows and tool-calling pipelines
  • Experience writing SKILL.md specifications
  • Strong judgment regarding deterministic rules versus LLMs
  • Expert Python, including concurrency and scale
  • Advanced SQL for complex transformations and performance optimisation
  • Extensive Airflow or equivalent experience
  • Experience shipping with agentic IDEs such as Claude Code or Cursor
  • Architecture judgment and product mindset
  • Current hands-on use of AI tools, structured outputs, evals, traces, and LLM tracing
  • Cloud data platforms such as Snowflake or Redshift highly valued
  • AWS pipeline deployment experience with Lambda, S3, ECS, or Glue highly valued
  • Vector databases, embeddings, or retrieval patterns highly valued
  • Eval frameworks such as Braintrust, Promptfoo, or Inspect highly valued
  • Data-quality frameworks with automated testing and anomaly detection highly valued
  • B2B data experience highly valued
  • Data privacy and compliance knowledge, including GDPR and CCPA, highly valued
  • Startup or scaleup experience highly valued

Benefits

Comp & perks
  • No fixed hours
  • Full autonomy on architecture choices