Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

Senior Data Engineer – Platform

NVIDIA

. Own systems end to end from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support .

Posted 9/21/2026full-timeRemote • United StatesSenior💰 $140,000 - $270,250 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating production software and data platforms, with a strong focus on Python, SQL, and distributed systems. Capable of leading architectural decisions, ensuring secure platform development, and implementing effective monitoring and incident response strategies.

Highest-signal resume keywords
Production Proficiency In PythonDistributed Data ProcessingDatabase Architecture And Operation At ScaleCloud Platform Systems ExperienceSecure Platform Development

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
PythonSQLDistributed SystemsETLChange-Data-CaptureStreaming SystemsData ModelingDebuggingIncident ResponseAI Agents
Soft Skills
Effective CommunicationMentoringProblem SolvingAdaptabilityCollaboration
Tools & Technologies
DatabricksApache SparkKafkaAWSAzureGCPKubernetesElasticsearchDelta LakeSlurm
Industry Keywords
Cloud ServicesData PlatformsOperational TelemetryProduction SoftwareFleet-Scale Telemetry

Tech Stack

Tools & technologies
ApacheAWSAzureCloudDistributed SystemsElasticSearchETLGoogle Cloud PlatformKafkaKubernetesPySparkPythonSparkSQLUnity

About the role

Key responsibilities & impact
  • Own systems end to end from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support
  • Design and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry
  • Build shared libraries, workflow and DAG abstractions, deployment tooling, data contracts, and paved-road platform patterns
  • Engineer reliable distributed workloads and diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services
  • Build for retries, idempotency, backfills, schema evolution, and partial failure
  • Apply least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices
  • Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership
  • Make trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications
  • Lead build reviews, communicate tradeoffs, mentor engineers, and improve architecture, testing, debugging, and operational practices

Requirements

What you’ll need
  • BS or MS in Computer Science, Engineering, or a related field, or equivalent experience
  • 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems
  • Production proficiency in Python or another backend or systems language, with ability and willingness to work primarily in Python and SQL
  • Hands-on experience in distributed data processing, database architecture and operation at scale, production ETL, change-data-capture, streaming/event-processing systems, backend or cloud-platform systems, or strong SQL and data-modeling
  • Ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments
  • Experience operating services or pipelines in a cloud or complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response
  • Working knowledge of secure platform development, identity and access management, least privilege, secret handling, trust boundaries, and safe multi-environment deployments
  • Ability to make architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields
  • Track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices
  • Experience with AI agents and LLM-supported workflow automation
  • Preferred experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, Kafka, change-data capture, event systems, Elasticsearch/OpenSearch, AWS, Azure, GCP, Kubernetes, Slurm, compute clusters, GPU infrastructure, fleet-scale telemetry, agentic systems, LLM-enabled workflow automation, harness engineering, or AI-agent evaluation and operational tooling

Benefits

Comp & perks
  • Equity
  • Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score