FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Data Engineer – Platform
NVIDIA. Own systems end to end from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating production software and data platforms, with a strong focus on Python, SQL, and distributed systems. Capable of leading architectural decisions, ensuring secure platform development, and implementing effective monitoring and incident response strategies.
Highest-signal resume keywords
Production Proficiency In PythonDistributed Data ProcessingDatabase Architecture And Operation At ScaleCloud Platform Systems ExperienceSecure Platform Development
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonSQLDistributed SystemsETLChange-Data-CaptureStreaming SystemsData ModelingDebuggingIncident ResponseAI Agents
Soft Skills
Effective CommunicationMentoringProblem SolvingAdaptabilityCollaboration
Tools & Technologies
DatabricksApache SparkKafkaAWSAzureGCPKubernetesElasticsearchDelta LakeSlurm
Industry Keywords
Cloud ServicesData PlatformsOperational TelemetryProduction SoftwareFleet-Scale Telemetry
Tech Stack
Tools & technologiesApacheAWSAzureCloudDistributed SystemsElasticSearchETLGoogle Cloud PlatformKafkaKubernetesPySparkPythonSparkSQLUnity
About the role
Key responsibilities & impact- Own systems end to end from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support
- Design and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry
- Build shared libraries, workflow and DAG abstractions, deployment tooling, data contracts, and paved-road platform patterns
- Engineer reliable distributed workloads and diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services
- Build for retries, idempotency, backfills, schema evolution, and partial failure
- Apply least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices
- Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership
- Make trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications
- Lead build reviews, communicate tradeoffs, mentor engineers, and improve architecture, testing, debugging, and operational practices
Requirements
What you’ll need- BS or MS in Computer Science, Engineering, or a related field, or equivalent experience
- 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems
- Production proficiency in Python or another backend or systems language, with ability and willingness to work primarily in Python and SQL
- Hands-on experience in distributed data processing, database architecture and operation at scale, production ETL, change-data-capture, streaming/event-processing systems, backend or cloud-platform systems, or strong SQL and data-modeling
- Ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments
- Experience operating services or pipelines in a cloud or complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response
- Working knowledge of secure platform development, identity and access management, least privilege, secret handling, trust boundaries, and safe multi-environment deployments
- Ability to make architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields
- Track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices
- Experience with AI agents and LLM-supported workflow automation
- Preferred experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, Kafka, change-data capture, event systems, Elasticsearch/OpenSearch, AWS, Azure, GCP, Kubernetes, Slurm, compute clusters, GPU infrastructure, fleet-scale telemetry, agentic systems, LLM-enabled workflow automation, harness engineering, or AI-agent evaluation and operational tooling
Benefits
Comp & perks- Equity
- Benefits 📊 Check your resume score for this job Improve your chances of getting an interview by checking your resume score before you apply. Check Resume Score