FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and maintaining data pipelines using SparkSQL and PySpark, with a strong focus on building ETL processes and data models at scale. Proficient in SQL, Python, and data orchestration tools, while ensuring high-quality data delivery for analytics and model training.
Highest-signal resume keywords
Data Pipeline DesignETL Pipeline DevelopmentSparkSQL and PySparkData Orchestration ToolsSQL and Python Proficiency
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data ModelingETL DevelopmentSparkSQLPySparkSQLPythonREST APIsVersion Control (Git/GitHub)Time Series DataData Transformation
Soft Skills
CollaborationProblem-SolvingCultural Principles Advocacy
Tools & Technologies
AirflowDagsterPrefectDatabricksDelta LakesAWSGCPAzure
Industry Keywords
Data EngineeringIoT DataUnstructured DataCausal InferenceModel TrainingDashboardingClient-Side Signals
Tech Stack
Tools & technologiesAirflowAWSAzureCloudETLGoogle Cloud PlatformIoTPySparkPythonSparkSQL
About the role
Key responsibilities & impact- Design and maintain data pipelines, primarily using SparkSQL and PySpark, within the central data lake
- Ingest and transform source data from IoT devices and software products into the core data model
- Build and maintain highly reliable computed tables using unstructured data, Samsara sensor and product data, and customer metadata
- Access, manipulate, and integrate external datasets with internal data
- Deliver high-quality data with strong uptime and reliability, including customer-facing datasets
- Collaborate with Data Science & Analytics, AI/ML, and other Data Engineers
- Provide high-quality data for causal inference, model training, and dashboarding
- Champion and embed Samsara’s cultural principles as the company scales globally
Requirements
What you’ll need- BA / MS degree in Computer Science, Statistics, or a related discipline
- 4+ years experience in a data engineering-focused role
- Demonstrated experience in designing data models at scale
- Proficiency in building ETL pipelines to handle large volumes of data
- Experience with Spark-based data platforms
- Strong command of at least one data orchestration tool, such as Airflow, Dagster, or Prefect
- Expertise in SQL, Python, and working with REST APIs
- Familiarity with software engineering fundamentals and reading backend development code
- Experience with version control systems such as Git/GitHub
- Relocation assistance will not be provided
- Must maintain the legal right to work at the company and in the specified work location
- Familiarity with time series data and late-arriving data
- Knowledge of Databricks, Delta Lakes, and Dagster
- Previous experience working in a public cloud, such as AWS, GCP, or Azure
- Exposure working on a data model for a product’s first-party data
- Exposure to complex data, including ML outputs and/or client-side signals
Benefits
Comp & perks- Initial RSU grant with no vesting cliff
- Ongoing equity refresh opportunities tied to performance
- Performance-based bonus/variable pay
- Equity for eligible roles
- Flexible, employee-led remote model
- Professional development stipend
- Comprehensive health plans
- Parental leave plans
- Above-market total compensation opportunities
