FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting and implementing automated test frameworks within Databricks using PySpark and Python, with a strong focus on data quality and validation techniques. Proficient in integrating automated tests into CI/CD pipelines and mentoring team members on QA standards.
Highest-signal resume keywords
Databricks Test Automation ArchitecturePython MasteryPySpark DataFrame ProcessingSQL Query OptimizationData Quality Validation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Automated Test FrameworksData Quality TestingBatch and Streaming Data ValidationAdvanced SQL QueriesBig-Data Validation LibrariesData Pipeline IntegrityPerformance TestingScalability TestingData Drift AnalysisSchema Evolution
Soft Skills
MentoringLeadershipCollaboration
Tools & Technologies
DatabricksAzure DevOpsGitHub ActionsDatabricks WorkflowsAirflowGreat ExpectationsPytestDelta Live TablesAWSGCP
Certifications & Qualifications
Databricks Certified Data Engineer ProfessionalDatabricks Certified Machine Learning Professional
Industry Keywords
Data EngineeringData QualityDataOpsEvent-Streaming ArchitecturesData Observability
Tech Stack
Tools & technologiesAirflowAWSAzureGoogle Cloud PlatformKafkaPySparkPythonSparkSQLUnity
About the role
Key responsibilities & impact- Architect, build, and scale automated test frameworks from scratch natively within Databricks using PySpark, Python, and SQL
- Design automated assertions for Delta Lake tables, including data drift, schema evolution, and historical data validation via time-travel functions
- Validate large-scale batch and real-time streaming data pipelines using Structured Streaming, ensuring source-to-target integrity
- Programmatically verify data lineage, audit logs, and access controls implemented via Databricks Unity Catalog
- Lead integration of automated data quality tests into enterprise CI/CD pipelines using Azure DevOps, GitHub Actions, Databricks Workflows, APIs, or Airflow
- Act as the subject matter expert for data quality, mentor junior team members, establish QA standards, and advocate for data quality principles
- Design and execute automated performance and scalability tests on Spark jobs, large clusters, and complex query optimizations
Requirements
What you’ll need- Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related quantitative field
- 8+ years of experience in data engineering, data QA, or software development engineering in test (SDET)
- At least 2+ years of dedicated experience architecting test automation in Databricks
- Mastery of Python and PySpark (DataFrames and SQL APIs) for processing and profiling large datasets
- Deep expertise in writing advanced SQL queries, optimization techniques, and understanding Spark query execution plans
- Hands-on mastery of big-data validation libraries, such as Great Expectations, pytest, and Delta Live Tables expectations
- Strong operational knowledge of Databricks deployment on AWS, Azure, or GCP
- Preferred: Databricks Certified Data Engineer Professional or Databricks Certified Machine Learning Professional
- Preferred: Experience validating real-time event-streaming architectures such as Kafka, Event Hubs, or Kinesis
- Preferred: Solid understanding of DataOps culture, testing infrastructure as code, and data observability principles
