FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and building ETL/ELT pipelines using Databricks and open-source strategies, with a strong focus on data curation, compliance with LGPD, and integration of databases with APIs. Proficient in automating processes and ensuring data quality within a medallion architecture.
Highest-signal resume keywords
Advanced PythonSQLDatabricksData ModelingLGPD Compliance
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
ETL/ELT Pipeline DesignData CurationVector IndexingData Quality TestingVersioningUnstructured Data ProcessingMedallion ArchitectureEmbedding GenerationIncremental Index UpdatesData Standardization
Tools & Technologies
DatabricksDelta LakePySparkSpark SQLGitCI/CDAzureAWSGCPMLflow
Certifications & Qualifications
Databricks Data Engineer AssociateDatabricks Data Engineer Professional
Industry Keywords
InsuranceFinancial ServicesSensitive Data HandlingUnity CatalogMosaic AIDelta Live TablesLakeflowAuto LoaderTerraformDatabricks Asset Bundles
Tech Stack
Tools & technologiesAWSAzureETLGoogle Cloud PlatformPySparkPythonSparkSQLTerraformUnity
About the role
Key responsibilities & impact- Design and build ETL/ELT pipelines in Databricks or using open-source strategies, ingesting and processing data within a medallion architecture on Delta Lake
- Ingest and standardize PDF documents and images with metadata
- Perform data curation and anonymization in compliance with Brazil’s LGPD, generating datasets for RAG, few-shot learning, and testing
- Build and operate vector indexing, including chunking, embedding generation, incremental index updates, and versioning
- Prepare versioned datasets (golden sets) for accuracy evaluation and benchmarking
- Structure feedback-loop persistence and feed the Quality and SLA Dashboard
- Integrate databases with inference services and APIs to other systems, taking performance, cost, monitoring, and alerting into account
- Automate processes using Workflows/Jobs, CI/CD, and data quality tests
Requirements
What you’ll need- Advanced Python and SQL
- Data modeling and medallion architecture / Lakehouse
- Experience with unstructured data pipelines (PDFs, images, and text) and preparing data for RAG
- Familiarity with vector databases and embeddings (Databricks Vector Search, pgvector, or similar)
- Data quality, data testing, and versioning
- Hands-on Databricks experience: Delta Lake, Workflows/Jobs, notebooks, PySpark, and Spark SQL
- Git and CI/CD
- Experience in a cloud environment (Azure, AWS, or GCP)
- Knowledge of LGPD and sensitive data handling
- Unity Catalog in a governed environment (preferred)
- MLflow, Databricks Model Serving, and Mosaic AI (preferred)
- Delta Live Tables / Lakeflow and Auto Loader (preferred)
- Experience in the insurance or financial services industry (preferred)
- Terraform or Databricks Asset Bundles (preferred)
- Databricks Data Engineer Associate or Professional certification (preferred)
Benefits
Comp & perks- Anywhere Office: geographic flexibility to work from the location that makes the most sense for you, depending on the project
- Discount partnerships
- Employee referral program with rewards
- Birthday day off
- TotalPass/Gympass
- Health and wellness initiatives
- Opportunities to relax at our offices, with beanbags, a pool table, video games, a lounge, a fully equipped kitchen, and coffee
