FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior AI Data Engineer
Strategic Systems International. Design and operate the data foundation supporting AI systems .
Tech Stack
Tools & technologiesAirflowAWSAzureGoogle Cloud PlatformJavaKafkaKubernetesPythonRustScalaSparkSQLTypeScriptVaultGo
About the role
Key responsibilities & impact- Design and operate the data foundation supporting AI systems
- Own movement, modeling, and quality of data from source systems through the warehouse into retrieval and feature layers
- Build and maintain data pipelines powering LLM pipelines, agentic workflows, and analytical products
- Model warehouse schemas and make decisions about grain, relationships, and denormalization
- Write production Python and maintain pipelines as versioned, tested, observable software
- Partner with the AI/ML Data Scientist, who owns model behavior and retrieval strategy
- Own the pipeline, schema, and data guarantees while collaborating on the algorithm, prompt, and evaluation boundary
Requirements
What you’ll need- 5-10+ years in software engineering or data engineering, with substantial time in production data platform work
- Demonstrable command of Kimball dimensional modeling, including choosing grain, resolving many-to-many relationships, and denormalization decisions
- Working knowledge of Data Vault, One Big Table, Inman, and tradeoffs against Kimball
- Expert-level SQL: window functions, CTEs, query plan reading, and performance tuning on a columnar warehouse
- Expert-level, production-grade Python: typing, packaging, dependency management, and testing
- Experience applying SOLID and domain-driven design in real systems
- Production experience with Airflow, Prefect, Dagster, or equivalent
- Expert-level experience with AWS, Azure, or GCP, including storage, compute, IAM, networking, and cost management
- Experience with containerization, Kubernetes (EKS/AKS/GKE), and CI/CD
- Experience with lakehouse table formats such as Iceberg, Delta Lake, or Hudi, including compaction, snapshot expiry, schema evolution, and partition evolution
- Preferred: experience building data layers beneath production RAG systems, including hybrid search infrastructure and index freshness guarantees
- Preferred: experience with Kafka, Kinesis, Flink, or Spark Structured Streaming
- Preferred: dbt or equivalent transformation and testing framework
- Preferred: data quality tooling such as Great Expectations or Soda, and catalog/lineage platforms
- Preferred: familiarity with MLflow, Weights & Biases, and feature stores
- Preferred: working knowledge of a second language: TypeScript, Java, Go, Scala, or Rust
- Preferred: experience with AI security, governance, and compliance frameworks
- Preferred: open-source contributions to data or AI infrastructure projects