FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Site Reliability Engineer
Intermedia Cloud Communications. Run and improve production environments supporting AI workloads, data pipelines, analytics applications, and customer-facing services .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing production environments for AI workloads and data pipelines, with a strong focus on automation, reliability, and performance optimization. Proficient in cloud infrastructure, CI/CD practices, and data quality management to ensure seamless operation of analytics applications and services.
Highest-signal resume keywords
Cloud Infrastructure ManagementCI/CD AutomationData Quality ManagementProduction Operations ExperienceTroubleshooting Across Distributed Systems
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Production OperationsSystems EngineeringSREDevOpsData Processing TechnologiesKubernetesKafkaSparkAirflowAutomated Testing
Soft Skills
Analytical Problem-SolvingCross-Functional CommunicationProactive Reliability Focus
Tools & Technologies
ContainersData WarehousesCloud ServicesMetrics MonitoringLogs AnalysisTraces CorrelationService-Level Indicators
Industry Keywords
AI WorkloadsData PipelinesAnalytics ApplicationsOperational MonitoringIncident Response
Tech Stack
Tools & technologiesAirflowCloudDistributed SystemsKafkaKubernetesLinuxSpark
About the role
Key responsibilities & impact- Run and improve production environments supporting AI workloads, data pipelines, analytics applications, and customer-facing services
- Build software and automation for cloud infrastructure, data platforms, model-serving infrastructure, and application services
- Define and measure SLIs, SLOs, and error budgets for AI and analytics services
- Build end-to-end observability correlating metrics, logs, traces, data-quality signals, AI-service performance, and customer impact
- Monitor and optimize reliability, performance, capacity, and cost of batch and streaming workloads, analytics queries, and inference services
- Partner with data engineering and machine-learning teams to productionize ingestion, transformation, feature, training, deployment, and reporting workflows
- Automate CI/CD and production-readiness checks for pipelines, model and prompt releases, schema changes, analytics applications, and dashboards
- Detect and resolve data-quality incidents involving missing, stale, delayed, or anomalous data, schema drift, and broken lineage or dependencies
- Design and test graceful degradation, dependency isolation, retry and fallback patterns, and recovery procedures
- Improve reliability of AI-powered Voice and Unified Communications capabilities including speech recognition, transcription, summarization, intelligent routing, conversational assistance, and text-to-speech
- Establish operational monitoring for model and AI-service behavior, including latency, throughput, error rates, drift indicators, and output quality
- Plan capacity and run performance, load, and resilience tests
- Lead incident response and post-incident improvement
- Support secure and reliable access to cloud storage, processing, and query services
- Reduce operational toil through platform tooling, runbooks, self-service automation, and operational standards
Requirements
What you’ll need- Bachelor's degree in computer science, data engineering, software engineering, or another technical or scientific discipline, or equivalent practical experience
- 4-7 years of experience in production operations, systems engineering, SRE or DevOps, software deployment, and maintenance of distributed production systems
- Experience with cloud infrastructure, containers, Kubernetes, distributed systems, and scalable compute and storage services
- Experience operating data processing, orchestration, storage, or analytics technologies such as Kafka, Spark, Airflow, dbt, data warehouses, or comparable cloud services
- Ability to use metrics, logs, traces, data-quality checks, freshness indicators, lineage, and service-level indicators to diagnose complex production issues
- Experience with CI/CD, DataOps or MLOps practices, automated testing, controlled rollout, and rollback of data and AI service changes
- Strong troubleshooting skills across Linux, applications, networks, APIs, data pipelines, and distributed service dependencies, with attention to security and access controls
- Strong analytical problem-solving and cross-functional communication skills, with a proactive approach to reliability, performance, and continuous improvement
- Must be located in Portugal
Benefits
Comp & perks- Equal opportunity employment practices
- Reasonable accommodations for identified disabilities or other limitations as required by applicable laws
- Diversity and inclusion commitment
- Promotion-from-within opportunities
- Primarily remote work with occasional office visits for collaboration and teamwork