FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing production infrastructure, particularly with Linux systems, Docker, and PostgreSQL, while ensuring security and compliance with ISO 27001. Capable of improving developer experience and operational resilience through effective incident response and observability practices.
Highest-signal resume keywords
Linux System AdministrationPostgreSQL ManagementInfrastructure as Code (Terraform, Pulumi)CI/CD Pipeline Development (GitHub Actions)Observability Stack Operation (Grafana, Prometheus)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Linux System AdministrationPostgreSQL ManagementDockerInfrastructure as Code (Terraform, Pulumi)CI/CD Pipeline DevelopmentObservability Stack OperationAPI Gateway ManagementIncident ResponseSecurity Practices (IAM, Secrets Management)Performance Analysis
Soft Skills
Fluent English CommunicationReliable Written CommunicationAccountabilityTechnical CuriosityCollaboration
Tools & Technologies
Docker ComposeGitHub ActionsGrafanaPrometheusLokiSentryTraefikKongVLLMHetzner
Certifications & Qualifications
ISO 27001 Compliance
Industry Keywords
SaaSInfrastructure SecurityObservabilityGenAIRAG Systems
Tech Stack
Tools & technologiesDockerGrafanaLinuxPostgresPrometheusTerraform
About the role
Key responsibilities & impact- Take technical ownership of MAIA’s production infrastructure
- Operate and evolve self-administered Linux systems across virtual machines and dedicated servers
- Manage containers, networks, reverse proxies, and API gateways
- Improve Infrastructure as Code, GitHub Actions workflows, deployment processes, automated checks, versioning, and rollback capabilities
- Make deployments safe, repeatable, and easy for engineers to use
- Operate and improve PostgreSQL in production, including performance analysis, connection pooling, capacity planning, backups, and tested restore procedures
- Develop the self-hosted observability stack connecting metrics, logs, traces, and actionable alerts
- Strengthen infrastructure security through IAM, least privilege, secrets management, TLS, vulnerability scanning, and patch management
- Implement technical controls for ISO 27001 with continuously generated, understandable, and auditable evidence
- Shape AI infrastructure by integrating and evaluating model and inference providers
- Evaluate providers based on reliability, latency, throughput, cost, and operational effort
- Explore self-hosted LLM inference, potentially including GPU infrastructure and vLLM
- Improve incident response and operational resilience through root-cause investigation, runbooks, documentation, and lasting improvements
- Improve internal developer experience by reducing manual work and creating clear interfaces and workflows
- Make infrastructure and inference costs transparent and inform build, buy, and hosting decisions
- Assess the existing platform, identify risks, explain options, and take improvements through to reliable production operation
- Collaborate with software engineers, the CTO, company leadership, and distributed team members
- Work as a hands-on senior individual contributor without people management
Requirements
What you’ll need- Several years of experience operating production SaaS systems on Linux servers administered directly by you or your team
- Experience limited to fully managed hyperscaler services is not sufficient
- Production experience with Docker and Docker Compose
- Understanding of reverse proxies or API gateways such as Traefik or Kong
- Practical experience operating PostgreSQL, including personally configured and tested backups and restores, performance analysis, and connection pooling
- Experience building and maintaining CI/CD pipelines using GitHub Actions, GitLab CI, or comparable systems
- Experience managing infrastructure through Terraform, Pulumi, or comparable Infrastructure as Code tooling
- Experience operating an observability stack using Grafana, Loki, Prometheus, Sentry, or comparable technologies
- Understanding of IAM, least privilege, secrets management, vulnerability scanning, and patching
- Ability to design infrastructure reliable for customers and straightforward for engineers to use
- Fluent English communication
- German helpful but not required
- Understanding of how RAG systems work and where they commonly fail in production
- Ability to discuss LLM serving trade-offs, including latency, throughput, reliability, cost, and operational complexity
- Familiarity with models, inference providers, and the wider GenAI market
- Strong platform engineering foundation and technical curiosity to develop deeper AI infrastructure expertise
- Ability to independently operate business-critical production systems
- Ability to identify and prioritise infrastructure work without detailed instructions
- Ability to lead technical improvements across team boundaries
- Reliable written and remote communication
- Accountability after changes are deployed
- Helpful but not required: experience with Hetzner, NixOS, self-hosted Supabase, ISO 27001 controls and audit evidence, SRE practices, GPUs, or LLM serving technologies such as vLLM
Benefits
Comp & perks- €75,000–€85,000 gross annual salary, depending on experience and scope
- Opportunity to participate in our VSOP
- Permanent, full-time position
- Flexible working hours
- Fully remote work from anywhere within Germany
- Regular opportunities to meet and work with the team in Leipzig, with travel and accommodation covered
- Access to a WellPass fitness membership
