FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and operating production distributed systems, with a strong focus on backend programming, cloud infrastructure, and observability. Proven ability to improve system reliability, scalability, and operational efficiency while mentoring teams in best practices for backend architecture.
Highest-signal resume keywords
Backend Programming Language ProficiencyDistributed Systems DesignCloud Infrastructure ManagementAPI Design and Data ModelingIncident Response and Post-Incident Analysis
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Asynchronous ProcessingMessage QueuesEvent-Driven SystemsIdempotencyConcurrencyDebugging SkillsSLOs and Error BudgetsWorkflow ExecutionObservability ToolsCapacity Planning
Soft Skills
MentoringCollaborationProblem-Solving
Tools & Technologies
KafkaPub/SubSQSGrafanaOpenTelemetrySentryKubernetesInfrastructure as CodeCloud-Native Deployment SystemsTemporal
Industry Keywords
HealthcareHIPAASOC 2SecurityPrivacyRegulated Environment
Tech Stack
Tools & technologiesCloudDistributed SystemsGrafanaKafkaKubernetes
About the role
Key responsibilities & impact- Design, build, and operate backend systems powering Ambient AI products across mobile and web
- Build inference workflows coordinating transcription, diarization, language models, clinical extraction, summarization, and other AI capabilities
- Develop orchestration systems for long-running, multi-stage workflows
- Implement scheduling, queueing, workload prioritization, retries, timeouts, fallbacks, dead-letter handling, idempotency, replay, recovery, model routing, versioning, progress tracking, and graceful degradation
- Improve reliability and scalability under variable, compute-intensive workloads
- Define and maintain SLOs for availability, processing latency, completion rates, and data durability
- Build observability through structured logging, metrics, distributed tracing, dashboards, alerting, and diagnostic tooling
- Lead incident response and post-incident analysis
- Identify systemic failure patterns and address them through abstractions, automation, testing, and architecture
- Improve cloud infrastructure, deployment systems, capacity planning, and operational tooling
- Design APIs and data models for safe, predictable interaction with long-running workflows
- Build tools for inspecting, testing, replaying, evaluating, and debugging inference pipelines
- Balance reliability, latency, quality, and infrastructure cost
- Partner with mobile, product, AI, security, and clinical teams
- Mentor engineers and establish practices for distributed systems, observability, operational readiness, and backend architecture
- Comply with information security policies and report confirmed or potential security events and risks
Requirements
What you’ll need- 8+ years of professional backend or infrastructure engineering experience
- Experience designing, building, and operating production distributed systems
- Strong proficiency in at least one modern backend programming language
- Experience with asynchronous processing, message queues, event-driven systems, or durable workflow execution
- Strong understanding of idempotency, consistency, concurrency, backpressure, retries, failure isolation, and eventual completion
- Experience operating services in a cloud environment, including deployment, monitoring, scaling, and incident response
- Experience designing APIs, service boundaries, and data models for complex product workflows
- Track record of improving system reliability, observability, scalability, or operational efficiency
- Strong debugging skills across application, infrastructure, data, and external dependency boundaries
- Ability to reason about real-time and long-running workloads with different latency and durability requirements
- Comfort working across backend, infrastructure, product, and AI systems
- Nice to have: speech-to-text, diarization, transcription, or audio-based machine-learning workflows
- Nice to have: large language models or multi-model inference pipelines in production
- Nice to have: Temporal or similar durable execution platforms
- Nice to have: Kafka, Pub/Sub, SQS, or similar messaging systems
- Nice to have: model routing, inference gateways, rate limiting, batching, caching, or GPU-backed workloads
- Nice to have: workflow inspection, replay, evaluation, or model debugging tooling
- Nice to have: SLOs, error budgets, alerting, and production incident response
- Nice to have: Grafana, OpenTelemetry, Sentry, or similar observability tools
- Nice to have: containers, Kubernetes, infrastructure as code, and cloud-native deployment systems
- Nice to have: throughput, tail latency, infrastructure cost, and workload isolation optimization
- Nice to have: offline clients, background synchronization, or delayed and duplicated event reconciliation
- Nice to have: healthcare, HIPAA, SOC 2, encryption, privacy, security, or regulated-environment experience
- Nice to have: AI-native, real-time, or agentic product experience
- Position requires being in the San Francisco, CA or Mountain View, CA office at least 3 days per week
- Must answer whether sponsorship to work in the US will be required
Benefits
Comp & perks- Equity
- Opportunity to build an AI healthcare product used in real clinical workflows
- Grow into broader backend, infrastructure, or technical leadership
- Official company communications exclusively from @commure.com email addresses
