Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
LILT

Machine Learning Engineer – Real-Time Speech Translation

LILT

. Build and manage high-throughput, real-time audio and text streaming services .

Posted 10/2/2026full-timeUnited StatesMid-LevelSenior💰 $129,147 - $161,434 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and managing high-throughput, real-time audio and text streaming services, with a strong focus on integrating machine learning models and optimizing system performance. Proficient in async programming and experienced in architecting GPU-accelerated infrastructure using Docker and Kubernetes.

Highest-signal resume keywords
Python ProgrammingAsync Programming (Asyncio)Machine Learning Model ServingKubernetes InfrastructureWebSocket or gRPC Streaming

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Signal ProcessingConcurrency ManagementReal-Time CorrectionsPerformance InstrumentationOptimizing ThroughputAudio Ingestion APIsBatching StrategiesLoad BalancingAutoscalingProfiling Low-Latency Systems
Tools & Technologies
DockerKubernetesRay ServeRabbitMQDatadogPrometheusClaude CodeCodexWebRTCLiveKit Agents
Industry Keywords
Streaming ASRMachine TranslationNLP ModelsCJK Text HandlingIncremental MTMachine Translation Quality Estimation

Tech Stack

Tools & technologies
DockerGRPCKubernetesPrometheusPythonRabbitMQRay

About the role

Key responsibilities & impact
  • Build and manage high-throughput, real-time audio and text streaming services
  • Handle signal processing, session lifecycles, and concurrency management
  • Integrate and serve streaming speech recognition and machine translation models
  • Develop confidence scoring, human-intervention routing, and real-time corrections
  • Architect and scale GPU-accelerated Kubernetes infrastructure
  • Implement batching, load balancing, and autoscaling strategies
  • Instrument performance, identify bottlenecks, optimize throughput, and reduce latency
  • Define audio ingestion and downstream integration APIs and technical contracts
  • Collaborate with research, frontend, platform engineering, and product teams

Requirements

What you’ll need
  • BS or MS in Computer Science or a related field, or equivalent practical experience
  • 3+ years building production backend or ML serving systems in Python
  • Strong async programming (asyncio) skills
  • Experience with WebSocket or gRPC bidirectional streaming, session state, backpressure, and connection lifecycle handling
  • Experience serving ML models in production on GPUs, with Docker and Kubernetes
  • Experience integrating speech or NLP models into production systems, ideally streaming ASR
  • Experience profiling, instrumenting, and optimizing a real-time or low-latency system
  • Effective use of AI coding agents such as Claude Code or Codex
  • US citizenship and residence in the United States
  • Preferred: Ray Serve, simultaneous or incremental MT, machine translation quality estimation, RabbitMQ, streaming text-to-speech, WebRTC, SFU concepts, LiveKit Agents, Pipecat, CJK text handling, Datadog, and Prometheus

Benefits

Comp & perks
  • Growth opportunities
  • Leading tools
  • Global collaboration
  • AI-first development environment with agentic coding and AI-driven PR reviews