Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
ASAPP

Speech Software Engineer

ASAPP

. Tune and optimize ASR and TTS models for real-world call center environments .

Posted 9/29/2026full-timeUnited StatesMid-LevelSenior💰 $215,000 - $235,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in tuning and optimizing ASR and TTS models for real-time applications, with a strong focus on improving transcription accuracy and managing low-latency, high-concurrency systems. Proficient in deploying ML services and collaborating with cross-functional teams to enhance voice infrastructure and ensure compliance with regulatory standards.

Highest-signal resume keywords
ASR And TTS SystemsGolang Or PythonLow-Latency Systems DesignKubernetes And DockerSpeech Quality Evaluation

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Model Performance EvaluationNoise Reduction TechnologiesSpeech Model TuningStreaming Data HandlingAudio FundamentalsConcurrency ManagementEvent-Driven SystemsSpeech Enhancement TechniquesMedia Stream HandlingStructured Experimentation
Tools & Technologies
AWSGCPAzureML ToolingMonitoring Tools
Industry Keywords
Voice InfrastructureTelephony ProvidersRegulatory Voice-Data StandardsWERCERMOSDiarizationEcho CancellationForced Alignment TechniquesReal-Time Audio Streams

Tech Stack

Tools & technologies
AWSAzureDistributed SystemsDockerGoogle Cloud PlatformKubernetesPythonGo

About the role

Key responsibilities & impact
  • Tune and optimize ASR and TTS models for real-world call center environments
  • Improve transcription accuracy, noise robustness, speaker variability, speech naturalness, prosody, pacing, pronunciation, and conversational flow
  • Balance latency and quality tradeoffs in streaming speech pipelines
  • Evaluate and integrate noise suppression, voice activity detection, and diarization technologies
  • Architect and modernize scalable, highly available voice infrastructure
  • Build multithreaded, low-latency server frameworks handling thousands of concurrent real-time audio streams
  • Design and operate streaming ASR → LLM → TTS pipelines for live AI-driven customer conversations
  • Develop robust media stream handling between telephony providers, clients, and ML services
  • Define speech quality evaluation frameworks using WER/CER analysis, latency tracking, and perceptual TTS metrics
  • Build monitoring tools and dashboards to detect production regressions
  • Create load-testing and simulation tools for high-concurrency voice traffic
  • Partner with Speech Scientists and ML Researchers to productionize ASR and TTS models
  • Work with Security and Compliance teams on enterprise and regulatory voice-data standards
  • Collaborate with Product teams to translate conversational quality requirements into measurable improvements

Requirements

What you’ll need
  • 5+ years of software engineering experience building and operating production-grade distributed systems
  • Strong proficiency in Golang or Python, or willingness to become an expert quickly
  • Experience designing low-latency, high-concurrency systems, ideally involving real-time media or streaming data
  • Practical experience working with ASR and/or TTS systems in applied or production environments
  • Understanding of adapting and tuning speech models for domain-specific use cases
  • Familiarity with WER, CER, MOS, latency, and streaming stability metrics
  • Strong understanding of audio fundamentals, including sample rates, Opus and G.711 codecs, buffering, packet loss, and jitter
  • Experience evaluating model performance and running structured experiments to improve transcription accuracy and speech naturalness
  • Comfort with modern ML tooling and model APIs for fine-tuning, adapting, or post-processing speech model outputs
  • Ability to balance model quality, compute cost, and real-time constraints
  • Experience with noise reduction, echo cancellation, VAD, diarization, or other speech enhancement technologies
  • Familiarity with forced alignment techniques or phoneme/word-level timing models
  • Hands-on experience deploying ML services with Kubernetes, Docker, and AWS, GCP, or Azure
  • Knowledge of event-driven and asynchronous systems, including async I/O, event loops, and streaming frameworks
  • Experience analyzing large-scale speech or conversation datasets

Benefits

Comp & perks
  • Performance bonus
  • Equal opportunity employment
  • Disability assistance during the application process