ASAPP
Speech Software Engineer
New York$215k–$235kfull-timemidAdded today
About this role
ASAPP seeks a Senior Speech Software Engineer to optimize and scale real-time voice AI infrastructure for enterprise call centers. You'll bridge cutting-edge speech research with production systems, tuning ASR/TTS models for low-latency, high-accuracy performance while architecting resilient, high-concurrency voice pipelines.
What you'll do
- Tune and optimize ASR and TTS models for accuracy, noise robustness, and real-world call center environments
- Architect scalable, high-availability voice infrastructure handling thousands of concurrent real-time audio streams
- Design and operate end-to-end streaming pipelines integrating ASR, LLM, and TTS for live conversations
- Define speech quality evaluation frameworks and build monitoring dashboards for production performance
- Evaluate and integrate emerging speech technologies like noise suppression, VAD, and diarization
- Collaborate with Speech Scientists, ML Researchers, and Product teams to productionize models and translate requirements
What they're looking for
- Golang or Python
- ASR and TTS systems (applied or production experience)
- Low-latency, high-concurrency distributed systems design
- Speech quality metrics (WER, CER, MOS, latency)
- Audio fundamentals (codecs, buffering, packet loss, jitter)
- ML model evaluation and fine-tuning
- Real-time media or streaming data handling
- Kubernetes, Docker, and cloud deployment
Benefits
- Work on cutting-edge AI-powered voice technology at scale
- Collaborate with world-class Speech Scientists and ML Researchers
- Fast-paced startup environment with continuous learning opportunities
- Global team with hubs in NYC, Mountain View, Latin America, and India
- Performance-based bonus compensation
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
ASAPP
- Website
- asapp.com
Likely interview questions
- Describe a production ASR or TTS system you've optimized—what quality/latency tradeoffs did you navigate?
- How would you design a low-latency streaming pipeline to handle thousands of concurrent call center audio streams?