Decagon
Research Engineer, Audio and Speech
About this role
Decagon is seeking a Research Engineer to develop production-grade audio and speech models powering real-time voice agents for enterprise customers. You'll design streaming agent systems, train multimodal models, and optimize end-to-end inference from research prototype to scalable deployment.
What you'll do
- Design agent harnesses optimized for streaming speech, turn-taking, interruptions, and continuous real-time interaction
- Research and train multimodal full-duplex models that jointly understand audio, reason, and generate speech
- Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers and environments
- Build evaluations using production call data to ship measurable improvements in accuracy, latency, and naturalness
- Optimize end-to-end inference for responsiveness, throughput, and cost in collaboration with platform teams
- Take research ideas from prototype through reliable production deployment
What they're looking for
- Speech and audio machine learning
- PyTorch or modern deep learning frameworks
- Streaming agent systems and low-latency inference
- Autoregressive, diffusion, flow-matching, or codec-based model development
- Python and signal processing
- Production model serving and evaluation
- Multimodal machine learning
- Real-world audio data handling
Benefits
- Equity compensation
- Work on cutting-edge conversational AI technology
- End-to-end ownership of impactful projects
- Collaboration with world-class research team
- In-office environment in San Francisco
- Exposure to enterprise-scale production systems
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Decagon
Decagon builds enterprise-grade conversational AI platforms that enable organizations to deploy AI agents for business impact. The company is hiring Strategic Solutions Engineers, Customer Engineers, Platform Engineers, and systems-focused engineers to deliver AI implementations, build internal infrastructure, and establish security practices across their growing platform.
View all jobs at DecagonLikely interview questions
- Describe your experience developing or adapting speech generation models—what architectures have you worked with and how did you optimize them for production?
- How have you approached latency optimization in streaming inference systems, and what trade-offs did you navigate?