Together AI
Forward Deployed Engineer (Inference & Post-Training)
About this role
Together AI seeks a Forward Deployed Engineer to serve as a technical expert for strategic customers deploying large language models at scale. You'll optimize inference engines, guide fine-tuning pipelines, and directly influence product development while partnering with customers on complex production AI challenges.
What you'll do
- Select and optimize inference engines (vLLM, TensorRT-LLM, SGLang) based on hardware, model architecture, and workload requirements
- Tune inference performance through KV cache optimization, speculative decoding, tensor parallelism, and quantization strategies to meet latency and throughput targets
- Guide customers through post-training workflows including LoRA, SFT, DPO, RLHF, and GRPO from experimentation to production
- Serve as primary technical contact for strategic accounts, monitoring configurations and ensuring platform adoption success
- Establish optimized inference and post-training configurations during customer onboarding to accelerate time-to-value
- Influence product roadmap by surfacing field insights and driving early adoption of new features with key customers
What they're looking for
- Inference engine optimization (vLLM, TensorRT-LLM, SGLang)
- LLM deployment and open-source model expertise
- Performance tuning (KV cache, speculative decoding, tensor/pipeline parallelism, quantization)
- Post-training pipelines (LoRA, SFT, DPO, RLHF, GRPO)
- Python programming in production environments
- Model selection and system design evaluation
- Customer-facing technical partnership
- AI infrastructure and hardware profiling
Benefits
- Competitive base salary
- Startup equity
- Health insurance
- Remote work flexibility
- Influence on AI infrastructure product direction
- Work with cutting-edge open-source AI research
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Together AI
Together AI builds GPU compute infrastructure and open-source model customization platforms for AI developers and enterprises. The company is hiring for infrastructure operations, ML systems engineering, go-to-market technology, customer success, and GPU research roles.
- Website
- together.ai
Likely interview questions
- Describe your hands-on experience optimizing inference engines like vLLM or TensorRT-LLM—what performance issues have you diagnosed and resolved?
- Walk us through a time you tuned KV cache, speculative decoding, or quantization strategies to hit specific throughput and latency targets.