Skip to main content

Snowflake

Senior/Staff System Research Engineer – LLM Inference Optimization

US-WA-Bellevue (Remote)$236k–$330kfulltimeseniorAdded 2 weeks ago

About this role

Snowflake's AI Research team seeks systems engineers and researchers to advance LLM inference optimization across the full stack—from distributed serving and runtime systems to GPU kernels. You'll develop high-performance inference systems using techniques like speculative decoding, adaptive parallelism, and KV-cache optimization, while building intelligent systems that automatically adapt to new models, hardware, and workloads.

What you'll do

  • Design and develop high-performance LLM inference systems spanning distributed serving, runtime systems, and GPU execution
  • Develop novel techniques to optimize inference latency, throughput, memory efficiency, scalability, and cost
  • Explore advanced inference methods including speculative/parallel decoding, prefill/decode disaggregation, and adaptive parallelism
  • Build adaptive inference systems that automatically optimize for new architectures, hardware, and workload characteristics
  • Apply AI-driven approaches to systems engineering including automated profiling, configuration search, and performance tuning
  • Analyze and optimize GPU kernels for attention, MoE, communication, and other performance-critical components

What they're looking for

  • LLM inference systems and optimization
  • CUDA/GPU kernel development and optimization
  • Distributed systems and parallel computing
  • Performance profiling and benchmarking
  • Python and systems programming languages (C++)
  • Model optimization techniques (quantization, pruning, speculative decoding)
  • Machine learning systems and ML frameworks
  • Networking and communication optimization
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Snowflake

Snowflake builds a cloud data platform with marketplace capabilities, analytics infrastructure, and AI-powered data solutions, supported by robust security and streaming systems. The company is hiring full-stack engineers, analytics engineers, solution engineers, security-focused software engineers, and principal engineers to enhance its platform and infrastructure.

View all jobs at Snowflake

Likely interview questions

  • Describe your experience optimizing inference performance—what bottlenecks have you identified and how did you resolve them?
  • How would you approach designing a distributed inference system to serve multiple LLMs with different latency and throughput requirements?