Skip to main content

Modal

Forward Deployed Engineer - ML

New York$180k–$250kfulltimemidAdded 1 month ago

About this role

Modal is seeking a Forward Deployed Engineer specializing in machine learning to collaborate with leading AI companies on optimizing production workloads. This role emphasizes hands-on engagement with customers to enhance their AI capabilities through technical support and innovative solutions.

What you'll do

  • Architect and optimize AI workloads with major clients
  • Contribute to open-source projects and publish technical content
  • Collaborate with product and sales teams on engineering solutions
  • Build relationships with technical leaders in AI sectors
  • Conduct technical demos and proof-of-concepts

What they're looking for

  • 2+ years of ML engineering experience
  • Experience in inference optimization and model training
  • Knowledge of GPU programming and ML infrastructure
  • Ability to communicate technical concepts to leadership
  • Interest in direct customer engagement
  • Familiarity with serving and training toolchains
  • Experience with open-source contributions
  • Willingness to work in-person in specified locations
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Modal

Modal builds a cloud platform for running large-scale AI workloads and infrastructure, enabling companies to deploy and optimize production machine learning systems. The company is hiring Forward Deployed Engineers to work directly with AI customers, Infrastructure Security Engineers to strengthen platform security, and Developer Relations Engineers to engage the developer community with technical content and best practices.

Website
modal.com
View all jobs at Modal

Likely interview questions

  • Walk us through a time you optimized an ML inference or training workload in production. What were the bottlenecks, and how did you measure improvement?
  • You're working with a customer running LLM inference who's hitting latency issues. How would you diagnose the problem, and what tools or frameworks would you reach for first?