Skip to main content

Point72

Machine Learning Infrastructure Engineer, GenAI Technology

New York, NY$180k–$300kmidAdded 1 month ago

About this role

Point72 is seeking a Machine Learning Infrastructure Engineer to enhance its AI capabilities by developing high-performance infrastructure for generative AI and machine learning workloads. The role involves collaborating with teams to optimize ML model training and deployment, while ensuring reliable and efficient workflows.

What you'll do

  • Design high-performance infrastructure for generative AI and ML workloads
  • Develop distributed systems for model training and data preprocessing
  • Collaborate with ML researchers to optimize model training
  • Automate deployment and CI/CD pipelines using infrastructure-as-code
  • Implement monitoring and cost-management strategies for compute environments
  • Evaluate emerging hardware and software technologies for scalability

What they're looking for

  • Bachelor’s or Master’s in Computer Science or related field
  • 3–7 years of scalable compute or ML infrastructure experience
  • Knowledge of distributed systems and container orchestration
  • Experience with ML operations tools like MLflow, Ray, and Airflow
  • Proficient in Python and systems-level programming (Go, C++, Rust)
  • Strong debugging and performance optimization skills
  • Excellent collaboration and communication skills
  • Understanding of reinforcement learning concepts

Benefits

  • Fully-paid health care benefits
  • Generous parental and family leave policies
  • Support for employee-led affinity groups
  • Mental and physical wellness programs
  • Tuition assistance
  • 401(k) savings program with employer match
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Point72

Point72 operates trading, data infrastructure, and investment management technology platforms that power systematic portfolio management and capital markets operations. The company is hiring infrastructure engineers, data reliability specialists, and software engineers to build and maintain mission-critical systems including data pipelines, network architectures, storage platforms, and research tools.

View all jobs at Point72

Likely interview questions

  • Walk us through a time when you designed infrastructure for a large-scale ML workload. What were the key performance bottlenecks you encountered, and how did you optimize for compute utilization and training throughput?
  • Describe your experience with Kubernetes and container orchestration in production ML environments. How have you handled resource allocation and scheduling for GPU-intensive workloads?