Verified current Job

Staff Machine Learning Systems Engineer

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Machine Learning Systems Engineer based in t

Job Remote Full source details
Jobgether Source published Sep 21, 2026 Verified 1 hour ago
✓ 100% verification score · Source: jobgether (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Machine Learning Systems Engineer based in t

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Machine Learning Systems Engineer based in the United States. This senior technical role focuses on building and scaling the infrastructure that powers large-scale machine learning systems. You will design end-to-end MLOps patterns that accelerate model development, experimentation, deployment, and iteration. The position combines platform engineering, distributed ML, data processing, performance optimization, and cloud infrastructure. You will help develop scalable graph machine learning platforms capable of handling billions of nodes and tens of billions of edges. Working closely with machine learning engineers, you will improve training performance, resource efficiency, reliability, and GPU utilization. The role offers significant technical ownership in a distributed, cloud-based environment where scalability and ease of use are critical. You will also serve as a strong advocate for platform users while shaping engineering practices and infrastructure that enable teams to move faster.

Design and implement end-to-end model lifecycle patterns and MLOps capabilities covering data preparation, model management, experiment tracking, deployment, and related workflows. Lead zero-to-one development of graph machine learning infrastructure and codebases that abstract common patterns and enable scalable model development. Collaborate with machine learning engineers to optimize model performance, training duration, memory utilization, and GPU costs across large distributed environments. Optimize batch data processing and data pipelines using technologies such as Apache Beam, Apache Spark, Ray Data, and cloud data warehouse services. Architect pipelines capable of building and maintaining massive graph data structures containing billions of nodes and tens of billions of edges. Build and evolve cloud-based infrastructure supporting machine learning platforms, with an emphasis on scalability, reliability, performance, and operational efficiency. Administer and integrate MLOps tools for experiment tracking, model serving, and model registries. Develop solutions that improve the overall machine learning development lifecycle and make platform capabilities easier and more efficient for engineering teams to use. Partner with cross-functional technical teams to understand infrastructure needs, identify bottlenecks, and deliver scalable solutions. Establish and promote engineering patterns that improve platform reliability, developer productivity, model iteration, and ease of use. Provide technical leadership across complex infrastructure initiatives and help shape the long-term architecture of machine learning systems. Requirements: 8+ years of experience working in machine learning infrastructure, including model training and model deployment environments. Hands-on experience optimizing machine learning systems, including memory profiling, GPU profiling, training performance, and resource utilization. Deep experience with cloud technologies used to support ML platforms, including Google Cloud Platform services such as BigQuery and Google Cloud Storage. Experience with infrastructure-as-code tools such as Terraform. Hands-on experience administering and integrating MLOps platforms for experiment tracking, model serving, and model registries, such as MLflow or Weights & Biases. Strong proficiency in Python and familiarity with machine learning frameworks such as PyTorch and TensorFlow. Deep experience with distributed machine learning and data processing frameworks, including Ray and Kubernetes. Strong understanding of scalable, reliable, high-performance platform architecture and the machine learning development lifecycle. Ability to advocate effectively for platform users and translate their needs into intuitive, scalable infrastructure solutions. Strong organizational, communication, and collaboration skills, with the ability to work effectively across technical teams. Experience with graph databases such as Neo4j, JanusGraph, or TigerGraph is highly valued. Experience with graph neural networks and graph ML frameworks such as PyTorch Geometric or Deep Graph Library is highly valued. Ability to operate effectively in a technically complex, rapidly evolving environment and take ownership of large-scale infrastructure initiatives. Benefits: Base salary range of $230,000–$322,000 USD , with final compensation determined by factors including skills, experience, and relevant credentials. Eligibility for equity in the form of restricted stock units. Potential eligibility for commission depending on the position offered. Medical, dental, and vision insurance for U.S.-based employees. 401(k) program with employer matching. Generous paid vacation and time-off benefits. Paid parental leave. Remote work opportunity within the United States. Opportunities to work on large-scale machine learning infrastructure and technically challenging systems. Opportunities for significant technical ownership and professional growth. For select roles and locations, interviews may be recorded, transcribed, and summarized using AI; candidates can opt out of these processes before scheduled interviews.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

jobgether (lever) ↗

Browse current Job and Scholarship listings from jobgether (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books