Verified current Job

Senior Machine Learning Systems Engineer

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Machine Learning Systems Engineer based in

Job Remote Full source details
Jobgether Source published Sep 21, 2026 Verified 1 hour ago
✓ 100% verification score · Source: jobgether (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Machine Learning Systems Engineer based in

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Machine Learning Systems Engineer based in United States. As a Senior Machine Learning Systems Engineer, you will help build and evolve the infrastructure powering large-scale machine learning systems. You’ll design end-to-end MLOps patterns that improve how models are developed, trained, deployed, and monitored. The role combines distributed systems, cloud infrastructure, data processing, and machine learning engineering at significant scale. You’ll contribute to graph machine learning platforms capable of supporting billions of nodes and tens of billions of edges. You’ll work closely with ML engineers to optimize training performance, GPU utilization, data pipelines, and infrastructure costs. This is a collaborative, technically demanding environment where scalability, reliability, and developer experience are core priorities. You’ll have the opportunity to influence platform architecture while solving complex infrastructure challenges across machine learning workflows.

Design and implement end-to-end model lifecycle and MLOps patterns covering data preparation, model management, experiment tracking, deployment, and related workflows. Develop and support graph machine learning platforms and codebases that abstract common patterns and enable scalable model development and iteration. Partner with machine learning engineers to improve training performance, reduce model training times, increase efficiency, and optimize GPU costs across distributed environments. Optimize large-scale batch data processing using cloud data warehouses and distributed processing technologies such as Apache Beam, Apache Spark, and Ray Data. Architect pipelines capable of creating and maintaining massive graph datasets containing billions of nodes and tens of billions of edges. Administer and integrate tooling for experiment tracking, model serving, model registries, and other components of the ML platform. Build infrastructure and platform capabilities that prioritize scalability, reliability, performance, maintainability, and ease of use. Collaborate closely with platform users to understand their development lifecycle and remove technical friction from machine learning workflows. Contribute to architectural and technical decisions across cloud infrastructure, distributed training, data processing, and machine learning systems. Help improve platform efficiency and cost management while maintaining reliable infrastructure for large-scale ML workloads. Requirements 5+ years of professional experience working with machine learning infrastructure, including model training and deployment. Hands-on experience optimizing machine learning workloads, including memory utilization, GPU profiling, training performance, and resource efficiency. Deep experience with cloud technologies supporting ML platforms, particularly services such as GCP BigQuery and Google Cloud Storage, as well as infrastructure-as-code tools such as Terraform. Practical experience administering and integrating MLOps technologies for experiment tracking, model serving, and model registries, such as MLflow or Weights & Biases. Strong programming skills in languages and frameworks commonly used for machine learning, including Python, PyTorch, and/or TensorFlow. Deep experience with distributed training and compute frameworks, including Ray and Kubernetes. Strong understanding of scalable data processing and distributed systems, with the ability to work effectively across complex ML infrastructure environments. A strong focus on scalability, reliability, performance, and developer experience, with an ability to understand and advocate for the needs of platform users. Strong organizational, communication, analytical, and problem-solving skills, with the ability to collaborate effectively across technical teams. Experience with graph databases such as Neo4j, JanusGraph, or TigerGraph is a plus. Experience with graph neural networks and graph ML frameworks such as PyTorch Geometric or Deep Graph Library is a plus. Benefits Base salary range of $216,700–$303,400 USD , with final compensation determined by factors including skills, experience, credentials, and other job-related considerations. Eligibility for equity in the form of restricted stock units. Potential eligibility for commission depending on the position offered. Medical, dental, and vision insurance options for U.S.-based employees. 401(k) program with employer matching. Generous paid time off and vacation benefits. Parental leave to support employees and their families. Remote work from within the United States. Opportunities to work on large-scale machine learning infrastructure, distributed systems, graph ML, and cloud technologies. A technically collaborative environment focused on innovation, scalability, reliability, and continuous improvement.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

jobgether (lever) ↗

Browse current Job and Scholarship listings from jobgether (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books