Verified current Job

ML Infrastructure Engineer

ABOUT THE ROLE

Job Full source details
Clera Source published Sep 14, 2026 Verified 6 days ago
✓ 100% verification score · Source: Clera (ashby) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.

Overview

ABOUT THE ROLE

Full job description

ABOUT THE ROLE This is a hands-on infrastructure engineering role at an early-stage enterprise AI company building a context and data governance layer for AI agents in highly regulated industries. You will own the inference and model-serving infrastructure end to end, ensuring AI agents run reliably, accurately, and at scale in production environments where performance is non-negotiable. WHAT YOU'LL DO

  • Design, build, and operate inference and model-serving infrastructure from development through production deployment.
  • Scale systems to support AI agents running reliably under increasing concurrency and production load.
  • Identify and resolve infrastructure bottlenecks in close collaboration with ML and platform engineering teams.
  • Optimize systems for latency, throughput, and reliability at scale. WHAT WE'RE LOOKING FOR
  • 5 or more years building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.
  • Hands-on experience designing and scaling inference serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.
  • Strong systems engineering fundamentals with expertise in distributed systems, containerization, and orchestration (Docker, Kubernetes).
  • Demonstrated ability to optimize production ML systems for latency, throughput, and reliability under high concurrency.
  • Experience with cloud infrastructure platforms such as AWS, GCP, or Azure for deploying and managing ML workloads.
  • Proficiency with monitoring, observability, and debugging tools such as Prometheus, Grafana, ELK, or distributed tracing frameworks.
  • Proficiency in at least one systems programming or backend language: Python, Go, Rust, C++, or Java.
  • Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune) is a plus.
  • Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines is a plus.
  • Experience with enterprise data infrastructure, data pipelines, or data integration platforms is a plus. LOCATION This role is on-site in San Mateo, California. Visa sponsorship is not available.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v2

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Clera (ashby) ↗

Browse current Job and Scholarship listings from Clera (ashby) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books