Verified current Job

Senior Site Reliability Engineer

Team: Infra Reliability · SF Bay Area / Remote (US)

Job Remote Full source details
Luma AI Source published Sep 20, 2026 Verified 12 hours ago
✓ 100% verification score · Source: Luma AI (ashby) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

Team: Infra Reliability · SF Bay Area / Remote (US)

Full job description

Team: Infra Reliability · SF Bay Area / Remote (US) You'll own the GPU infrastructure Luma's research and product run on — thousands of NVIDIA and AMD GPUs across on-prem and multi-cloud (AWS and OCI). As a Senior SRE, you keep training and inference clusters reliable and fast, and you help redesign them for the next level of scale. This is a hands-on, close-to-the-metal role for a first-principles Linux engineer. You'll be the final escalation for the hardest GPU, networking, and kernel-level failures, sometimes debugging directly with NVIDIA. It fits someone who thrives on low-level problems in a fast, less-structured environment. If you want a narrow, well-bounded ops role, this isn't it. What You'll Own

  • Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
  • Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
  • Tune Linux performance deeply, at the OS and kernel level.
  • Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
  • Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA.
  • Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices. First 90 Days One way the first 90 could unfold.
  • Days 1–30 — Immerse & Diagnose: Learn the current clusters across on-prem, AWS, and OCI, and where reliability and performance hurt most.
  • Days 30–60 — Ship & Validate: Take ownership of a production cluster and ship automation or tuning that measurably improves availability or performance.
  • Days 60–90 — Scale & Systemize: Contribute to the next-gen re-architecture and harden security and compliance practices. What You Bring
  • 5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
  • Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
  • Working experience with Terraform, Airflow, and Ray.
  • Strong experience with AWS or OCI.
  • Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
  • Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
  • Comfort in a less-structured, fast-paced environment. Nice to Have
  • Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm).
  • Experience managing large-scale GPU clusters for AI/ML training or inference.
  • Familiarity with Kubernetes or orchestration frameworks like Ray.
  • Deep expertise in data pipelines and infrastructure. About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v2

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Luma AI (ashby) ↗

Browse current Job and Scholarship listings from Luma AI (ashby) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books