Verified current Job

Senior Infrastructure Engineer - GPU Compute

Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heteroge...

Job Remote Full source details
Boundless Networks, Inc. United States Source published Aug 3, 2026 Verified 2 weeks ago
✓ 92% verification score · Source: Boundless Networks Inc 1 Careers · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull Time
Work modeRemote / location-flexible
CountryUnited States
DepartmentBoundless Team
Job functionEngineering
IndustryInformation Technology and Services

Overview

Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization. You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action. What You'll Do GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) u

Full job description

Full Job Description

Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization.

You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.

What You'll Do

GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably.

Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the "always-on inference substrate" saturated and economical.

Bare-Metal & GPU Optimization: Go deep on GPU performance — PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology — to push throughput per node.

Reliability, Access & Observability: Build secure fleet access (Tailscale, Teleport), robust observability and alerting, and zero-downtime rollouts across a distributed node fleet.

Cost Optimization: Drive down $/GPU-hr through spot instance management, intelligent workload placement between on-prem and cloud, and resource scheduling — without sacrificing reliability.

Requirements

  • 5+ years of infrastructure/DevOps experience operating large-scale production systems
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale
  • Strong Linux systems administration skills
  • Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
  • Track record of managing mission-critical, high-throughput systems
  • Strong infrastructure-as-code background in heterogeneous environments
  • Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
  • Comfort navigating ambiguity with a strong bias for action

Nice to Have

  • Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
  • Experience operating ML training or other large-scale distributed compute infrastructure
  • Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
  • Familiarity with fleet access and networking tooling (Tailscale, Teleport)
  • Knowledge of network optimization and topology design
  • Experience with multi-region, globally distributed systems
  • Proficiency in Rust or low-level systems programming
  • Experience with on-premises data center operations

Additional Requirements

  • Candidates must include a public GitHub profile in their application.
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history.
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.

Benefits

At Boundless, we take care of our people, because building the future of AI compute starts with an empowered team. Here's what you can expect when you join us:

  • Competitive salary (proposed band b/t US$175k and $250k annually) + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment

We are a global team, and applicants from around the world are welcome to apply.

Qualifications And Requirements

Mid-Senior level

Requirements & qualifications

Mid-Senior level

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

Discovered directly from the employer’s public Workable account endpoint with details=true where supported. Public description, responsibilities, requirements and benefits were normalized into complete candidate-facing sections.

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Boundless Networks Inc 1 Careers ↗

Browse current Job and Scholarship listings from Boundless Networks Inc 1 Careers →

Related opportunities

Other current verified records you may want to review.

Job

Mitarbeiter (m/w/d) Burger-Crew für Street Food-Court im Bistro von “Moulin Rouge! Das Musical“

Atgde Careers

Current Mitarbeiter (m/w/d) Burger-Crew für Street Food-Court im Bistro von “Moulin Rouge! Das Musical“ opening at Atgde Careers in Hamburg. Full...

Job

Assistenz / Stellv. Teamleitung für Street-Food-Court „Bistro“ von "Moulin Rouge! Das Musical"

Atgde Careers

Current Assistenz / Stellv. Teamleitung für Street-Food-Court „Bistro“ von "Moulin Rouge! Das Musical" opening at Atgde Careers in Hamburg. Full...

Job

Recruiter:in / Personalreferent:in – Teilzeit 20h / 30h(m/w/d)

Assentio Careers

Current Recruiter:in / Personalreferent:in – Teilzeit 20h / 30h(m/w/d) opening at Assentio Careers in Leipzig. Full employer-published role secti...

Job

Personalberater / 360° Recruiter (m/w/d)

Assentio Careers

Current Personalberater / 360° Recruiter (m/w/d) opening at Assentio Careers in Leipzig. Full employer-published role sections have been imported...

Job

Infrastructure Technical Program Manager, Global Connectivity Infrastructure Delivery - Fiber Deployment

Amazon Data Services, Inc. · United States

AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people...

Job

Infra-Delivery Install Technician, Infrastructure Delivery

Amazon Data Services, Inc. · United States

AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people...

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books