Verified current Job

Technical Lead - GPU Infrastructure

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Technical Lead - GPU Infrastructure based in Polan

Job Remote Full source details
Jobgether Source published Sep 15, 2026 Verified 3 hours ago
✓ 100% verification score · Source: jobgether (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Technical Lead - GPU Infrastructure based in Polan

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Technical Lead - GPU Infrastructure based in Poland. This is a hands-on technical leadership role responsible for building and delivering a full-stack GPU infrastructure platform. You will own architecture and implementation across bare-metal GPU scheduling, Kubernetes, managed inference, and platform observability. The role combines deep infrastructure engineering with leadership of a distributed team spanning backend, frontend, DevOps, QA, and documentation. You will help operate GPU compute environments for research, model training, and inference workloads at scale. A key focus will be reliable GPU fleet operations, high-performance networking, multi-tenant infrastructure, and production-grade platform services. You will also serve as the primary technical interface for infrastructure partners while translating internal workload requirements into robust platform capabilities. The position is fully remote and offers the opportunity to shape a critical AI infrastructure platform from architecture through production delivery.

Own the end-to-end platform architecture, including architecture proposals, high-level and low-level designs, technical reviews, and maintaining the architecture baseline. Lead and line-manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation, setting engineering standards and overseeing code reviews, design reviews, release gates, one-to-ones, and performance development. Design, build, and operate a managed Slurm service for research users, covering controllers, accounting, partitions, login nodes, node onboarding, NVIDIA drivers and CUDA baselines, job and node-health monitoring, autohealing, storage visibility, identity, and workload isolation. Own Kubernetes cluster bootstrap and lifecycle on bare-metal infrastructure, including NVIDIA GPU and Network Operators, KubeVirt and VFIO-based GPU isolation, upgrades, backup and recovery, and node replacement. Define and deliver managed inference architecture covering serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable infrastructure. Establish observability and operational practices across the control plane, GPU fleet, and application tiers, including metrics, logging, alerting, SLOs, incident response, post-incident reviews, and a sustainable on-call model. Act as the primary technical interface with infrastructure partners and vendors, translating requirements into specifications and acceptance tests, managing escalations, and contributing to capacity planning and hardware sourcing. Work directly with research, model-training, and product teams to understand workloads, translate requirements into platform capabilities, and manage capacity constraints. Hire additional members of the platform team and establish the technical standards and expectations for future engineering hires. Maintain a hands-on contribution to architecture, technical reviews, implementation decisions, and infrastructure delivery rather than operating solely in a management capacity. Requirements: Bring 8+ years of hands-on engineering experience, including at least 3 years leading teams that build and operate infrastructure platforms used by other teams. Hold a Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or demonstrate equivalent practical experience. Have hands-on experience operating Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node-health scripting, and upgrades while jobs remain active; experience with HPC or GPU training clusters is highly desirable. Demonstrate deep experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycle management, Fabric Manager, NVSwitch on SXM systems, DCGM-based health and utilization, MIG, node burn-in, and acceptance processes. Have strong knowledge of high-performance interconnects, including InfiniBand fabrics, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues. Possess advanced Linux systems knowledge covering kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning for compute-intensive workloads. Have production Kubernetes operations experience beyond deployment, including control planes, upgrades, CNI and CSI, operators, custom controllers, and multi-tenancy architecture. Understand HPC storage and large-scale data movement, including shared filesystems such as VAST, Lustre, or NFS, node-local NVMe caching, and distributing large model weights and datasets across multiple nodes. Be experienced with observability and production operations using Prometheus, Grafana, Loki, or equivalent technologies, together with SLO management, incident response, and post-incident reviews. Have working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions, without requiring feature-development specialization. Have shipped a platform with real users, such as a multi-tenant IaaS/PaaS, research computing service, or comparable infrastructure platform involving resource isolation, quotas, usage metering, and user-facing APIs or CLI surfaces. Demonstrate leadership that remains technically engaged, including people management across time zones, cross-track technical reviews, documented architecture decisions, and the confidence to challenge partners or executives with clear technical reasoning. Have excellent written and spoken English, particularly for technical, partner, and leadership communication. Be based within the UTC to UTC+5:30 time-zone range to provide working-hour overlap with teams and partners in Europe and India, with availability for occasional travel to partner sites and team events. Experience with Slurm operators on Kubernetes, Kubernetes-native schedulers, modern AI serving stacks such as vLLM, SGLang, or TensorRT-LLM, and GPU parallelism or quantization strategies is a plus. Experience with KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, or GPU autohealing technologies is desirable. Background working on GPU cloud platforms, university or national HPC centers, AI lab infrastructure teams, peer-to-peer or distributed systems, or hardware-provider relationships is also considered an advantage. Benefits: Fully remote work within the specified UTC to UTC+5:30 time-zone range. The opportunity to lead architecture and delivery of a full-stack GPU infrastructure platform spanning bare-metal compute, Kubernetes, Slurm, and managed AI inference. A hands-on technical leadership position combining engineering, architecture, people management, and infrastructure-partner engagement. Collaboration with a distributed international team across Europe and India. Exposure to advanced GPU infrastructure, high-performance networking, AI workloads, multi-tenant compute, and production inference systems. Opportunities to shape engineering standards, platform architecture, hiring, and long-term infrastructure capabilities. Occasional travel opportunities to partner sites and team events.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

jobgether (lever) ↗

Browse current Job and Scholarship listings from jobgether (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books