Overview
ABOUT THE ROLE
Full job description
ABOUT THE ROLE This is an early-team Platform Engineer role at a seed-stage AI infrastructure startup building an open-source, GitOps-native distributed operating system on top of Kubernetes. You will work closely with the founding team to design and evolve core platform features, helping take companies from bare metal to production-grade AI clusters in days. Your work directly shapes the architecture that makes multi-cluster Kubernetes management intuitive for developers running demanding GPU workloads. WHAT YOU'LL DO
- Build and extend core platform features in Go, including custom Kubernetes operators and controllers.
- Design and implement GitOps workflows using ArgoCD to make continuous deployment feel seamless.
- Develop infrastructure-as-code patterns with Terraform and Helm to provision and manage clusters.
- Work on distributed storage solutions using Ceph and WEKA for high-performance, scalable cluster storage.
- Create observability and monitoring systems with Prometheus and Grafana to surface cluster health and performance.
- Build and optimize container networking with Cilium for network security and deep observability.
- Design and implement federated Kubernetes architectures for multi-cluster management.
- Build automation tooling that reduces operational overhead for developers running production workloads.
- Contribute to open-source components and help establish platform architecture patterns as an early team member. WHAT WE'RE LOOKING FOR
- 2+ years of software development experience, with approximately 5 years preferred.
- Hands-on proficiency in Go and Kubernetes, including writing Go in a Kubernetes environment (operators, controllers, CRDs).
- Experience managing Kubernetes clusters at meaningful production scale.
- Experience designing federated Kubernetes architectures for multi-cluster management.
- Proficiency with GitOps workflows and ArgoCD.
- Experience with infrastructure as code using Terraform and Helm or comparable tools.
- Experience with distributed storage solutions such as Ceph or WEKA.
- Familiarity with container networking via Cilium, CNI plugins, or service mesh technologies such as Istio or Linkerd.
- Experience with cloud platforms (AWS, GCP, or Azure) and their managed Kubernetes offerings.
- Bonus: experience with GPUs, bare metal infrastructure, observability tooling, or contributions to Kubernetes ecosystem open-source projects.
- Self-directed, strong communicator, and comfortable working in a fast-moving early-stage environment. COMPENSATION & BENEFITS Base salary range: $150,000 to $200,000 USD annually. Visa sponsorship is available. LOCATION On-site in San Francisco, California, United States. Candidates must be able to work full-time in person.
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →