Verified current Job

Site Reliability Engineering Lead

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Lead based in the Uni

Job Remote Full source details
Jobgether Source published Oct 2, 2026 Verified 2 hours ago
✓ 100% verification score · Source: jobgether (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Lead based in the Uni

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineering Lead based in the United States. This leadership role is responsible for building reliable, scalable, and secure cloud platforms while developing a high-performing SRE team. You will lead engineers, establish priorities, and drive initiatives that improve service reliability, resilience, automation, and operational efficiency. The role combines people leadership with hands-on technical direction across modern cloud and infrastructure environments. You will work closely with Development, Security, Product, and other engineering teams to strengthen platform performance and incident response. A key focus will be reducing operational toil through automation, observability, infrastructure as code, and self-healing capabilities. You will also guide post-incident reviews, root-cause analysis, and continuous improvement efforts across services and infrastructure. This is an opportunity to shape SRE practices at scale while supporting engineers in their technical and professional growth.

Lead, mentor, and develop a small to medium-sized team of Site Reliability Engineers through regular 1:1s, performance reviews, career planning, and ongoing coaching. Own hiring, onboarding, team capacity, resourcing, and workforce planning decisions to ensure the team can effectively support business and platform priorities. Establish team objectives, prioritize the engineering backlog, coordinate planning, and ensure projects and operational tasks remain aligned with reliability goals. Lead reliability initiatives across infrastructure and services, improving availability, scalability, resilience, security, and operational performance. Drive incident response activities and facilitate blameless post-incident reviews, ensuring timely root-cause analyses and actionable follow-up. Partner with Development, Security, Product, and other engineering teams to resolve cross-functional issues and strengthen collaboration. Champion automation and operational excellence by reducing manual work, eliminating recurring toil, and introducing self-healing systems and infrastructure automation. Support the design and evolution of scalable, secure, cloud-native environments and continuously identify opportunities to improve performance, reliability, and cost efficiency. Requirements Demonstrated experience in SRE, DevOps, infrastructure engineering, or a related discipline, including experience leading engineering teams. Expert-level knowledge of Kubernetes, including cluster architecture, upgrades, autoscaling, security hardening, and large-scale troubleshooting. Advanced experience with Terraform, including modular infrastructure-as-code design, state management, multi-environment provisioning, and policy-as-code. Deep knowledge of Azure Cloud services, including compute, networking, identity and access management, storage, and cost optimization. Experience designing and scaling CI/CD pipelines using GitHub Actions, release strategies, and automated rollback approaches. Strong knowledge of observability platforms such as Prometheus, Grafana, and OpenTelemetry, along with experience managing SLOs, SLAs, and error budgets. Strong automation capabilities and advanced proficiency in Python, Bash, and/or PowerShell for infrastructure tooling and operational automation. Deep understanding of networking fundamentals, including TCP/IP, DNS, load balancing, VPNs, and cloud-native networking. Proven experience leading incident response, conducting root-cause analysis, and implementing measurable reliability improvements. Strong people leadership, communication, prioritization, and problem-solving skills, with the ability to support engineers while coordinating effectively across technical teams. Benefits U.S. national base salary range of $118,300–$219,800 , with geographic differentials potentially applying depending on location. Eligibility for an annual incentive bonus . Country- and location-specific employee benefits designed to support overall health and well-being. Support for an accessible and inclusive hiring process, including reasonable accommodations where required. Remote/home-based opportunities available in multiple U.S. locations, including Florida, Connecticut, New Jersey, New York, and Pennsylvania. Opportunity to lead a technically sophisticated SRE function focused on cloud platforms, automation, resilience, and operational excellence. Professional development opportunities through team leadership, cross-functional collaboration, and exposure to large-scale cloud environments.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from jobgether (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books