Verified current Job

Director, Site Reliability Engineering

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Director, Site Reliability Engineering based in Ca

Job Remote Full source details
Jobgether Canada Source published Sep 23, 2026 Verified 3 hours ago
✓ 100% verification score · Source: jobgether (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible
CountryCanada

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Director, Site Reliability Engineering based in Ca

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Director, Site Reliability Engineering based in Canada. As Director, Site Reliability Engineering, you will lead the teams and technical strategy responsible for keeping critical infrastructure reliable, scalable, and secure for millions of users. You will work across software, systems, automation, cloud infrastructure, and operational processes to solve complex reliability challenges at scale. The role combines strategic leadership with hands-on technical depth, including troubleshooting production systems and partnering closely with software engineering teams. You will help shape the future architecture and deployment practices of a large-scale, privacy-focused technology environment. You will also drive improvements in automation, observability, incident response, and engineering efficiency. This is a remote-first leadership opportunity with significant ownership, autonomy, and impact.

Lead and develop Site Reliability Engineering teams responsible for the reliability, scalability, performance, and operational health of large-scale systems. Define and execute the technical direction for infrastructure, deployment, reliability engineering, automation, and operational practices. Lead high-impact and complex initiatives from initial proposal and planning through implementation, measurement, and postmortem. Investigate and resolve sources of instability across high-traffic, distributed systems, identifying root causes and implementing sustainable remediation. Establish and improve tools, services, monitoring, alerts, incident-response processes, and operational practices that identify and mitigate reliability risks. Partner closely with software engineers to troubleshoot production issues, evaluate performance considerations, and implement appropriate code-level or infrastructure-level solutions. Drive automation for infrastructure provisioning and configuration management to improve efficiency, scalability, consistency, and reliability. Leverage cloud-native architectures and services to strengthen system resilience and support continued growth. Help ensure products and infrastructure meet established reliability standards while minimizing user impact during failures and incidents. Identify emerging technical needs and opportunities to guide the long-term evolution of deployment and infrastructure architecture. Support a culture of ownership, continuous improvement, measurable outcomes, and effective post-incident learning. Requirements: 10+ years of relevant professional experience in Site Reliability Engineering, platform engineering, infrastructure engineering, software engineering, or related fields. 4+ years of experience leading SRE or comparable engineering teams. Experience participating in or managing 24/7 on-call operations for large-scale production environments. Advanced programming experience and the ability to read, write, troubleshoot, and deploy software across production systems. Strong experience with Linux administration and troubleshooting, web technologies, distributed systems, and high-traffic production environments. Demonstrated ability to lead complex technical projects from ambiguous initial requirements through execution and postmortem. Experience developing effective reliability tooling, services, monitoring, alerting, and incident-response capabilities. Strong investigative and root-cause analysis skills, particularly within distributed and high-scale systems. Experience designing and implementing infrastructure automation, provisioning, and configuration-management solutions. Hands-on experience with cloud-native services and architectures, including application packaging and deployment using Docker and Docker Compose. Experience with high-level programming languages such as Go, Perl, TypeScript, Python, or comparable technologies. Experience with AI-driven software development, including the design and implementation of agentic workflows. Strong ability to turn ambiguous or complex problems into practical, innovative solutions with measurable outcomes. Strategic thinking and technical foresight, with the ability to anticipate future infrastructure and reliability requirements. Excellent communication and collaboration skills, with the ability to work effectively across engineering teams and technical disciplines. Strong sense of ownership, autonomy, and accountability in a remote-first working environment. Benefits: Annual compensation of $243,800 USD , plus stock options. Transparent compensation structure, with team members at the same professional level and within the same global region receiving the same compensation regardless of functional team, location, gender, educational background, or years of experience. Fully remote, flexible working arrangement with no core working hours. Average full-time commitment of approximately 40 hours per week. Company-sponsored health benefits for eligible team members based in the United States; these benefits do not extend to team members based in Canada or other countries. Paid parental leave. Support for home-office setup. Co-working allowances. Opportunities to participate in company-wide and team gatherings, with travel expected at least twice per year for an all-hands meeting and a team retreat. Remote-first environment centered on trust, inclusivity, ownership, and empowered project management. Equal employment opportunities and a commitment to an accessible, inclusive hiring process. Reasonable accommodations are available for candidates who require support during the application process. Successful candidates must complete a background check as a condition of employment. The role requires participation in video meetings with cameras enabled.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from jobgether (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books