Verified current Job

Site Reliability Engineer

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in India.

Job Remote Full source details
Jobgether Source published Oct 5, 2026 Verified 56 minutes ago
✓ 100% verification score · Source: jobgether (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in India.

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in India. This is a fully remote engineering role focused on building resilient, highly available, and scalable production systems. You will operate at the intersection of software engineering, cloud infrastructure, observability, and platform reliability. The role involves designing automation, strengthening monitoring, reducing operational toil, and improving the reliability of cloud-native services. You will work extensively with AWS and GCP, Kubernetes, Terraform, and modern GitOps practices while supporting mission-critical systems. You will also contribute to disaster recovery, incident response, and reliability engineering standards across distributed services. This opportunity is well suited to an engineer who enjoys solving complex production problems through code and continuously improving operational maturity.

Design, deploy, and maintain highly available, reliable, and performant production systems and APIs across AWS and GCP. Define and operationalize SLIs, SLOs, and error budgets in partnership with application engineering teams. Build and improve end-to-end observability across microservices and cloud infrastructure using Datadog or comparable monitoring platforms. Implement actionable monitoring around the Golden Signals—latency, traffic, errors, and saturation—to improve detection times and reduce unnecessary alerts. Participate in on-call rotations, incident response, troubleshooting, and blameless post-incident reviews, using incidents to drive systemic improvements. Manage production Kubernetes environments, including EKS clusters, container orchestration, and GitOps delivery workflows using tools such as Argo CD and Kargo. Provision, manage, and secure multi-cloud infrastructure using modular Terraform and Infrastructure-as-Code practices. Develop and maintain disaster recovery dashboards, runbooks, failover automation, and validation tests aligned with defined RTO and RPO targets. Build production-grade Python or Go automation tools and scripts to eliminate repetitive operational work and reduce engineering toil. Use AI-assisted development tools to accelerate scripting, runbook creation, troubleshooting, and incident triage. Monitor cloud infrastructure costs, support resource optimization, and improve cost visibility through appropriate allocation and FinOps practices. Configure and troubleshoot production service meshes and high-availability proxy solutions. Contribute to CI/CD, DevSecOps, resilience testing, secrets management, and other platform engineering initiatives where required. Proactively identify reliability, performance, security, and operational improvements across the production environment. Requirements 5+ years of professional software engineering experience in Site Reliability Engineering, DevOps, Platform Engineering, or a related discipline supporting 24/7 mission-critical systems. Strong hands-on programming skills in Python or Go, with experience building SRE tools, automation, scripts, and cloud integrations. Production experience operating Kubernetes clusters and containerized workloads, including experience with GitOps tools such as Argo CD. Strong Infrastructure-as-Code experience using Terraform, including writing, maintaining, and modularizing configurations. Direct experience operating cloud workloads on AWS or GCP, with knowledge of services such as EKS, IAM, VPC networking, Route 53, ALB/NLB, or comparable cloud infrastructure. Practical FinOps experience, including cost-allocation tagging, resource right-sizing, cloud cost analysis, or development of cost-visibility dashboards. Experience designing or operating disaster recovery solutions, conducting failover exercises, and monitoring recovery metrics. Strong observability and incident-management experience using Datadog or similar tools, PagerDuty, alerting systems, and SLI/SLO frameworks. Hands-on experience configuring and troubleshooting production service meshes such as Istio or equivalent technologies. Experience managing high-availability proxy solutions such as HAProxy, NGINX, or comparable platforms. Strong troubleshooting and problem-solving skills, with a demonstrated ability to improve operational efficiency through engineering and automation. Experience with CI/CD platforms such as GitHub Actions or GitLab Pipelines is preferred. Familiarity with chaos engineering or resilience testing in staging or production environments is a plus. Knowledge of secrets-management solutions such as HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator is beneficial. Basic understanding of DevSecOps practices and Infrastructure-as-Code security scanning and remediation is preferred. Strong collaboration and communication skills, with the ability to work effectively across distributed engineering teams. Fluent written and spoken English, as the role involves working with global teams and conducting interviews and business communication primarily in English. Willingness and ability to participate in an on-call rotation and respond during assigned shifts. Benefits Fully remote position for candidates based in India. Opportunity to work on highly available, mission-critical cloud-native systems at scale. Hands-on exposure to AWS, GCP, Kubernetes, Terraform, GitOps, observability, and modern reliability engineering practices. Opportunity to use AI-assisted engineering tools to improve automation, incident response, and development productivity. Significant scope to influence platform resilience, operational maturity, disaster recovery, and engineering efficiency. Collaboration with experienced, globally distributed engineering teams. Fast-paced SaaS environment offering challenging technical problems and opportunities for continuous learning. Opportunity to contribute to reliability standards and engineering practices across distributed systems. Supportive environment that values collaboration, innovation, continuous improvement, and technical ownership.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from jobgether (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books