Verified current Job

Senior Site Reliability Engineer - India

All roles at JumpCloud® are Remote unless otherwise specified in the Job Description.

Job Remote Full source details
Jumpcloud Bangalore, Bangalore, India - Remote Source published Sep 29, 2026 Verified 2 hours ago
✓ 100% verification score · Source: Jumpcloud (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull Time
Work modeRemote / location-flexible

Overview

All roles at JumpCloud® are Remote unless otherwise specified in the Job Description.

Full job description

All roles at JumpCloud® are Remote unless otherwise specified in the Job Description. About JumpCloud® JumpCloud® is the AI-powered unified IT management platform designed to secure the modern workforce. By consolidating identity, device, and access management, JumpCloud provides intelligent, secure IT that scales from human users to autonomous AI agents. We help organizations around the globe eliminate complexity and turn AI risk into an optimized advantage, ensuring the right people and agents have secure access to the right resources at all times.

JumpCloud is Intelligent, Secure IT.

Architect, scale, and continuously improve the reliability, availability, and performance of JumpCloud’s multi-region microservices, APIs, and authentication infrastructure (AWS/GCP). Architect, build, and maintain Disaster Recovery (DR) process, multi-region failover automation, and business continuity strategies to ensure rapid recovery against strict RTO and RPO objectives. Lead the design and enforcement of SLIs, SLOs, and Error Budget frameworks across multi-disciplinary engineering teams. Drive end-to-end observability strategy using Datadog, implementing actionable Golden Signals monitoring to drastically reduce MTTD/MTTR and eliminate alert fatigue. Lead on-call escalation, major incident management, and drive strict adherence to 99.99% availability SLAs. Facilitate blameless post-incident reviews, executing systemic root-cause remediations to prevent recurring failure modes. Architect, manage, and scale production Kubernetes (EKS) clusters, implementing advanced GitOps workflows (Argo CD, Kargo) and deployment patterns. Design and maintain modular, enterprise-grade Infrastructure-as-Code using Terraform across multi-account, multi-region cloud environments. Design, build, and maintain interactive FinOps and cost-optimization dashboards to provide engineering and leadership teams with actionable insights into multi-cloud spend, unit economics, and resource utilization. Eliminate complex operational toil by writing production-grade Python or Go tooling, platform automation, and custom integrations. Champion AI-assisted software development workflows (Cursor, Claude Code, GitHub Copilot) to accelerate automation, runbook creation, and incident triage across the team. Author operational runbooks, architecture decision records, and mentor mid-level/junior engineers to raise the overall technical bar.

8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical, highly available distributed systems. Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline. Strong Python/Go Capabilities: Advanced software engineering skill set for writing internal SRE platforms, tools, and API integrations. Deep Kubernetes Expertise: Hands-on experience with production EKS/GKE cluster lifecycles, ingress/egress, networking, RBAC, and GitOps tooling (Argo CD). Advanced IaC & AWS/GCP: Deep Terraform proficiency (module architecture, state management refactoring) across complex multi-account AWS environments (IAM, VPCs, Transit Gateway, ALB/NLB, Route53). FinOps & Cost Optimization Leadership: Demonstrated experience driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and building FinOps dashboards to embed financial accountability into engineering workflows. Disaster Recovery & High Availability: Proven background in designing and testing multi-region Disaster Recovery architectures, automating failover systems, and monitoring recovery health via DR dashboards. Observability & Reliability Architecture: Track record of defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms. Experience designing and operating enterprise service meshes (Istio, Linkerd, or similar) and production ingress/proxy systems (HAProxy, NGINX, or similar). Technical Mentorship: Demonstrated ability to lead technical discussions, write architectural design docs/RFCs, and mentor engineering peers. Strong problem-solving, communication, and collaboration skills with a passion for solving complex distributed systems challenges at scale. A strong team player who helps us live by our core values: building connections, thinking big, and getting 1% better every day.

Basic understanding of chaos engineering principles or testing resilience in staging/production. Experience with secrets management architectures (Vault, AWS Secrets Manager, External Secrets Operator, Cert-Manager). Background in DevSecOps practices, service meshes (Istio), and automated vulnerability remediation within cloud infrastructure code. Background supporting identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from Jumpcloud (lever) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books