Overview
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Platform Operations Engineer based in United State
Full job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Platform Operations Engineer based in United States. This role provides first-line operational support for a production SaaS platform within a DevOps/SRE organization. You will help ensure reliable day-to-day operations across production and software delivery environments. Your responsibilities will include monitoring systems, responding to alerts, managing deployments and builds, and troubleshooting operational issues. You will work with Linux, Kubernetes, cloud platforms, CI/CD pipelines, infrastructure-as-code, and observability tools. The position is designed for an engineer who can work independently while knowing when and how to escalate complex incidents. You will collaborate with globally distributed engineering teams and participate in scheduled operational coverage and on-call responsibilities. As you grow in the role, you will have opportunities to take on broader infrastructure, automation, reliability, and SRE responsibilities.
Monitor production and SaaS environments and proactively identify potential operational issues. Respond to monitoring alerts, conduct initial investigations, and perform appropriate remediation. Execute routine application and platform deployments using established processes and tooling. Create, trigger, and monitor software builds and release workflows. Handle incoming operational requests from engineering and other internal teams. Take ownership of operational requests through resolution or ensure effective handoff to the appropriate team. Follow documented procedures and runbooks for common operational tasks and incidents. Perform first-line troubleshooting using logs, metrics, Kubernetes tooling, and other diagnostic information. Escalate complex or higher-risk incidents to senior DevOps/SRE or development engineers when appropriate. Initiate incident or war-room coordination when required and ensure the appropriate technical teams are engaged. Perform known and approved production remediation activities, such as restarting or scaling workloads, when appropriate. Participate in follow-the-sun operational coverage and provide clear handoffs for active incidents, deployments, and unresolved requests. Participate in an on-call rotation and fulfill assigned operational coverage responsibilities. Create, maintain, and improve operational documentation and runbooks based on recurring issues and evolving procedures. Identify opportunities to improve operational processes, documentation, and recurring troubleshooting workflows. Progressively take on greater ownership across infrastructure, Kubernetes, CI/CD, automation, observability, and reliability engineering as experience develops. Requirements Approximately 1–3 years of experience in DevOps, SRE, cloud operations, infrastructure operations, production support, or a related technical role. Strong entry-level candidates with relevant hands-on experience and solid technical fundamentals may also be considered. Hands-on experience with Linux and Kubernetes. Working familiarity with most of the following: Helm, Git, CI/CD pipelines, deployment workflows, Terraform, infrastructure-as-code concepts, public cloud platforms, monitoring and logging systems, networking fundamentals, and Bash scripting. Experience with at least one major cloud platform such as AWS, Azure, or GCP; transferable cloud and infrastructure fundamentals are valued over experience with a specific provider. Familiarity with monitoring, logging, and alerting tools such as Datadog or similar platforms. Basic understanding of networking and technical troubleshooting concepts. Familiarity with databases, storage, IAM, DNS, and cloud networking is helpful but not required. Comfortable working with production systems and following controlled operational procedures. Ability to investigate technical issues, collect useful diagnostic information, and recognize when escalation is appropriate. Strong written and verbal communication skills, particularly for incident documentation, technical handoffs, and escalations. Ability to work independently during assigned shifts while collaborating effectively with a globally distributed engineering organization. Willingness to participate in on-call responsibilities and provide operational coverage as required. Strong learning mindset and willingness to develop deeper DevOps/SRE expertise over time. No specific degree or professional certification is required. Must already be authorized to work in the United States, as visa sponsorship is not specified for this position. Benefits 100% remote work opportunity. Structured operational coverage designed to support teams across regions. Participation in an on-call rotation, with on-call arrangements compensated separately or supported through time off in lieu according to the source role terms. Opportunity to develop hands-on experience across Linux, Kubernetes, cloud infrastructure, CI/CD, Terraform, monitoring, and production operations. Clear career growth path toward broader DevOps and Site Reliability Engineering responsibilities. Increasing opportunities to work on infrastructure, automation, observability, reliability engineering, and production architecture. Exposure to a globally distributed engineering environment. Opportunity to contribute to improved operational processes, runbooks, and reliability practices.
Tips for this job
Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
jobgether (lever) ↗Browse current Job and Scholarship listings from jobgether (lever) →