Overview
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior MLOps Platform Engineer based in United Sta
Full job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior MLOps Platform Engineer based in United States. This role is an opportunity to build and operate the infrastructure powering next-generation Agentic AI services at scale. You will own the end-to-end MLOps foundation across on-premise Kubernetes environments and AWS. Working alongside data scientists, product engineers, and SREs, you will turn experimental AI models into reliable production services. Your work will shape architecture, automation, observability, security, and performance for mission-critical applications. You will help enable thousands of concurrent AI agents through highly available and governed platform capabilities. The environment emphasizes strong engineering practices, reusable tooling, and continuous improvement. A flexible schedule is also available within the pay period to support work-life balance.
Design, implement, and operate a unified MLOps platform spanning on-premise Kubernetes clusters and AWS, enabling rapid onboarding and consistent governance for Agentic AI services. Develop reusable GitLab CI and CI/CD pipelines covering model packaging, containerization, automated testing, canary deployments, and rollbacks. Build and maintain observability, monitoring, and alerting capabilities using tools such as Prometheus, Grafana, OpenTelemetry, and CloudWatch to track latency, throughput, resource utilization, and data drift. Create self-service CLIs, SDKs, and dashboards that allow data science and product teams to register models, configure inference endpoints, and manage versions with minimal DevOps support. Architect and maintain robust data pipelines for training data, model artifacts, and inference logs across S3 and on-premise object storage. Partner with research and product engineering teams to transform AI prototypes into production-grade services with strong reproducibility, security, compliance, and maintainability. Optimize inference performance through GPU/CPU scaling, model quantization, batching strategies, and other techniques for high-throughput workloads. Champion security, cost efficiency, disaster recovery, and operational best practices across hybrid infrastructure, including IAM, network policies, and secret management. Mentor junior engineers and contribute to technical documentation, knowledge sharing, skills development, and engineering reviews. Requirements: Bachelor’s degree in Computer Science, Engineering, or a related technical field. 5+ years of experience building and operating production-grade software infrastructure, preferably across hybrid on-premise and cloud environments. Deep expertise with Kubernetes, including cluster provisioning, Helm, operators, custom resources, and container runtimes such as Docker and OCI. Hands-on experience with AWS services including EKS, SageMaker, S3, IAM, CloudWatch, and Step Functions, with the ability to connect on-premise resources to AWS through VPN or Direct Connect. Strong Python software engineering skills plus proficiency in at least one compiled language such as Go, Rust, or Java for developing platform components and SDKs. Experience with CI/CD and GitOps tooling such as Argo CD, Flux, GitLab, or GitHub Actions. Strong understanding of distributed systems, including consensus, fault tolerance, load balancing, and performance tuning for high-throughput, low-latency inference pipelines. Experience with data engineering technologies such as Airflow, Prefect, Kafka, Spark, or Flink and the ability to build robust, versioned data pipelines. Familiarity with observability platforms including Prometheus, Grafana, OpenTelemetry, or ELK, along with experience defining meaningful SLIs and SLOs for AI services. Demonstrated ability to collaborate with research and product teams to transition experimental code into scalable, maintainable production services. Strong problem-solving skills, excellent written and verbal communication, and enthusiasm for building scalable AI infrastructure. Working knowledge of Scrum and Agile software development methodologies is preferred. Ability to work remotely from within the United States. No visa sponsorship is available, and access to export-controlled information may require U.S. Person status or an applicable special license. Successful completion of required pre-employment screening, including credit, background, and drug screening. Benefits: Remote position within the United States, supporting teams in Aurora, Colorado and King of Prussia, Pennsylvania. Flexible scheduling options within the pay period to support work-life balance. Comprehensive medical, dental, and vision insurance. Employer contributions to eligible HSA accounts. 401(k) retirement plan with strong employer contributions. Three weeks of vacation accrual per year, plus sick leave and time off for unscheduled life events. 13 paid holidays. Upfront tuition assistance for approved degree programs. Annual bonus program based on company and employee performance. Company-paid life insurance, AD&D, short-term disability, and long-term disability coverage. Four weeks of paid parental leave. Employee Assistance Program (EAP).
Tips for this job
Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
jobgether (lever) ↗Browse current Job and Scholarship listings from jobgether (lever) →