Overview
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and oper
Full job description
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster. We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI. We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe. ABOUT THIS ROLE: Crusoe Cloud is seeking a Senior Manager Engineering, Network Operations to own production reliability for our global network infrastructure, spanning edge, backbone, data center fabric, and GPU cluster interconnects. As our footprint of high-performance compute (HPC) and GPU-based AI infrastructure scales, you will build and lead the team responsible for incident response, root cause analysis, and operational excellence across a growing, 24x7, multi-region network. You're accountable for the org's operational output: uptime, mean-time-to-resolution, and the rigor of every postmortem. You build and scale the operations team, set the standards for how incidents are handled and prevented, and are a trusted voice in planning conversations with sister Network Operations, Network Deployment, Architecture, and Site Reliability leadership. WHAT YOU'LL BE WORKING ON:
- Build and Lead the Organization: Own the operations org's structure and growth. Hire and develop Senior and Staff Network Production Operations Engineers across multiple sites and time zones, and build the management layer beneath you as the on-call footprint and device fleet scale.
- Own Production Reliability: Be accountable for uptime across Crusoe's global edge, backbone, data center, and GPU cluster network, directly supporting AI workloads running across thousands of GPUs worldwide.
- Own Incident Response: Set the org's incident command structure, escalation paths, and on-call model. Serve as an executive escalation point during high-severity network events, and ensure stakeholder communication is fast, clear, and accurate.
- Own Root Cause Analysis and Remediation: Hold the org accountable for rigorous RCAs and for tracking remediation plans through to closure. Run cross-team post-incident reviews and drive systemic fixes in partnership with Architecture and Deployment.
- Drive Operational Automation: Champion the shift from manual triage to automated detection, diagnosis, and remediation. Direct investment in tooling built on top of Crusoe's observability stack (streaming telemetry, SNMP, NetFlow, Kentik, Grafana, Prometheus, ThousandEyes) to reduce toil and accelerate MTTR.
- Own SLIs/SLOs: Partner with Architecture and SRE to define, track, and report on network reliability metrics. Hold the org to those targets and escalate when they're at risk.
- Run Executive-Level Cadence: Own operational reporting and risk escalation to Deployment, Architecture, and Site Reliability leadership. Represent the operations org in cross-functional and executive planning.
- Champion Operational Standards: Ensure runbooks, escalation playbooks, and SOPs are current, followed, and continuously improved, with an eye toward what can be scripted or automated rather than performed manually.
- Develop the Next Layer of Leaders: Coach engineers and emerging leads into stronger technical and operational leaders, and build the org's bench strength for future scale.
- Partner Across the Network Organization: Work closely with Deployment leadership on handoff quality from build to production, and with Architecture on feedback from the field that should shape reference designs and standards. WHAT YOU'LL BRING TO THE TEAM:
- 10+ years of experience in network engineering or network operations, including 4+ years managing engineers or engineering managers in a large-scale, 24x7 infrastructure environment.
- Proven People and Org Leadership: Track record of building, scaling, and retaining high-performing operations teams. Experience managing managers or team leads is strongly preferred.
- Incident and Reliability Ownership at Scale: Demonstrated ability to own production reliability across large, multi-region device fleets, with the reporting and risk management to match at an executive level.
- Strong Technical Fluency: Expert working knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments; familiarity with RDMA/RoCE lossless fabrics (PFC, ECN, DCQCN) for GPU/HPC workloads; comfort with Arista (EOS) and Juniper (Junos) platforms in leaf-spine CLOS architectures, enough to credibly challenge your own team and partner teams without needing to be the org's deepest technical expert.
- Automation Fluency: Working understanding of Python-based tooling, observability automation, and auto-remediation practices, sufficient to guide investment decisions and evaluate tooling proposals.
- Executive Communication: Comfortable setting context, managing expectations, and delivering hard news to senior stakeholders across Deployment, Architecture, and Supply Chain, especially during and after high-severity incidents.
- Operational and Strategic Judgment: Able to balance near-term incident pressure against longer-term org health, reliability standards, and team sustainability.
- Education: Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience in hyperscale or ISP environments. BONUS POINTS:
- Experience with NVIDIA/Mellanox networking platforms in GPU cluster environments.
- Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility.
- Experience defining SLIs/SLOs in partnership with SRE or product teams.
- Exposure to operating 10K+ device fleets across hyperscale or cloud environments.
- Background building or scaling post-incident learning programs org-wide. Benefits:
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location Compensation Range Compensation will be paid in the range of up to $230,000 - $280,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data. Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →Browse current JobOpportunity listings from Crusoe (ashby) →