Overview
ABOUT US:
Full job description
ABOUT US: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now. Our customers include category-defining companies like Lovable https://modal.com/blog/lovable-case-study, Ramp https://modal.com/blog/how-ramp-built-a-full-context-background-coding-agent-on-modal, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. We recently raised a $355M Series C https://modal.com/blog/modal-series-c at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September. Our team includes creators of popular open-source projects (e.g.,Seaborn https://github.com/mwaskom/seaborn,Luigi https://github.com/spotify/luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. THE ROLE: We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll make cold starts feel local when the data is hundreds of milliseconds away, designing the caching, preloading, and peer-to-peer layers that hide object-store latency and keep public ingress off saturated uplinks. You'll own durability and cost at petabyte scale, from streaming and batch replication between origins, to garbage collection over billions of objects. You'll work across the stack, from local disk and page cache to distributed blob storage and garbage collection and you'll help shape what storage becomes next as we push storage closer to workloads. REQUIREMENTS:
- 5+ years of experience writing high-quality production code
- Experience building high-performance distributed storage or caching systems at a large scale (the more challenges you've worked through, the better)
- Strong cloud skills, including deep familiarity with object storage (S3 or similar), CDNs, and their consistency, throughput, and cost characteristics
- Strong knowledge of low-level operating system foundations (Linux kernel, file systems, page cache, containers, etc.)
- Willingness to step into the thick of it with our on-call rotation and respond to production incidents NICE-TO-HAVES:
- Experience with replication, content addressing, and consistency models in multi-region or multi-cloud systems
- Experience operating storage systems at scale (petabyte-scale datasets, high-throughput read/write paths, large-scale garbage collection or data migration)
- Experience with data engineering at petabyte-scale.
- Prior experience with Rust KEY THINGS THE TEAM IS WORKING ON:
- P2P sharing of data across workers within a single datacenter to dramatically reduce ingress
- Replicating data across multiple blob storage providers
- Automating garbage collection across hundreds of petabytes of data
- Deploying colocated storage clusters to datacenters to accelerate high-throughput customer workloads
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Modal (ashby) ↗