Modal Labs
Member of Technical Staff - Storage
San Francisco · Posted 1h ago
Job Description
About Us:
AI needs a new infrastructure layer. We're building it at Modal.
Every era of computing brought new workloads that previous infrastructure couldn't support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.
Our customers include category-defining companies like Lovable, Ramp, Cognition, DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.
We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We've crossed $300M+ ARR and grown fivefold since September.
Our team includes creators of popular open-source projects (e.g.,Seaborn,Luigi), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.
The Role:
We are looking for strong engineers with experience and interest in designing, building, and maintaining the novel, high-performance systems that make up our serverless platform. Specifically, you'll be working on the distributed object storage system that underpins every container image, volume, and checkpoint on Modal: hundreds of petabytes of data, replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter. You'll make cold starts feel local when the data is hundreds of milliseconds away, designing the caching, preloading, and peer-to-peer layers that hide object-store latency and keep public ingress off saturated uplinks. You'll own durability and cost at petabyte scale, from streaming and batch replication between origins, to garbage collection over billions of objects. You'll work across the stack, from local disk and page cache to distributed blob storage and garbage collection and you'll help shape what storage becomes next as we push storage closer to workloads.
Requirements:
5+ years of experience writing high-quality production code
Experience building high-performance distributed storage or caching systems at a large scale (the more challenges you've worked through, the better)
Strong cloud skills, including deep familiarity with object storage (S3 or similar), CDNs, and their consistency, throughput, and cost characteristics
Strong knowledge of low-level operating system foundations (Linux kernel, file systems, page cache, containers, etc.)
Willingness to step into the thick of it with our on-call rotation and respond to production incidents
Nice-to-Haves:
Experience with replication, content addressing, and consistency models in multi-region or multi-cloud systems
Experience operating storage systems at scale (petabyte-scale datasets, high-throughput read/write paths, large-scale garbage collection or data migration)
Experience with data engineering at petabyte-scale.
Prior experience with Rust
Key Things the Team Is Working On:
P2P sharing of data across workers within a single datacenter to dramatically reduce ingress
Replicating data across multiple blob storage providers
Automating garbage collection across hundreds of petabytes of data
Deploying colocated storage clusters to datacenters to accelerate high-throughput customer workloads
More jobs like this
Staff Product Manager, AI Product
PathAI · Boston, MA, NYC, or Remote · $146-224K
Sr. Staff Engineer, Endpoint Security
Netskope · Santa Clara, California, United States · $125-253K
Staff Computer Vision Engineer - Search AI Product Engineering
Coupang · Mountain View, USA · $174-299K