Verified current Job

Member of Technical Staff - ML Systems & Inference

About us

Job Full source details
Gimlet Labs San Francisco, San Francisco, CA Source published Mar 10, 2026 Verified 2 hours ago
✓ 100% verification score · Source: Gimlet Labs (ashby) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time

Overview

About us

Full job description

About us Gimlet is building the first multi-silicon neocloud designed for fast, efficient AI inference. We combine large-scale compute infrastructure with an execution platform that partitions AI workloads and maps each stage to the hardware best suited to run it. We work with foundation labs, hyperscalers, and AI-native companies, giving our team access to technical problems spanning frontier models, production infrastructure, and emerging hardware. About the role As a Member of Technical Staff focused on ML Systems, you will build the inference systems that execute models end-to-end in production. You will work on the systems that determine how inference executes across that pipeline: how requests are batched and scheduled, how stages are placed and scaled, how KV cache and intermediate state move between accelerators, and how the system balances latency, throughput, and utilization across different hardware characteristics. You will work across model serving, batching, scheduling, concurrency, KV cache management, and memory placement. You will help bring up models on novel hardware. You will support new model architectures and inference techniques, improve performance under real production workloads, and partner with compiler, kernel, networking, and distributed systems engineers to optimize the full execution path. What success looks like In the first 12-18 months, you will:

  • Improve the latency, throughput, and efficiency of production inference workloads
  • Design execution strategies across batching, scheduling, concurrency, and resource utilization
  • Improve KV cache management, memory efficiency, and execution under load
  • Enable new models, accelerator architectures, and inference techniques to run efficiently in production You may be a good fit if you have
  • Strong software engineering fundamentals
  • Experience building or operating ML inference or model serving systems
  • Comfort reasoning about performance, memory usage, and system behavior under load
  • Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience. Strong candidates may also have
  • Experience with inference runtimes such as TensorRT-LLM, vLLM, or custom serving systems
  • Deep understanding of modern model architectures and attention mechanisms
  • Experience with batching, scheduling, and concurrency control in inference systems
  • Familiarity with KV cache management and memory placement strategies
  • Experience profiling and tuning latency- and throughput-critical systems
  • Software development experience in Python and C++ Why join now? Gimlet is expanding from its core technology into a production neocloud spanning new hardware, customers, and data centers.
  • Solve hard problems.
  • Own meaningful work.
  • Build for production.
  • Help define what’s next. Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Gimlet Labs (ashby) ↗

Browse current Job and Scholarship listings from Gimlet Labs (ashby) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books