Verified current Job

Member of Technical Staff (AI Inference Engineer)

We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures a

Job Full source details
Perplexity London Source published Apr 13, 2026 Verified 4 hours ago
✓ 100% verification score · Source: perplexity (ashby) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time

Overview

We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures a

Full job description

We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL. RESPONSIBILITIES:

  • New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.
  • GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.
  • Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
  • Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernels interleaving.
  • Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents. WHO WE'RE LOOKING FOR:
  • Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus.
  • You understand modern LLM architectures and are able to bring them up reliably in a production environment.
  • You've built and operated production distributed systems under real load - ideally performance-critical ones.
  • Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.
  • You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday.
  • Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you. NICE-TO-HAVE:
  • ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.
  • Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.
  • Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.
  • Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis.
  • Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads. QUALIFICATIONS:
  • 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
  • Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow).
  • Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores).
  • Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation). Final offer amounts are determined by multiple factors including experience and expertise. Equity: In addition to the base salary, equity may be part of the total compensation package.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

perplexity (ashby) ↗

Browse current Job and Scholarship listings from perplexity (ashby) →

Related opportunities

Other current verified records you may want to review.

Job

Referendariat Wahlstation Arbeitsrecht

Bauindustrieverband Ost Careers

Current Referendariat Wahlstation Arbeitsrecht opening at Bauindustrieverband Ost Careers in Magdeburg. Full employer-published role sections hav...

Job

Buyer II - AMZ24149.3

Amazon.com Services LLC - A57 · United States

Employer: Amazon.com Services LLC Position: Buyer II - AMZ24149.3 Location: Atlanta, GA Multiple Positions Available: Delivering improved financi...

Job

Software Development Engineer, Agentic AI, Velocity Labs

Amazon Development Center U.S., Inc. · United States

The Velocity Labs team mission is to think beyond the confines of the normal product-orientated approach and to discover new ways to apply and em...

Job

Senior Marketing Manager, Global Executive Marketing, AWS Global Executive Marketing

Amazon Web Services, Inc. · United States

AWS Global Executive Marketing (GEM) is responsible for how AWS engages its most senior customer leaders. We design and deliver the strategy, pro...

Job

Sr. Product Manager - Tech, Devices and Services FinTech

Amazon.com Services LLC · United States

Are you interested in working with the teams that developed the Kindle, FireTV and Alexa? The Amazon Devices and Service Finance organization is...

Job

Software Development Manager, Agentic AI, Velocity Labs

Amazon Development Center U.S., Inc. · United States

The Velocity Labs team mission is to think beyond the confines of the normal product-orientated approach and to discover new ways to apply and em...

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books