Verified current Job

Member of Technical Staff | ML Systems

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Member of Technical Staff | ML Systems based in Br

Job Remote Full source details
Jobgether Source published Sep 23, 2026 Verified 5 days ago
✓ 100% verification score · Source: jobgether (lever) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time
Work modeRemote / location-flexible

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Member of Technical Staff | ML Systems based in Br

Full job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Member of Technical Staff | ML Systems based in Brazil. This role sits at the intersection of machine learning, distributed systems, data infrastructure, and production engineering. You’ll join an ML Systems team responsible for turning research outputs into reliable, reproducible, and governed model releases. Your work will support researchers and platform engineers by creating the infrastructure and tooling they need to experiment and ship efficiently. You’ll work on GPU performance, distributed training, data materialization, experiment tracking, model lineage, and release governance. The role combines deep technical ownership with a strong internal-product mindset, where platform users and measurable engineering outcomes guide priorities. You’ll help ensure models can run reliably across cloud, production, batch, and customer environments while remaining fully auditable. It’s an opportunity for a strong systems engineer to have broad technical scope and directly shape the foundations of an ML platform.

Build high-performance CUDA kernels and compute primitives supporting the training and serving of graph neural networks. Evolve distributed sampling and training infrastructure, including neighbor sampling and performance improvements for multi-node workloads. Develop and maintain the systems used to efficiently train and serve machine learning models at scale. Define binary and columnar data formats, including Lance, Arrow, and CSR/CSC representations, and own data materialization and feature backfills for training and evaluation. Establish data contracts and consumption requirements in collaboration with teams responsible for customer and proprietary datasets. Build and operate experiment tracking, checkpointing, and evaluation infrastructure with reproducibility as a default requirement. Own model registry, lineage, versioning, and compatibility across models, embeddings, and downstream models. Define and operate release gates that ensure every production, batch, or on-premise model corresponds to an authorized and governed release. Make model releases fully auditable by maintaining traceability across the data, code, configuration, and supporting evidence used to produce them. Improve time-to-experiment by making data, compute, and experiment tracking rapidly accessible to research teams. Improve time-to-governed-release by creating efficient and reliable paths from validated candidate models to production-ready releases. Optimize training throughput and GPU utilization across large-scale foundation model workloads. Build internal platform capabilities with the mindset of a product, focusing on the needs and productivity of researchers and platform engineers. Maintain high standards for reliability, reproducibility, performance, and operational quality across ML systems. Requirements: Strong systems engineering background combined with production-quality Python development skills. Professional experience with distributed machine learning training, including technologies such as Ray or PyTorch Distributed. Experience operating or developing multi-node GPU workloads and understanding the challenges of distributed compute. Experience working with columnar data formats and large-scale data materialization pipelines. Familiarity with ML lifecycle infrastructure, including experiment tracking, model registries, evaluation systems, checkpoints, and reproducibility tooling. Strong understanding of software engineering principles for building reliable, maintainable, production-grade infrastructure. Ability to think of internal platforms as products with real users, requirements, feedback loops, and measurable outcomes. Ability to take ownership of systems and outcomes rather than focusing narrowly on individual implementation tasks. Strong analytical and problem-solving skills, with the ability to work across data, compute, model, and infrastructure layers. A product-oriented mindset and interest in enabling researchers and engineers to work more effectively. Data science expertise is not required, provided you have strong systems engineering capabilities and an understanding of ML infrastructure. Experience developing CUDA kernels or optimizing GPU performance is a strong advantage. Knowledge of graph neural networks, graph sampling, or large-scale graph workloads is a plus. Experience with Lance, Arrow, or comparable columnar or indexed storage technologies is beneficial. Familiarity with multi-cloud GPU infrastructure, including tools such as SkyPilot, is advantageous. Experience with model governance, auditability, or regulated production environments—particularly financial services—is a plus. Benefits: Fully remote position based in Brazil. Full-time opportunity within an engineering organization focused on high-impact ML infrastructure. Broad technical ownership across ML systems, distributed computing, GPU performance, data infrastructure, and model governance. Opportunity to work on foundational machine learning systems supporting research and production workloads. Direct impact on experiment velocity, training performance, model reproducibility, and governed releases. Opportunity to work with advanced technologies including CUDA, distributed GPU workloads, graph neural networks, columnar data systems, and ML lifecycle infrastructure. Environment that values technical depth while giving engineers ownership of systems and outcomes. Internal-platform mindset, with researchers and platform engineers treated as real users whose productivity drives priorities. Opportunity to shape reliable ML infrastructure designed to operate across cloud, production, batch, and customer environments.

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from jobgether (lever) →

Related opportunities

Other current verified records you may want to review.

Job

Principal Solutions Architect, Storage, WWSO Storage SA Team

Amazon Web Services, Inc. · United States

Amazon Web Services (AWS) is proud to be the pioneer and recognized leader in Cloud Computing. Our web services provide IT infrastructure used by...

Job

Site EHS Manager II, Field Workplace Health and Safety

Amazon.com Services LLC · United States

Join Amazon’s mission to become Earth’s safest place to work! At Amazon, we’ve set the ambitious goal to become the benchmark of safety excellenc...

Job

DCO Site HRP, AWS Builder Experience Team

Amazon Web Services Canada, Inc. · Canada

Join Amazon's Builder Experience Team (BeXT) as a Data Center Operations Site HR Partner and make every day better for the Builders who power AWS...

Job

AI & Creative Ops Lead, XCM Global Studio

Amazon.com Services LLC · United States

XCM Global Studio is evolving how Amazon approaches creative content at scale. After a successful first year establishing foundational workflows...

Job

Senior Applied Scientist - Optimization, Fulfillment Planning and Execution Science - Fulfillment Optimization

Amazon.com Services LLC · United States

Have you ever wondered how Amazon predicts when your order will arrive and how we ensure that it actually arrives on at the promised date/time? H...

Job

TIPM, Network Delivery, Global Network Delivery

Amazon Data Services, Inc. · United States

AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people...

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books