Verified current Job

MTS - Engineering (Data Infrastructure)

ABOUT THE ROLE

Job Full source details
Collinear.ai San Francisco, San Francisco, CA, Sunnyvale, California Source published Oct 1, 2026 Verified 54 minutes ago
✓ 100% verification score · Source: Collinear.ai (ashby) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time

Overview

ABOUT THE ROLE

Full job description

ABOUT THE ROLE As a Member of Technical Staff - Engineering (Data Infrastructure), you will own the systems that turn large, real-world datasets into data Collinear can build on. Our environments are grounded in terabytes of data, spread across archives, spreadsheets, email, PDFs, and scanned documents. The pace at which we process these determines how fast we deliver to frontier labs. This is a hands-on role at the intersection of algorithms, distributed systems, and data quality. Many of our hardest problems, such as linking related records across millions of files, don't split up neatly, and you will define how we solve them at scale. WHAT YOU'LL DO

  • Build pipelines that process multi-terabyte datasets in parallel across archives, spreadsheets, email, PDFs, and scanned documents
  • Design graph-based systems that link related records, such as the same person or company appearing across millions of files
  • Build fast string and pattern search over large, heterogeneous datasets
  • Profile and remove bottlenecks, and decide how to split work that doesn't parallelize neatly
  • Transform data for downstream use, including consistently replacing sensitive fields across files
  • Define how we measure data quality, and build review tools so the team can catch and fix errors without reprocessing everything
  • Assess new datasets and filter out low-quality data before it reaches our environments ABOUT YOU
  • You have 5+ years of experience building data-intensive systems in production
  • You have processed large datasets in parallel with frameworks such as Apache Spark, Ray, or Dask, and know when to design your own
  • You have a strong command of graph algorithms, and experience using them to transform large amounts of data
  • You have built efficient string matching and pattern search at scale, such as fuzzy matching or indexing
  • You make sound tradeoffs between accuracy, speed, and cost, and can explain them clearly NICE TO HAVE
  • Experience with entity resolution or record linkage
  • Experience with NLP or LLM-based information extraction
  • OCR or document processing experience, including poor scans and handwriting
  • Experience in a systems language such as Rust, C++, or Go
  • Experience with regulated or sensitive data, such as financial or healthcare records

Tips for this job

Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v3

Original authoritative source

JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Apply through JobOpportunity →

Browse current JobOpportunity listings from Collinear.ai (ashby) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books