Verified current Job

Research Engineer - Evals

RESEARCH ENGINEER - EVALS

Job Full source details
Firecrawl San Francisco, San Francisco HQ; Toronto Hub Source published Sep 20, 2026 Verified 9 hours ago
✓ 100% verification score · Source: Firecrawl (ashby) · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
EmploymentFull-time

Overview

RESEARCH ENGINEER - EVALS

Full job description

RESEARCH ENGINEER - EVALS You'll build the evaluation systems that tell us whether Firecrawl actually works. That sounds simple. It isn't. Our core promise, convert any URL into clean, structured, LLM-ready data reliably, is hard to measure rigorously across millions of different websites, formats, and edge cases. As the systems we're measuring get more complex, the question "did that work?" gets harder, not easier. This isn't an eval role where you inherit a framework and run benchmarks. You'll design the metrics, build the pipelines, generate the datasets, and own the feedback loop from output quality back to model and product decisions. If you care about what "good" actually means and have the engineering depth to measure it, this is the role. Salary Range: $250,000–$290,000 USD/year (SF) / $210,000–$224,000 CAD/year (Toronto) Equity Range: Competitive equity. Details shared during the process. Location: San Francisco, CA (SF HQ) or Toronto, ON (Toronto Hub). On-site, five days a week. Job Type: Full-Time Experience: 4+ years in ML, research engineering, or data-heavy backend, with real evaluation work Work Authorization: Must be authorized to work in the United States or Canada. We're not able to sponsor US visas right now. For Canada, we'll consider sponsorship on a case-by-case basis through our Toronto Hub. ABOUT FIRECRAWL Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data. It's the boring-hard problem everyone building with LLMs eventually hits, solved. We hit 8 figures in ARR in year one and more than doubled it in year two. We have 180k+ GitHub stars, putting us in the top 50 repositories of all time, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started. We're a small team punching far above our weight, working out of SF HQ and our new Toronto Hub. Everyone here owns a real piece of the product and company, end to end, and runs it themselves. No hiding behind process or headcount. This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web. WHAT YOU'LL DO

  • Design the metrics that define what "good output" actually means across millions of sites, formats, and edge cases
  • Build the pipelines and harnesses that measure quality rigorously and at scale
  • Generate and curate the datasets that make evaluation trustworthy
  • Own the feedback loop from output quality back to model and product decisions
  • Turn "did that work?" into an answer the whole team can act on WHAT WE'RE LOOKING FOR
  • You have the engineering depth to build real evaluation systems, not just run existing ones
  • You care deeply about what "good" means and how to measure it rigorously
  • You're comfortable owning ambiguous problems where the metric itself has to be invented
  • You move fast and close the loop. You'd rather ship, measure, and iterate than perfect on paper WHAT WE'RE NOT LOOKING FOR
  • Someone who only wants to run benchmarks someone else designed
  • A pure researcher who won't build the systems, or a pure engineer who won't think about methodology
  • Someone who needs a fully-specced ticket to start A NOTE ON PACE We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings, but this role probably isn't for you. BENEFITS & PERKS AVAILABLE TO ALL EMPLOYEES
  • Salary that makes sense: $250,000–$290,000 USD/year (SF) / $210,000–$224,000 CAD/year (Toronto), based on impact, not tenure
  • Own a piece: Gain competitive equity in what you're helping build
  • Generous PTO: 15 days mandatory, anything after 24 days, just ask (holidays excluded). Take the time you need to recharge
  • Parental leave: 12 weeks fully paid, for all parents
  • Wellness stipend: $100 USD/month for the gym, therapy, massages, or whatever keeps you human
  • Learning & Development: Expense up to $1,000 USD/year toward anything that helps you grow professionally
  • Team offsites: A change of scenery, minus the trust falls
  • Sabbatical: 3 paid months off after 4 years, do something fun and new AVAILABLE TO US-BASED FULL-TIME EMPLOYEES
  • Full coverage, no red tape: Medical, dental, and vision (100% for employees, 50% for partner and kids). No weird loopholes, just care that works
  • Life & Disability insurance: Employer-paid basic life and AD&D, short-term disability, and long-term disability. Coverage for life's curveballs
  • Virtual care and a health guide: Teladoc for the couch doctor visit, plus Rightway to answer coverage questions and fight billing errors for you
  • Mental health: Talkspace, therapy and psychiatry on your schedule
  • Fertility and family building: Carrot, covering you and your partner
  • EAP: Free confidential counseling, legal and financial consults, and online will prep through Guardian
  • 401(k) plan: Retirement might be a ways off, but future-you will thank you
  • Pre-tax benefits: HSA, FSA, and commuter benefits to help your wallet out a bit
  • Supplemental options: Extra life and AD&D, accident, critical illness, hospital indemnity, plus pet, legal, and identity protection through MetLife AVAILABLE TO CANADA-BASED FULL-TIME EMPLOYEES
  • Full coverage, no red tape: Extended health, dental, and vision through Manulife (Diamond, the top tier), 100% employer-paid for you, your partner, and your kids
  • Life & Disability insurance: Employer-paid life, AD&D, short-term disability, and long-term disability. Coverage for life's curveballs
  • Virtual care: Dialogue Premium, so you can see a doctor or nurse from your couch, any hour
  • Mental health: Talkspace Elite, therapy and psychiatry on your schedule
  • Fertility and family building: Carrot, covering you and your partner
  • Retirement: Group RRSP through Wealthsimple, so future-you can thank you AVAILABLE TO SF-BASED EMPLOYEES
  • SF HQ perks: Snacks, drinks, team lunches, intense ping pong, and peak startup energy
  • E-Bike transportation: A loaner electric bike to get you around the city, on us AVAILABLE TO TORONTO-BASED EMPLOYEES
  • Toronto Hub perks: Snacks, drinks, team lunches, glass-walled views down University Avenue, and a home base steps from Union Station
  • Transit, covered: A PRESTO card loaded for GO Transit, subway, and streetcar, plus station parking if you drive to the train. Winter-proof, on us INTERVIEW PROCESS Application Review: Send us your work and a quick note on why this excites you. Show us what you've built: eval systems, metrics you designed, datasets you created, quality problems you measured. We care about what you've shipped, not where you went to school. Intro Chat (~25 min): A quick conversation to get to know each other before we go deep. We'll talk about what you've been working on, what drew you to Firecrawl, and what you're looking for in your next role. Time for your questions too. Technical Chat (~45 min): We'll dig into a real problem from our world: how you'd measure whether messy web output is actually "good," and build the system to prove it. Come ready to think out loud. We care how you reason, not whether you memorized the answer. Founder Chat (~25 min): Culture, pace, ownership, and how you like to work. Time for your questions too. Paid Work Trial (~40 Hours): Work with the team on a real, scoped evaluation problem, paid at a contractor rate. It's the truest signal for both sides. You see what building at Firecrawl actually feels like, and we see how you ship. Remote-friendly, and we'll flex around your current commitments. Decision: We move fast after the trial. If you care about what "good" actually means and have the engineering depth to measure it, you should join us. 👉 Apply now.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

laptop-ats-crawler v2

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Firecrawl (ashby) ↗

Browse current Job and Scholarship listings from Firecrawl (ashby) →

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books