Verified Remote Opportunities

Search worldwide by keyword, country, category, job type and location. Select any result to review its complete source-backed details.

Clear filters
70,594 verified opportunitiesSelect a card to preview the complete details
Research Scientist (Remote/US/LATAM)✓ Verified
Anyone Ai Careers
Mexico City, Mexico
Job Contract Remote
2mos ago
Full-Stack Developer (Uruguay)✓ Verified
Anyone Ai Careers
Montevideo, Uruguay
Job Part Time Remote
3mos ago USD 45 – 80
Full-Stack Developer (Peru)✓ Verified
Anyone Ai Careers
Lima, Peru
Job Contract Remote
3mos ago USD 45 – 80
Full-Stack Developer (Mexico)✓ Verified
Anyone Ai Careers
Ciudad de Mexico, Mexico
Job Part Time Remote
3mos ago USD 45 – 80
Full-Stack Developer (Ecuador)✓ Verified
Anyone Ai Careers
Quito, Ecuador
Job Part Time Remote
3mos ago USD 45 – 80
Full-Stack Developer - AI Trainer (Poland)✓ Verified
Anyone Ai Careers
Warsaw, Poland
Job Contract Remote
2mos ago USD 45 – 80
Full-Stack Developer (Colombia)✓ Verified
Anyone Ai Careers
Bogota, Colombia
Job Part Time Remote
3mos ago USD 45 – 80
Full-Stack Developer (Chile)✓ Verified
Anyone Ai Careers
Santiago de Chile, Chile
Job Part Time Remote
3mos ago USD 45 – 80
Full-Stack Developer (Brazil)✓ Verified
Anyone Ai Careers
São Paulo, Brazil
Job Contract Remote
3mos ago USD 45 – 80
Full-Stack Developer (Argentina)✓ Verified
Anyone Ai Careers
Bogota, Colombia
Job Contract Remote
3mos ago USD 45 – 80
Data & Operations Specialist✓ Verified
Anyone Ai Careers
Global / location varies
Job Full Time Remote
2mos ago
Full-Stack Developer - AI Trainer (Czechia)✓ Verified
Anyone Ai Careers
Prague, Czechia
Job Contract Remote
2mos ago USD 45 – 80
Software Engineering AI Trainer (Uruguay)✓ Verified
Anyone Ai Careers
Montevideo, Uruguay
Job Contract Remote
3mos ago USD 45 – 80
Software Engineering AI Trainer (Peru)✓ Verified
Anyone Ai Careers
Lima, Peru
Job Part Time Remote
3mos ago USD 45 – 80
Software Engineering AI Trainer (Paraguay)✓ Verified
Anyone Ai Careers
Asunción, Paraguay
Job Part Time Remote
3mos ago USD 45 – 80
Software Engineering AI Trainer (Mexico)✓ Verified
Anyone Ai Careers
Ciudad de Mexico, Mexico
Job Contract Remote
3mos ago USD 45 – 80
Software Engineering AI Trainer (Ecuador)✓ Verified
Anyone Ai Careers
Quito, Ecuador
Job Part Time Remote
3mos ago USD 45 – 80
Software Engineering AI Trainer (Colombia)✓ Verified
Anyone Ai Careers
Bogota, Colombia
Job Contract Remote
3mos ago USD 40 – 80
Software Engineering AI Trainer (Chile)✓ Verified
Anyone Ai Careers
Santiago de Chile, Chile
Job Part Time Remote
3mos ago USD 45 – 80
Software Engineering AI Trainer (Brazil)✓ Verified
Anyone Ai Careers
São Paulo, Brazil
Job Part Time Remote
3mos ago USD 45 – 80
Software Engineering AI Trainer (Argentina)✓ Verified
Anyone Ai Careers
Buenos Aires, Argentina
Job Contract Remote
3mos ago USD 40 – 80
Mathematics Expert - (Chile)✓ Verified
Anyone Ai Careers
Santiago de Chile, Chile
Job Contract Remote
3mos ago USD 40
Mathematics Expert - (Argentina)✓ Verified
Anyone Ai Careers
Buenos Aires, Argentina
Job Contract Remote
3mos ago USD 40
Mathematics Expert - (Uruguay)✓ Verified
Anyone Ai Careers
Montevideo, Uruguay
Job Contract Remote
3mos ago USD 40
Loading opportunity details…
Job ✓ 92% verified

Research Scientist (Remote/US/LATAM)

Anyone Ai Careers
⌖ Mexico City, Mexico ▣ Contract Remote Posted 2 months ago
CountryMexico
Job typeContract
Work / event modeRemote
DeadlineNot specified

About this job

Research Scientist, LLM Evaluations & Benchmarking Anyone AI Labs Reports to: CEO · Remote / LatAm / US The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or measuring the wrong thing. You'll own that problem at Anyone AI, measuring frontier model capability. This is a research role at heart: you decide what a good evaluation is , design the benchmarks that prove it, and defend the methodology under lab scrutiny. You'll build frontier-grade evaluation packages across reasoning, coding, agents, tool use, and multi-modal — grounded in expert-verified truth, validated against multiple models, and QC'd to survive buyer-side review. Responsibilities ● Evaluation research. Turn eval targets into original benchmark designs. Own

Department: Anyone AI Internal Team

Responsibilities & complete job details

Description

Full Job Description

Research Scientist, LLM Evaluations & Benchmarking

Anyone AI Labs Reports to: CEO · Remote / LatAm / US

The role Evaluation is one of the hardest open problems in AI: we still don't have reliable ways to measure what frontier models can and can't do, and the field mostly runs on benchmarks that are saturated, contaminated, or measuring the wrong thing. You'll own that problem at Anyone AI, measuring frontier model capability.

This is a research role at heart: you decide what a good evaluation is , design the benchmarks that prove it, and defend the methodology under lab scrutiny. You'll build frontier-grade evaluation packages across reasoning, coding, agents, tool use, and multi-modal — grounded in expert-verified truth, validated against multiple models, and QC'd to survive buyer-side review.

Responsibilities

● Evaluation research. Turn eval targets into original benchmark designs. Own the hard measurement questions: construct validity, item discrimination, headroom, reliability, contamination, and capability elicitation. Push toward evals that stay informative as models improve.

● Benchmark development. Build evaluation packages with subject-matter experts, each with expert-verified ground truth, multi-model headroom results, and rigorous QC (calibration layers, severity-weighted rubrics, deterministic verifiers).

● Experts. Recruit, calibrate, and review a pool across coding, agentic/tool-use, and STEM/reasoning. Be the final arbiter of correctness and frontier difficulty.

● Lab relationships. Be a technical point of contact for labs, with CEO support. Understand what they're trying to measure and translate it into an evaluation design.

● Delivery & dissemination. Turn lab requests into winning sample packages and own pilots end to end. Where the work generalizes, help turn it into public benchmarks and papers: we support publishing at venues like NeurIPS Datasets & Benchmarks, ICLR, and ACL.

What we're looking for

● Research background in ML evaluation or benchmarking (a track record of published or open benchmarks, eval/measurement research, or equivalent hands-on work that labs have relied on).

● Deep LLM/frontier-model benchmarking expertise, with real strength in code-model and agentic evaluation.

● Fluency with the measurement problem itself: construct validity, psychometrics, rubrics, pass rates, headroom, contamination, and what makes a task genuinely discriminate a model.

● Interest in the safety side of evaluation, capability elicitation, robustness, and measuring the things that are hardest to measure honestly.

● Proven ability to hold a team or expert pool to a rigorous standard.

● Comfort with the full research loop: framing the question, running the study, and writing it up.

● Fluent English; Spanish a plus.

About Anyone Ai Careers

Anyone Ai Careers is the organization associated with this verified listing. JobOpportunity keeps the original authoritative source attached to every record so applicants can verify final requirements directly.

Source-first verification. JobOpportunity helps you discover and organize opportunities. Always confirm final eligibility, compensation/funding, dates and application instructions on the official source before submitting.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books