Overview
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior AI Systems Quality Engineer based in United
Full job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior AI Systems Quality Engineer based in United States. This role focuses on engineering quality, reliability, and trust into production-grade AI systems used in mission-critical healthcare environments. You will design automated testing frameworks, evaluation pipelines, and scalable quality practices for agentic and LLM-driven applications. The position goes beyond traditional QA, emphasizing quality-by-design, safe failure behavior, measurable evaluation signals, and continuous validation throughout the development lifecycle. You will help build an AI testing platform integrated with Databricks and MLflow to support traceability, lineage, and auditability at scale. Working closely with AI engineering, platform, security, and delivery teams, you will influence release readiness and operational confidence. This is an automation-first engineering opportunity for someone passionate about making advanced AI systems reliable, governed, and production-ready.
Build and deploy production-grade automated validation frameworks, test harnesses, and evaluation pipelines across the full AI development lifecycle. Design and evolve an AI testing platform integrated with Databricks and MLflow to enable repeatable testing, traceability, lineage, and auditability. Create large-scale, scenario-based test suites covering hundreds or thousands of cases, including edge cases, long-tail scenarios, and system failure modes. Validate agentic orchestration behaviors such as tool usage, memory, decision logic, and non-deterministic outputs before production deployment. Embed quality-by-design principles by defining system contracts, guardrails, safe-degradation patterns, and validation requirements at key system boundaries. Define measurable quality signals for LLM systems, including grounding, hallucination rates, relevance, latency, cost, accuracy, and explainability. Integrate automated quality gates into CI/CD pipelines and ensure validation runs continuously following model, prompt, or code changes. Build reusable testing libraries, frameworks, and components that enable engineering teams to adopt consistent AI quality practices. Establish measurable release-readiness criteria and support go/no-go decisions based on defined quality thresholds. Partner with AI, platform, security, and delivery teams to translate business and mission requirements into clear quality criteria, trade-offs, and confidence levels. Evaluate system behavior, reliability, security, privacy, and operational risk in regulated and mission-critical environments. Requirements: 7+ years of software engineering experience, primarily focused on backend or platform systems. Proven experience designing and implementing automated AI testing and validation solutions in production environments. Demonstrated ability to build custom testing, validation, or evaluation frameworks for complex and distributed systems. Strong proficiency in Python and/or TypeScript within modern AI engineering environments. Hands-on experience with AI-powered systems, including LLM-based or agentic workflows and non-deterministic behavior. Experience designing AI testing at scale, including regression frameworks, long-tail evaluations, and broad test coverage. Deep understanding of CI/CD practices and experience embedding automated tests and quality gates into deployment pipelines. Solid knowledge of AWS cloud-native architectures. Strong track record of engineering for quality, reliability, governance, safety, and operational resilience as core system principles. Working knowledge of security, privacy, and operational risk within regulated or mission-critical environments, including failure modes and recovery strategies. Experience with AI testing methodologies such as non-deterministic output evaluation, drift detection, bias and fairness testing, and robust regression strategies. Ability to establish measurable trust thresholds and operationalize metrics such as query accuracy, hallucination limits, explainability, and PHI-safe behavior as release criteria. Experience collaborating with domain experts to define correctness and real-world validation scenarios that reflect genuine production use cases. Experience with Databricks and Medallion architecture is preferred but not required. Familiarity with MLflow for model evaluation, lineage, and auditability is a plus. Exposure to observability tools such as Datadog, Prometheus, or Grafana is desirable. Familiarity with LLM evaluation techniques, guardrails, and policy enforcement frameworks is beneficial. Experience evaluating AI performance, latency, and cost regressions is a plus. Ability to clearly communicate system behavior and quality trade-offs to both technical and business audiences. Formal AI/ML training or certifications, such as ISTQB AI Testing, AWS ML Specialty, or Google ML Engineer, are welcome. Experience designing prompts, agent behaviors, and orchestration logic as versioned, testable artifacts is advantageous. Familiarity with using AI systems to generate and expand diverse, adversarial, and large-scale test scenarios is a plus. Benefits: Compensation based on experience, skills, and location, including base salary, performance bonus eligibility, and equity grants. Unlimited paid time off. Work-from-anywhere flexibility. Comprehensive health coverage with multiple plan options. Equity for every employee. Growth-focused environment with opportunities for professional development. One-time home office setup allowance. Monthly cell phone allowance.
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →Browse current JobOpportunity listings from jobgether (lever) →