Overview
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer / AI Builder based in Canada.
Full job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer / AI Builder based in Canada. This is a senior technical leadership role focused on designing and delivering production-grade AI and ML systems at enterprise scale. You’ll architect agentic workflows, RAG pipelines, evaluation frameworks, and cloud-native services that solve real business problems. The role combines deep hands-on engineering with architectural ownership and technical mentorship across a cross-functional team. You’ll work on live, integration-heavy projects where reliability, security, scalability, and measurable outcomes are essential. You’ll help establish engineering standards and shape how AI systems are designed, evaluated, deployed, and operated over time. The environment is fast-moving and collaborative, with production AI systems expected to move from concept to deployment in approximately 30–45 days. This is an opportunity to influence both the technology and engineering practices behind sophisticated AI solutions while remaining close to implementation.
Design and build agentic AI workflows, including reasoning loops, tool and function calling, orchestration, and single- or multi-agent architectures based on the needs of each use case. Develop and maintain production RAG pipelines covering chunking strategies, embeddings, vector search, re-ranking, knowledge freshness, and dynamic content updates. Integrate AI solutions with AWS services such as Bedrock and Agent Core, including MCP-based tools and precise tool descriptions that enable reliable orchestration. Develop production-grade system prompts using structured role definitions, output constraints, few-shot examples, and other prompt-engineering techniques. Build robust evaluation and observability capabilities using golden datasets, LLM-as-judge approaches, RAGAS-style metrics, tracing platforms, and runtime guardrails. Design for failure and degradation through retries, backoff and jitter, circuit breakers, fallback models, and clear user-facing recovery experiences. Build backend and full-stack services using Python and Node.js, including AWS Lambda, API Gateway, RESTful APIs, and serverless architectures. Design DynamoDB single-table schemas for conversation state, agent memory, and session history, while supporting event-driven workflows through Step Functions, SQS, and EventBridge. Contribute to frontend integration points, cloud deployment, monitoring, troubleshooting, and production hardening across AWS and Docker-based environments. Provide technical leadership by evaluating architectural trade-offs, challenging assumptions, mentoring engineers, and establishing strong practices for agentic AI development. Partner with product, design, and delivery stakeholders across global teams to scope, build, launch, and continuously improve AI-powered features. Own assigned systems and releases end-to-end, including debugging, reliability improvements, operational support, and maintaining production trustworthiness. Requirements 6+ years of professional software engineering experience, including substantial hands-on experience delivering GenAI or LLM-powered systems into production. Demonstrated expertise in agentic AI, including reasoning loops, tool/function calling, multi-agent orchestration, and the ability to determine when single-agent or multi-agent architectures are most appropriate. Strong practical experience with RAG, including chunking, embeddings, vector databases such as OpenSearch, similarity search, re-ranking, and knowledge-refresh strategies. Experience implementing LLM evaluation and observability, including golden datasets, LLM-as-judge, RAGAS or comparable metrics, and tracing platforms such as LangFuse or LangSmith. Strong production prompt-engineering capabilities, including the ability to create structured system prompts with meaningful constraints and examples. Hands-on experience with the AWS GenAI ecosystem, particularly Bedrock, Agent Core, Lambda, DynamoDB, S3, SQS, EventBridge, and Step Functions. Strong Python and Node.js development skills, with experience building full-stack applications and RESTful APIs. Solid understanding of reliability patterns for LLM-backed systems, including retries, backoff, circuit breakers, fallback models, and graceful degradation. Experience with Docker, cloud-native deployment, production monitoring, and troubleshooting. Ability to take ownership of ambiguous, integration-heavy technical problems and communicate architectural decisions with clarity and depth. Strong collaboration and communication skills, with the ability to work effectively with product, design, delivery, and engineering teams. Experience mentoring engineers and contributing to technical standards and engineering practices. Helpful additional experience includes Amazon Bedrock Agent Core, MCP tool integration, agent registration, memory and session management, and gateway provisioning. Experience working in regulated or high-stakes environments such as education, healthcare, or financial services is advantageous. Familiarity with workflow orchestration and state-machine-based systems, as well as practical experience running LLM observability tooling in production, is a plus. A willingness to travel for required business activities and interviews is expected, including approximately 20% travel to the United States. Benefits Competitive salary range of $176,612–$243,680 CAD . Remote position based in Canada. Opportunity to work on production-grade AI systems with significant technical ownership and influence. Exposure to advanced agentic AI, GenAI, RAG, AWS, evaluation, and observability technologies. Collaborative environment with cross-functional and globally distributed teams. Opportunities to mentor engineers and influence engineering standards and architectural direction. Work on meaningful enterprise AI solutions where reliability, security, scalability, and measurable impact are core priorities. Opportunities for professional growth through challenging technical problems and hands-on exposure to emerging AI technologies. Required travel and in-person collaboration opportunities, including final interviews and onboarding as applicable.
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
JobOpportunity.info helps you discover and organize source listings. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →Browse current JobOpportunity listings from jobgether (lever) →