Jobs

Search worldwide by keyword, country, category, job type and location. Select any result to review its complete source-backed details.

Clear filters
313,990 source-listed opportunitiesSelect a card to preview the complete details
Senior Software Engineer — Infra Agent Systems Source-linked
Together AI (greenhouse)
Global / location varies
Job Full-time
3w ago
Sr. Staff Physical Design Engineer Source-linked
Lightmatter (greenhouse)
Canada
Job Full-time
3w ago
Account Management, Manager Source-linked
Jobgether
India
Job Remote
1mo ago
Customer Success Manager Northern Europe Source-linked
pigment (lever)
United Kingdom
Job Full Time
3w ago
Senior Software Engineer, Customer Insights Source-linked
Together AI (greenhouse)
Global / location varies
Job Full-time
3w ago
Sr. Staff Physical Design Engineer Source-linked
Lightmatter (greenhouse)
United States
Job Full-time
3w ago
Customer Success Manager - French Mid Market Source-linked
pigment (lever)
France
Job Full Time
3w ago
Account Executive, Spokeo for Business Source-linked
Jobgether
United States
Job Remote
3w ago
Senior Program Manager, Data Center Delivery Source-linked
Together AI (greenhouse)
Global / location varies
Job
3w ago
Sr. Staff Photonics Systems Engineer Source-linked
Lightmatter (greenhouse)
United States
Job Full-time
3w ago
Sr. Staff Photonics Systems Engineer Source-linked
Lightmatter (greenhouse)
Canada
Job Full-time
3w ago
Customer Success Manager - French Mid Market Source-linked
pigment (lever)
United Kingdom
Job Full Time
3w ago
Senior Product Manager, Model APIs & Developer Experience Source-linked
Together AI (greenhouse)
Global / location varies
Job Full-time
3w ago
Account Executive, Regional Source-linked
Jobgether
United States
Job Remote
1mo ago
Customer Success Manager Source-linked
pigment (lever)
France
Job Full Time
3w ago
Senior Product Engineer, Fullstack Source-linked
Together AI (greenhouse)
Global / location varies
Job Full-time
3w ago
Signal Integrity Engineer Source-linked
Lightmatter (greenhouse)
United States
Job Full-time
3w ago
Account Executive, Enterprise (Remote Midwest) Source-linked
Jobgether
United States
Job Remote
3w ago
Customer Marketing Manager Source-linked
pigment (lever)
United Kingdom
Job Full Time
3w ago
Senior Product Engineer Source-linked
Together AI (greenhouse)
Global / location varies
Job
3w ago
Senior Photonics Design Automation Engineer Source-linked
Lightmatter (greenhouse)
United States
Job Full-time
3w ago
Customer Marketing Manager Source-linked
pigment (lever)
France
Job Full Time
3w ago
Account Executive, Enterprise (Acquisition) Source-linked
Jobgether
United Kingdom
Job Remote
4w ago
Senior Network Engineer (Amsterdam) Source-linked
Together AI (greenhouse)
Amsterdam, Netherlands
Job
3w ago
Loading opportunity details…
Job Source-linked

Senior Software Engineer — Infra Agent Systems

Together AI (greenhouse)
⌖ San Francisco ▣ Full-time Added 3 weeks ago
CountryGlobal
Job typeFull-time
Work / event modeSee source
DeadlineNot specified

About this job

About the Role Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure. We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling. You’ll work across two areas: Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack. Core Agent Platform — Build the platform that power

Responsibilities & complete job details

Full Job Description

About the Role

Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure.

We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling.

You’ll work across two areas:

Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack.

Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve.

We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure.

This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation, solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure.

Why this Role

You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective.

You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure.

Responsibilities

  • Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.

  • Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.

  • Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions.

  • Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs.

  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.

  • Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops.

  • Turn what agents learn in production into reliable, reviewed software and automation.

Requirements

  • 5+ years of experience building production backend systems, distributed systems, or infrastructure platforms.

  • Strong systems design skills and experience owning significant systems from design through production.

  • Depth in at least one of the following:

  • AI agent systems, orchestration, tool use, evaluation, or grounding

  • Knowledge graphs or graph data modeling

  • Search, retrieval, ranking, RAG, or semantic search systems

  • Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems.

  • Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms.

  • Comfortable working across languages such as Go, TypeScript, Python, or Rust.

Experience in the following is a plus:

  • GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers

  • Graph databases

  • Event-driven systems and messaging platforms such as NATS or Kafka

  • Observability platforms such as Prometheus and Grafana

  • Building evaluation frameworks or improving the quality and reliability of LLM-powered systems

About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $250,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy.

About Together AI (greenhouse)

Together AI (greenhouse) is the organization associated with this source-listed listing. JobOpportunity keeps the original authoritative source attached to every record so applicants can verify final requirements directly.

Source-first verification. JobOpportunity helps you discover and organize opportunities. Always confirm final eligibility, compensation/funding, dates and application instructions on the official source before submitting.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books