Overview
ABOUT FORMAL
Full job description
ABOUT FORMAL Formal is building a network security platform for the AI era. As agents gain access to sensitive data and production systems, organizations need control over the actions they take and the data they receive. We enforce those controls directly in the network path, with a shared policy layer for humans and agents across databases, infrastructure, APIs, and AI tools. We are trusted by leading companies such as Cursor http://cursor.com, Decagon https://decagon.ai, Notion http://Notion.com, and more to solve problems across data security and compliance, data quality management, and infrastructure access. Formal is backed by top-tiers VCs including Thrive Capital and Y Combinator with angel investors that include executives and founders from Datadog, Clickhouse, Plaid, and Vanta. ABOUT THE ROLE Formal secures billions of requests daily. Join a small team tackling the infrastructure behind that scale: Kafka-to-Quickwit log ingestion, fast search and APIs, and reliable cloud and on-prem deployments. You'll help lead our move to Kubernetes and make the platform easier to deploy, operate, and upgrade in our cloud and customer environments. WHAT YOU'LL DO
- Lead our migration to Kubernetes and operate reliable clusters, networking, access controls, and data services.
- Scale Kafka-to-Quickwit ingestion and search; improve throughput, indexing lag, API latency, and storage costs.
- Build repeatable on-prem installations and upgrades, including configuration, data migrations, and rollback.
- Establish useful alerts and reliability targets; lead incident response and eliminate recurring failures.
- Automate provisioning with Terraform and Helm; improve CI/CD, capacity planning, backups, and recovery. WHAT YOU NEED
- Experience building and operating production cloud infrastructure on AWS or Google Cloud using Terraform.
- Experience operating production Kubernetes clusters, including deployments, ingress, autoscaling, and upgrades.
- Strong Linux and networking fundamentals, including DNS, TLS, and load balancing.
- Experience with Go, Python, Rust, or another relevant programming language.
- Experience with CI/CD, observability, and operating data pipelines or distributed APIs in production.
- A track record of automating operational work and owning services through incidents and upgrades. NICE-TO-HAVES
- Experience deploying and supporting on-prem software across customer networks and infrastructure.
- Experience with Quickwit or similar search systems, Kafka, PostgreSQL, or Temporal.
- Experience with Helm, Bazel, self-hosted CI runners, or environments for AI coding agents. COMPENSATION
- This role offers cash compensation and a stock options grant.
- The positioning of offers within a certain range depends on various factors, including: candidate experience, qualifications, skills, business requirements and geographical location. BENEFITS (FOR U.S.-BASED FULL-TIME EMPLOYEES)
- 100% medical, dental & vision insurance coverage for you
- Partially covered for your dependents
- Flexible PTO
Tips for this job
Practical JobOpportunity guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v3
JobOpportunity is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Apply through JobOpportunity →Browse current JobOpportunity listings from formal (ashby) →