Verified current Job

Database Reliability Engineer

We are hiring a Senior Database Reliability Engineer to join the Infrastructure DBA cell. This is a hands-on production ownership role, not a narrow ticket-processing DBA position. You will keep critical database services reliable...

Job Remote Full source details
Alex Staff Agency Serbia Source published Jul 29, 2026 Verified 10 minutes ago
✓ 92% verification score · Source: Alex Staff Agency Careers Careers · Always confirm final requirements on the original source.
Complete source information imported The available role or programme description, requirements, benefits and source facts were imported from the public official endpoint and formatted for reading.
Database Reliability Engineer at Alex Staff Agency
EmploymentFull Time
Work modeRemote / location-flexible
CountrySerbia
Job functionEngineering
IndustryInformation Technology and Services

Overview

We are hiring a Senior Database Reliability Engineer to join the Infrastructure DBA cell. This is a hands-on production ownership role, not a narrow ticket-processing DBA position. You will keep critical database services reliable, automate repeated work, support engineering teams, and reduce single-person dependency in our PostgreSQL, ClickHouse, MongoDB, and Redis operations. PostgreSQL is the main requirement. ClickHouse experience is a strong plus, but it is not a day-one blocker. We need a senior engineer with enough database, Linux, automation, and incident-response depth to learn our ClickHouse environment quickly and operate it safely. Your Responsibilities: Own production PostgreSQL reliability: HA design, Patroni, PgBouncer, replication, failover, upgrades, vacuum/bloat control, query tuning, locks, indexes, capacity, backups, PITR, and restore validation. Improve disaster reco

Full job description

Full Job Description

We are hiring a Senior Database Reliability Engineer to join the Infrastructure DBA cell. This is a hands-on production ownership role, not a narrow ticket-processing DBA position. You will keep critical database services reliable, automate repeated work, support engineering teams, and reduce single-person dependency in our PostgreSQL, ClickHouse, MongoDB, and Redis operations.

PostgreSQL is the main requirement. ClickHouse experience is a strong plus, but it is not a day-one blocker. We need a senior engineer with enough database, Linux, automation, and incident-response depth to learn our ClickHouse environment quickly and operate it safely.

Your Responsibilities:

Own production PostgreSQL reliability: HA design, Patroni, PgBouncer, replication, failover, upgrades, vacuum/bloat control, query tuning, locks, indexes, capacity, backups, PITR, and restore validation. Improve disaster recovery and operational evidence: tested restores, documented recovery paths, measurable RTO/RPO targets, runbooks, and safe maintenance plans. Support the wider database estate: ClickHouse, MongoDB, and Redis. You will troubleshoot incidents, review access and data-safety changes, improve monitoring, and learn the production ClickHouse patterns already in use. Automate DBA workflows with Ansible, Terraform/OpenTofu, GitLab CI/CD, scripts, and reproducible runbooks for provisioning, grants, backups, restores, health checks, and ownership metadata. Help build DBaaS-style self-service capabilities so engineering teams can request databases, access, credentials, and operational checks with less manual DBA intervention. Improve observability and incident response through Grafana, metrics, logs, SLOs, alert rules, Opsgenie routing, and clear communication during production issues.

What Success Looks Like:

PostgreSQL clusters have tested backup and restore paths, useful dashboards, clear ownership, and documented failover procedures. Repeated DBA tickets become automation or self-service workflows. ClickHouse operational knowledge is no longer a single-person dependency. Database incidents have owners, runbooks, evidence, and measurable recovery paths. Product and engineering teams get database help faster without sacrificing safety, auditability, or reliability.

Why Us?

You will work on real production infrastructure used across products. You will have a direct impact on reliability, incident response, developer experience, and operational resilience. You will also work in an AI-assisted engineering culture where automation, documentation, Claude, Codex, and careful human verification are part of the daily operating model.

Requirements

What We Expect From You:

Deep hands-on PostgreSQL experience in business-critical production environments, typically 5+ years or equivalent depth. Strong understanding of PostgreSQL internals and operations: MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, major upgrades, backups, PITR, and restore testing. Proven experience with highly available databases and the ability to reason about quorum, split-brain risk, failover, rollback, and recovery. Strong Linux and infrastructure fundamentals: systemd, networking, storage, filesystems, CPU/memory/disk bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting. Automation skills with Ansible and scripting. Terraform/OpenTofu, GitLab CI/CD, and merge-request based delivery are strong advantages. Ability to support more than one database engine. You do not need to be a ClickHouse expert on day one, but you must be ready to learn it quickly and take responsibility for it. Practical use of AI engineering assistants such as Claude and Codex. We expect you to use them to improve speed and quality, while personally verifying generated SQL, commands, scripts, and operational conclusions. English - upper-intermediate or higher - to ensure clear communication of progress within the teams.

Nice to Have:

ClickHouse operations: replication, Keeper/ZooKeeper, MergeTree engines, distributed DDL, grants, row policies, backups, query troubleshooting, and cluster recovery. MongoDB replica sets and Percona Backup for MongoDB. Redis/Sentinel and broker/cache failure modes. Database observability, SLOs, golden signals, alert tuning, and executable incident runbooks. Building internal platforms, self-service portals, or DBaaS workflows for engineering teams.

Benefits

A focus on professional development. Interesting and challenging projects. Fully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwide. Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves. Compensation for private medical insurance. Co-working and gym/sports reimbursement. Budget for education. The opportunity to receive a reward for the most innovative idea that the company can patent.

Tips for this job

Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.

  1. Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
  2. Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
  3. Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
  4. Apply through the original employer or official recruitment destination shown on this page.

Verification notes

Discovered directly from the employer’s public Workable account endpoint with details=true where supported. Public description, responsibilities, requirements and benefits were normalized into complete candidate-facing sections.

Original authoritative source

Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.

Alex Staff Agency Careers Careers ↗

Browse current Job and Scholarship listings from Alex Staff Agency Careers Careers →

Related opportunities

Other current verified records you may want to review.

Job

Onsite Medical Representative

Amazon.com Services LLC · United States

Join Amazon’s mission to become Earth’s safest place to work! At Amazon, we’ve set the ambitious goal to become the benchmark of safety excellenc...

Job

Sr. Technical Account Manager, ES - WWPS

Amazon Web Services, Inc. · United States

As part of the AWS Applied AI Solutions organization, we have a vision to provide business applications, leveraging Amazon’s unique experience an...

Job

Principal Business Developer, Strategic Vendor Acceleration

Amazon.com Services LLC - A57 · United States

We are looking for a Principal Business Developer (External Business Developer / EBD) to serve as the face of Amazon's global relationship with a...

Job

Senior Partner Solutions Architect, Global Strategic Partners (GSP)

AWS EMEA SARL (UK Branch) · United Kingdom

This is an ideal role for someone who has experience working for a global systems integrator, a large IT consulting firm, or a Fortune 1000 compa...

Job

Associate Director, Delivery Lead - Content Innovation

Audible Limited (UK) - B14 · United Kingdom

At Audible, we believe stories have the power to transform lives. It’s why we work with some of the world’s leading creators to produce and share...

Job

Front End Engineer, Amazon Security Experiences

Amazon.com Services LLC · United States

Amazon Security Experiences is building the unified security platform for every builder at Amazon. Our product brings together security reviews,...

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books