yoinka

Senior Database Reliability Engineer (DBRE)

Trumid

RemoteSenior$225k – $265k/yr
Sign in to applyVerified 2h ago
Location
Remote
Work model
Remote
Level
Senior
Salary
$225k – $265k/yr
Posted
2h ago

Skills

BigQueryGrafanaKafkaKubernetesLinuxPostgreSQLPrometheusPythonSQLShellTerraform

About this role

About us

Trumid is a dynamic fintech revolutionizing the landscape of fixed income trading. With intelligent, easy-to-use, electronic solutions, we are rapidly growing and seeking exceptional talent to help redefine the boundaries of technology and finance.

Founded in 2014 by a team of fixed income market experts, Trumid has quickly become one of the top three corporate bond e-trading platforms in the U.S. Today, over 1,300 traders from an extensive and expanding client network of 890+ buy-and sell-side institutions transact on Trumid monthly.

With a rich history of innovation and a unique ability to innovate at scale, we collaborate closely with our clients, iterating quickly toward optimal solutions. With market share and client engagement at all-time highs and our pace of product development faster than ever, this is an exciting and transformative time at Trumid.

Our business model thrives on participation, and so does our company culture. We rely on every team member’s contribution to help us accomplish our goals. To succeed at Trumid, you must be curious, passionate about your craft, ambitious, collaborative, and driven. Learn more at www.trumid.com

The opportunity

This role owns data resilience and continuity, and the scope test is simple: if losing it loses data, or makes data unavailable, it’s yours. The job exists so that data-layer failure modes — storage contention, replica lag, region loss — are found and retired in drills, not discovered in production. Our reliability doctrine is to assume failure and concentrate statefulness into a small core of systems proven against specific failure modes. Postgres is the heart of that core: everything around it gets to be disruptible because the data layer is not.

You’d join the databases side of our SRE team, working alongside deep Postgres expertise. Some of what you’d walk into:

• A production RDS fleet backing a live trading venue, with performance work that goes deep: we’ve characterized WAL-write contention under concurrent commits down to the fsync level, and are weighing group-commit tuning, dedicated log volumes, and storage-class changes against actual measurements.

• Disaster recovery as an engineering discipline: automated cross-region failover with promotion measured in minutes, and restore paths (snapshot, point-in-time, logical) validated by timed, documented drills on a fixed cadence.

• Database observability and AI-assisted tooling: engine performance telemetry exported into Prometheus and Grafana, modular database health-check skills, and an automated reviewer for schema-migration PRs.

What you’ll do?

• Own the durability, recoverability, and performance of the Postgres/RDS fleet across every environment: replication, failover, backup and restore, and storage behavior under load.

• Make recovery provable: restore-tested coverage of every production database, timed failover drills, and measured RPO/RTO per tier — evidence, not assertion.

• Own the data lifecycle end to end: retention and cleanup policies that preserve recoverability, access control at the data layer, encryption posture, and knowing where sensitive data lives.

• Hunt performance pathologies at the engine level: lock contention, WAL throughput, replica lag, bloat, index hygiene, write amplification.

• Build database observability with the team, and extend our AI-assisted operations tooling (health-check skills, migration PR review).

About you

• Five or more years running production PostgreSQL at meaningful scale — ideally on RDS or Aurora — with depth in the internals: replication, WAL mechanics, MVCC and vacuum behavior, query planning and performance.

• You’ve owned the full lifecycle of

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Senior Database Reliability Engineer (DBRE) at Trumid, Remote | Yoinka