Senior Site Reliability Engineer
Cross River
- Location
- Remote
- Work model
- Remote
- Level
- Senior
- Salary
- $160k – $200k/yr
- Posted
- 51m ago
Skills
About this role
Who We Are
Cross River builds the infrastructure behind the world’s most innovative financial products. Our technology and capital solutions power payments, cards, lending, and digital asset capabilities that move money safely, instantly, and inclusively — trusted by leading fintechs, enterprises, and disruptors across the globe.
Our mission is simple: to build the financial infrastructure that expands access and opportunity for all. Guided by a culture of collaboration, curiosity, and purpose, Cross River has been named one of American Banker’s Best Places to Work in Fintech year after year. Whether you’re designing code, solving regulatory puzzles, or developing strategy, you’ll join a team where innovation and integrity drive everything we do — and where your work helps shape the future of finance.
What We're Looking For
We are seeking a highly skilled and motivated Senior Site Reliability Engineer with 8+ years of hands-on experience ensuring the reliability, scalability, and performance of mission-critical systems. The ideal candidate brings deep expertise in building and maintaining production infrastructure, establishing DevOps best practices, and driving operational excellence across engineering teams. We're looking for someone who takes ownership of system reliability, thrives in a collaborative and fast-paced environment, and is passionate about building resilient financial infrastructure.
Responsibilities
• Define and enforce DevOps guardrails, standards, and best practices to ensure consistency, security, and compliance across engineering teams
• Enable Engineering teams to design, implement, and maintain CI/CD pipelines best to enable fast, safe, and repeatable deployments across all environments
• Co-Develop and maintain Infrastructure as Code (IaC) with Application teams using tools such as Terraform
• Establish and govern deployment strategies including blue/green, canary, and rolling deployments
• Build and maintain developer self-service tooling and internal platforms that accelerate delivery while maintaining governance
• Champion a "shift-left" culture by embedding reliability, security, and observability practices early in the software development lifecycle
• Help define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical services
• Build and maintain comprehensive observability stacks including centralized logging, metrics, distributed tracing, and alerting using tools such as New Relic, ELK, Prometheus, and Grafana
• Lead incident response and management, including on-call rotations, root cause analysis (RCA), and blameless post-mortems
• Perform capacity planning and performance engineering to ensure systems scale efficiently with business growth
• Identify and eliminate toil through automation, reducing manual operational overhead
• Conduct reliability reviews and chaos engineering exercises to proactively identify and mitigate failure modes
• Manage and optimize cloud infrastructure to balance reliability, cost, and performance
• Collaborate with software engineering teams to improve system architecture, resiliency patterns, and fault tolerance
• Support and maintain .NET-based services running across Windows and Linux environments
Qualifications
Area
<td