Senior Engineer, Reliability
LPL Financial
- Location
- Austin, TX
- Work model
- On-Site
- Level
- Senior
- Posted
- Sep 11, 2026
Skills
About this role
Where Ambition Meets Innovation Build a career that matches all your initiative with an impressive dose of innovation. From cutting-edge resources and a collaborative environment to the freedom to make an impact and more, you’ll find the ingredients you need at LPL Financial to shape your success while helping clients pursue their financial goals.
Job Overview
The Senior Engineer, Observability and Platform Stability is responsible for ensuring the reliability, availability, and performance of enterprise observability platforms and supporting applications. This role drives operational excellence through proactive monitoring, incident response, platform maintenance, automation, and continuous improvement initiatives. The position partners closely with Site Reliability Engineering (SRE), Platform Engineering, DevOps, Operations, and Incident Management teams to enhance platform stability, mature CI/CD practices, and support strategic technology initiatives. The ideal candidate brings strong cloud technology experience and a proven ability to support production environments while delivering actionable insights through observability capabilities.
Responsibilities
Delivery Support Partner with Product Owners and engineering teams to provide operational and release support for technology initiatives. Ensure platform changes meet operational readiness requirements, including rollback procedures, runbook documentation, integration standards, and support handoffs. Maintain production stability throughout platform upgrades, enhancements, and enterprise initiatives. Support platform ownership transitions and operational readiness activities across global delivery teams. Production Support & Incident Response Troubleshoot application and platform issues to restore services and minimize business impact. Serve as an escalation point for complex production incidents and operational challenges. Participate in incident triage, root cause analysis, corrective action planning, and resolution activities. Collaborate with engineering teams to implement long-term solutions that reduce recurring incidents. Support mission-critical environments through timely incident response and service restoration. Platform Stability & Proactive Operations Conduct health checks, configuration reviews, and performance assessments to identify operational risks. Support reliability initiatives focused on improving recovery times, reducing incident recurrence, and minimizing change-related defects. Validate vendor releases, hotfixes, and configuration changes prior to production deployment. Partner with observability and analytics teams to enhance monitoring, alerting, and issue detection capabilities. Release Execution & CI/CD Maturation Execute platform changes through established SDLC, change management, and release management processes. Collaborate with Platform Engineering, DevOps, Quality Engineering, and Scrum teams to improve release and deployment practices. Support the adoption of source control, environment separation, release automation, and CI/CD capabilities. Ensure solutions are testable, deployable, and operationally supported before and after production implementation. Documentation & Operational Excellence Maintain runbooks, support documentation, configuration records, incident playbooks, and release procedures. Document incident findings, lessons learned, and process improvement opportunities. Contribute to the development of standardized, repeatable, and scalable operational practices. What are we looking for? We seek professionals who pursue greatness , act with integrity , are driven to help our clients succeed , win together , and create and share joy . The ideal candidate brings strong technical expertise in observability platforms, site reliability engineering, production support, release management, and operational excellence while demonstrating a commitment to platform reliability, cross-functional collaboration, and continuous