Principal Site Reliability Engineer
Fidelity Investments
- Location
- Durham, NC
- Work model
- On-Site
- Level
- Principal
- Posted
- Sep 9, 2026
Skills
About this role
Job Description
Note: Fidelity will not provide immigration sponsorship for this position.
Position
Description : Deploys and supports distributed, multi-tiered systems at scale while ensuring high availability and fault tolerance across multiple environments. Builds and operates resilient platforms in Amazon Web Services (AWS) using Elastic Compute Cloud (EC2), Simple Storage Service (S3), and Auto Scaling Groups for dynamic resource management. Designs, develops, and executes performance tests using Java-based frameworks, Apache JMeter, k6, and Rush-hour to validate system behavior under day-to-day traffic patterns. Defines and implements observability practices to monitor system health, latency, and error rates through metrics, logs, and distributed tracing using Datadog, Grafana, Splunk, and the Elasticsearch, Logstash, and Kibana (ELK) stack. Automates operational workflows with Python and Shell scripting to enhance efficiency and reduce manual tasks. Supports consistent build, deployment, and orchestration processes using cloud computing and DevOps technologies -- Continuous Integration and Continuous Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets, and implementing proactive monitoring and incident response strategies. Builds and refines methodologies for performance, load, stress, and chaos testing and develops analytics and reports aligned with business needs to improve system resilience and optimization. Primary Responsibilities : Defines and leads enterprise-level reliability strategies. Architects resilient systems and infrastructure. Creates and publishes performance test results report with recommendations on quality improvement. Maintains scalability and resiliency of complex environment. Implements advanced observability practices and techniques at scale. Manages and interprets large datasets using query languages and visualization tools. Advises senior leadership on reliability engineering best practices. Mentors junior engineers. Performs independent and complex technical and functional analysis for multiple divisional initiatives. Develops innovative solutions to improve system availability, scalability, and performance. Designs, implements, and maintains performance test frameworks. Education and Experience : Bachelor’s degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment. Or, alternatively, Master’s degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and three (3) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment. Skills and Knowledge : Candidate must also possess: Demonstrated Expertise (“DE”) performing software performance benchmarking and engineering for online financial web applications, Application Programming Interfaces (APIs), and mobile transactions according to DevOps practices, using performance benchmarking tools Rushhour, Locust, K6, and JMeter; and configuring CI/CD and test automation, using Jenkins, Sonar, Ant, Maven, Artifactory, and Terraform in AWS. DE solutioning, designing, architecting, and building scalable and resilient enterprise-grade software platforms using cloud-based architecture and AWS services (EC2, Elastic Container Service (ECS), Lambda, Elastic MapReduce (EMR), and CloudFormation); developing microservices on Elastic Kubernetes