Java Senior Software Engineer —Vice President
Citigroup
- Location
- Irving Texas United States
- Work model
- On-Site
- Level
- Staff
- Posted
- Sep 18, 2026
Skills
About this role
We are seeking a highly skilled and results-driven Senior Software Engineer with deep expertise in observability, AIOps, and Java/Python application engineering to join our team and play a pivotal role in advancing the firm's Operating Model AI (OMAI) strategy. This role is central to our transition from traditional operational models to a predictive, efficient, and scalable AI-driven framework. The successful candidate will design, build, and ship production-grade applications, standardize telemetry across the estate using Open Telemetry (OTel), harness AIOps platforms to intelligently correlate and remediate operational events, and integrate AI/ML capabilities into observability pipelines delivering a unified, intelligent view of our systems. We are looking for an active, hands-on software engineer who thrives not only in designing, but in implementing the solution as well, takes pride in shipping production-grade code, and brings an energizing, collaborative presence to everything they do. Responsibilities : Design, develop, and deploy Java and Python applications across microservices and distributed architectures, with observability engineered in from the start Standardize telemetry (metrics, logs, and traces) across applications using OpenTelemetry (OTel) to build a consistent, AI-ready data foundation Advance AIOps capabilities by leveraging BigPanda to correlate operational alerts, suppress noise, and build automated remediation workflows Harness Google Cloud Observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace — for advanced cloud-native monitoring and performance analysis Integrate AI/ML into observability and automation pipelines to enable predictive failure detection and self-healing systems Analyze operational data to identify high-value automation opportunities that maximize reliability, efficiency, and engineering velocity Define and own SLIs, SLOs, and error budgets in partnership with reliability and product engineering teams Partner with AI/ML, platform, and architecture teams to co-innovate and deploy scalable automation solutions aligned to the OMAI strategy Act as Subject Matter Expert (SME) for observability and AIOps — providing hands-on technical guidance and mentorship across the team Participate in all SDLC phases — analysis, design, construction, testing, deployment, and maintenance Proactively troubleshoot and resolve complex application and infrastructure issues, driving long-term stability through collaborative root cause analysis Stay ahead of the curve on emerging practices in observability, AIOps, cloud engineering, and agentic AI Qualifications : 6+ years of hands-on software engineering experience Proven, hands-on Java development skills designing and delivering robust production applications using Spring Boot or equivalent frameworks, independently and at pace Practical Python experience for application development, automation, and tooling Demonstrated expertise in OpenTelemetry (OTel) — instrumenting Java and Python applications and standardizing metrics, logs, and distributed traces across services Exposure to tools like BigPanda for alert correlation, event management, and automated remediation workflows Proficiency with Google Cloud Observability — Cloud Monitoring, Cloud Logging, Cloud Trace, and related tooling — for cloud-native monitoring and performance analysis Strong understanding of SRE principles — including SLIs, SLOs, error budgets, and incident management Experience integrating AI/ML or agentic capabilities into operational workflows to enable predictive and self-healing systems (highly desirable) Solid grasp of RESTful API design and hands-on experience integrating monitoring and automation tools with backend services Strong version control and CI/CD experience — Git, pipeline tooling, and modern release practices Familiarity with containerization and orchestration technologies — Docker and Kubernetes — in cloud-native environments