L2 Production Support Genesis
Citigroup
- Location
- DLF CYBERCITY 12B
- Work model
- On-Site
- Level
- Mid
- Posted
- Sep 2, 2026
Skills
About this role
Key Responsibilities
1. Big Data & EAP (Enterprise Application/Analytics Platform) Support Support Distributed Environments: Provide Level 2 (L2) and Level 3 (L3) support for applications hosted on Big Data platforms and Citi's Enterprise Application/Analytics Platform (EAP). Troubleshoot Data Pipelines: Diagnose and resolve failures in complex data ingestion and processing pipelines, including distributed processing frameworks (e.g., Apache Spark, Hadoop MapReduce). Cluster & Resource Monitoring: Monitor cluster resource utilization (using YARN, Cloudera Manager, or similar tools) to identify and resolve memory bottlenecks, queue congestion, and job failures (e.g., Spark Out-Of-Memory errors). Data Querying & Validation: Query and validate large-scale datasets stored in distributed data warehouses and file systems (e.g., HDFS, Hive, Impala, or HBase). Message Queue Management: Monitor and troubleshoot real-time streaming and messaging platforms (e.g., Apache Kafka), managing consumer groups, offsets, and partition lags. 2. Batch Management & Job Scheduling (Autosys) Monitor and manage batch execution: Oversee the execution of critical daily, weekly, and monthly batch processing cycles scheduled via Autosys. Troubleshoot batch failures: Rapidly diagnose and resolve Autosys job failures, analyzing log files, identifying dependency issues, and performing necessary job overrides, force-starts, or hold/release actions to minimize business impact. Optimize job flows: Collaborate with development and engineering teams to define, configure, and optimize Autosys job definitions using JIL (Job Information Language). 3. Automation & Process Enhancement (Toil Reduction) Identify and eliminate manual bottlenecks: Actively analyze daily support activities to identify repetitive, manual tasks ("toil") and design automated solutions to eliminate them. Develop automation scripts: Write, test, and deploy robust scripts (using Python, Bash, or PowerShell) to automate routine operations, such as daily health checks, application restarts, log archiving, and data reconciliation. Drive process improvements: Evaluate existing support workflows, runbooks, and escalation paths, implementing enhancements to streamline operations and reduce Mean Time to Repair (MTTR). 4. Incident Management & Production Recovery Own and drive the end-to-end resolution of L2/L3 production incidents, ensuring strict adherence to corporate Service Level Agreements (SLAs) and Service Level Objectives (SLOs). Lead technical triage during Major Incidents (MIM) and high-severity outages. Coordinate effectively with cross-functional global teams (Infrastructure, Database, Networks, Development, and Business Operations) to restore services rapidly. Act as the primary technical escalation point during incidents, translating complex technical issues into clear, concise, and business-friendly updates for senior leadership and stakeholders. Ensure accurate and timely logging, categorization, and tracking of incidents within ServiceNow. 5. Problem Management & Root Cause Analysis (RCA) Lead proactive Problem Management initiatives by analyzing incident trends, identifying systemic patterns, and pinpointing recurring failure points. Conduct deep-dive technical investigations—including log analysis, database queries, and infrastructure health checks—to perform comprehensive Root Cause Analysis (RCA). Author high-quality Post-Incident Reviews (PIRs) and RCA documents, detailing the timeline, root cause, impact, and preventative actions. Collaborate closely with Development and Engineering teams to prioritize, track, and implement permanent bug fixes, structural workarounds, and long-term remediations. 6. Hands-on Unix/Linux & Application Troubleshooting Perform deep-dive technical troubleshooting directly within Unix/Linux production environments (analyzing system resources, CPU/memory bottlenecks, process states, and network connectivity). Conduct advanced log analysis