DE-RCE-AIProject-PythonAI-Manager-GDSN02
EY
- Location
- Pune, MH, IN, 411014
- Work model
- On-Site
- Level
- Senior
Skills
About this role
At EY, you’ll have the chance to build a career as unique as you are, with the global scale, support, inclusive culture and technology to become the best version of you. And we’re counting on your unique voice and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all.
Manager - Python PySpark | Python | SQL | Hadoop | Impala Experience
7+ years
Domain preference BFSI / Banking / Capital Markets / FSO Employment Full-time EY GDS - Delivery Excellence Manager Position In today's fast-paced technology landscape, cost-effective delivery and engineering excellence are critical to building successful global businesses. Our teams apply modern data engineering practices, automation, and scalable platforms to deliver tailored solutions for clients across industries. As part of our global team, you will contribute to GDS projects as a data engineering leader. You will solve complex data challenges, guide engineering teams, and help create next-generation data processing and analytics platforms. The Opportunity We are offering a challenging Manager-level role for a hands-on data engineering professional. You will design and deliver scalable batch and distributed data solutions, work in an international environment, and collaborate with business and technology stakeholders to create reliable, high-performing data products. Role Overview We are looking for a highly skilled Manager to lead the design, development, and delivery of enterprise-grade data platforms. The ideal candidate will have strong hands-on expertise in PySpark, Python, SQL, Hadoop, and Impala, along with proven leadership, stakeholder management, and delivery capabilities. Exposure to Flask and basic REST API development will be an advantage. Key Responsibilities
Lead the end-to-end design, development, testing, deployment, and support of scalable data engineering solutions. Build and optimize distributed data processing pipelines using PySpark and the Hadoop ecosystem. Develop reusable, maintainable, and well-tested data processing components in Python. Design complex SQL transformations and tune queries for performance across large data volumes. Use Impala for interactive analytics, data validation, troubleshooting, and performance-focused querying. Drive data ingestion, transformation, quality, lineage, reconciliation, and monitoring practices. Partner with product owners, architects, analysts, and client stakeholders to translate requirements into technical solutions. Provide technical leadership, conduct code and design reviews, and mentor a team of data engineers. Ensure adherence to engineering standards covering security, privacy, coding quality, version control, and release management. Drive Agile delivery, estimation, risk management, dependency tracking, and continuous improvement. Participate in client discussions, solution design, proposal support, and technical presentations.
Primary Skills - Must Have
7+ years of software or data engineering experience, including experience leading technical teams or workstreams. Strong hands-on experience with PySpark for large-scale distributed data processing and transformation. Domain - FSO, banking, capital markets, or other regulated industries. Advanced Python programming skills, including modular design, exception handling, logging, testing, and performance optimization. Strong SQL skills, including complex joins, window functions, aggregations, query optimization, and data validation. Solid experience with Hadoop components and distributed storage and processing concepts. Hands-on experience with Impala, including query development, troubleshooting, and performance tuning. Strong understanding of data modeling, data quality, metadata, partitioning, file formats, and performance optimization. Effective communication,