Network Development Engineer, Capacity Restoration Team
Amazon
- Location
- IN, TS, Hyderabad
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- May 21, 2026
Skills
About this role
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we’re looking for talented people who want to help. You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion. As an Network Development Engineer on the Capacity Restoration Team, you are an experienced Builder. You will own the design and delivery of automation systems that transform capacity restoration from a manual, labor-intensive process into a scalable, self-service operation. You will architect solutions, lead technical projects end-to-end, mentor NDEs, and partner with service teams to integrate restoration automation into the broader tooling ecosystem. This role demands strong software engineering skills, deep networking knowledge, and the ability to drive results across organizational boundaries. Technical strategy is defined; component design is not. You are trusted with autonomy and are expected to make pragmatic trade-off decisions at product and component levels, identify and eliminate patterns affecting reliability and availability, and define and simplify team processes. Key job responsibilities Automation & System Design - Own design and delivery of automation systems that restore out-of-service capacity with minimal manual intervention. - Develop end-to-end link automation frameworks that transform manual troubleshooting into automated, system-guided restoration workflows. - Build and improve centralized device health validation services for border, backbone, and regional network layers. - Design self-service frameworks and next-step engines for automated troubleshooting and remediation. - Architect solutions that are scalable, secure, maintainable, and extensible across a growing multi-timezone team. Metrics & Monitoring - Create and maintain the metrics and monitoring infrastructure for CRT: capacity out-of-service dashboards, burndown tracking, TTR, restoration success rate, and SLA compliance. - Build systems to track and optimize bandwidth utilization trends, available capacity vs. projected peak, and redundancy coverage. - Design automation to improve in-team resolution percentage and reduce escalation to Operations and Engineering teams. Technical Leadership & Collaboration - Lead technical projects end-to-end; drive engineering best practices for development, testing, deployment, and operational excellence. - Mentor NDEs; actively participate in hiring and conducting technical assessments. - Partner with Engineering, Operations, Tooling, and Software teams to integrate restoration automation into the broader ecosystem. - Identify and proactively address architectural or process deficiencies affecting restoration performance, reliability, and scalability. Operations - Participate in on-call rotations and operational reviews for follow-the-sun coverage. - Lead complex incident response for capacity events; drive root-cause analysis and remediation for systemic failures. A day in the life - Review the capacity out-of-service dashboard; triage restoration priorities by customer impact and network health blocking status. - Design and implement scalable automation to systematically restore capacity at fleet level (e.g., end-to-end link