Systems Engineer, Region Services
Amazon
- Location
- AU, VIC, Melbourne
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- Nov 18, 2025
Skills
About this role
Applicants must be Australian citizens and hold or be eligible to obtain an Australian Government Security Clearance with the ability to successfully complete an Organisational Suitability Assessment. For more information regarding security clearances please visit (https://www.agsva.gov.au/). The AWS Region Services team combines AWS global cloud leadership with Australian security expertise to deliver highly secure, scalable environments for sensitive workloads. We’re creating innovative ways to use cloud computing, artificial intelligence, and machine learning while maintaining the highest standards of security and operational excellence. The Engineering organisation within Region Services is structured across core capability pillars: Compute & Machine Learning, Security Identity & Compliance, Storage & Databases, and a growing capability domain. Collectively these pillars encompass a team of varying technical skillsets, including Engineers; Technical Program Managers and Subject Matter Experts, organised into focused sub-teams. This is an opportunity to make a lasting impact on Australia’s digital future. You’ll work with AWS services, implement innovative solutions, and help customers succeed in their most important missions. We’re committed to helping our builders grow through continuous learning, mentoring, and collaboration with industry experts. Are you ready to build the future of secure cloud computing in Australia? Key job responsibilities - Define and/or refine hardware requirements, participate in the development and delivery of operability-related features such as system health monitoring, diagnostics, repair, and other self-healing automation - Develop or further existing application and system management tools and processes that reduce manual efforts and increase overall efficiency - Adapt and improve operations management systems and processes to accommodate rapid and increasing growth in systems and traffic - Participate in the design and execution of production acceptance tests and new hardware evaluations - Monitor the health of the fleet, automating system health, maintenance tasks, and reporting systems as needed - Participate in “on-call” rotations to resolve incidents occurring out-of-hours. - Must be an Australian Citizen and hold or be able to attain an Australian Government Security Vetting Agency clearance (see https://www.agsva.gov.au) A day in the life Your morning begins with a fleet health review — scanning dashboards you've built, checking automated alerts, and confirming that overnight self-healing routines performed as expected. A quick stand-up with your sub-team surfaces a capacity threshold approaching in one of the compute clusters. You pull up the growth projections and begin sketching an automation enhancement that will handle the scaling gracefully. Mid-morning, you're deep in code. Today it's refining a diagnostic tool that identifies degraded hardware components before they impact workloads. You test against real fleet telemetry, iterate on detection thresholds, and push a change that will save hours of manual investigation across the team. After lunch, you join a hardware evaluation session. A new server platform is being assessed for production readiness, and your role is to design acceptance test criteria — thermal performance under load, firmware compatibility, integration with existing monitoring frameworks. Your input directly determines whether this hardware earns its place in the fleet. Late afternoon brings a collaborative design review. A colleague proposes a new approach to automated repair workflows, and you offer feedback drawn from patterns you've observed in incident data. The discussion is rigorous, respectful, and energising — this is a team that sharpens each other. Before you wrap up, you check in on a long-running automation project — a self-healing pipeline that's reduced manual intervention by 40% since you deployed it last quarter. You note