yoinka

AMHS Site Reliability Engineer

Samsung

1530 FM 973 Taylor, TX, USAMid$90k – $115k/yr
Sign in to applyVerified 3h ago
Location
1530 FM 973 Taylor, TX, USA
Work model
On-Site
Level
Mid
Salary
$90k – $115k/yr
Posted
23h ago

Skills

DockerGrafanaJavaKubernetesLinuxPrometheusPythonRabbitMQRedis

About this role

About Samsung Austin Semiconductor Samsung is a world leader in advanced semiconductor technology, founded on the belief that the pursuit of excellence creates a better world. At Samsung Austin Semiconductor, we are Innovating Today to Power the Devices of Tomorrow.  Come innovate with us!

Position

Summary The AMHS Site Reliability Engineer (SRE) is a critical member of the software operations organization, responsible for guaranteeing the uncompromising 24x7 reliability and scalability of the software controlling Samsung's automated material transport ecosystem. In this role, you will ensure the seamless performance of the Material Control System (MCS) and OHT Control System (OCS), tackling complex server-side bottlenecks and network communications (IPC, RPC) to prevent fab operational disruptions. Your primary objective is to transform reactive incident response into proactive reliability. Beyond immediate operational support, you will spearhead the modernization of AMHS infrastructure by leveraging Python, Java, and modern SRE practices. You will design scalable data pipelines, implement comprehensive observability frameworks (Prometheus, Grafana), and manage the orchestration of services via Docker and Kubernetes. Working at the intersection of software engineering and manufacturing execution systems (MES), you will build sophisticated automation tooling to reduce manual toil and optimize critical middleware (e.g., RabbitMQ, Redis). Ultimately, your data-driven optimizations will ensure the AMHS control software scales seamlessly, directly driving the throughput and productivity of our high-stakes semiconductor fabrication site. Role and Responsibilities Here’s What You’ll Be Responsible For: AMHS Control Software Reliability:  Ensure the high availability, performance, and scalability of the AMHS controlling software suite (e.g., Vehicle Control Systems, fleet routing logic, and material tracking applications) rather than the physical hardware. Back-end Data Engineering & Monitoring:  Design, build, and maintain scalable back-end data pipelines to process high-throughput logs and events generated by fab control systems. Develop proactive monitoring and observability solutions. Incident Response & Software Troubleshooting:  Act as the primary escalation point for AMHS software anomalies. Troubleshoot complex software bottlenecks, network communications (IPC, RPC), and server-side issues to minimize fab downtime. Operations Automation:  Reduce manual toil by developing robust automation scripts and tools for deployment, configuration management, and system recovery.

Skills and Qualifications

Here's what you'll need: Bachelor’s degree in computer science, Software Engineering, or a related technical field. Minimum of 3+ years of work experience in semiconductor-related industries.  Knowledge of semiconductor manufacturing processes, OR direct experience working as a Software Engineer, SRE, or DevOps Engineer within a semiconductor fabrication (Fab) environment. Proven hands-on experience in Software Engineering, Site Reliability Engineering (SRE), or DevOps roles. Strong background in  Back-end Data Engineering , including experience with ETL processes, relational/NoSQL databases, and managing large-scale data pipelines. Observability (o11y) & Middleware- Proficiency with monitoring, alerting, and visualization tools such as  Prometheus and Grafana , as well as message brokers and in-memory data stores like  RabbitMQ  and  Redis Capability to document technical details in depth with expectation of strong English language skills. Strong proficiency in navigating, troubleshooting, and administering Linux/Unix environments. Experience with equipment control protocols (e.g., SECS/GEM) or manufacturing execution systems (MES) is a strong plus. The current base salary range for this role is between $90,000.00 - $115,000.00. Individual base pay rates will depend on factors including duties, work location,

AMHS Site Reliability Engineer at Samsung, 1530 FM 973 Taylor, TX, USA | Yoinka