yoinka

Software Engineer — Distributed LLM Inference Systems

Intel

PRC, ShanghaiMidH-1B sponsor company
Sign in to applyVerified 55m ago
Location
PRC, Shanghai
Work model
On-Site
Level
Mid
H-1B history
1,112 approvals (FY2023)
Posted
Aug 14, 2026

Skills

Deep LearningLLMMachine LearningPyTorchPython

About this role

Job Details

Job Description: The Role and Impact: As a Software Engineer on Intel’s Artificial Intelligence Frameworks team, you will contribute to designing, developing, and optimizing distributed inference systems for large language models. Your day-to-day work will involve implementing distributed inference algorithms, optimizing model execution and communication, transforming neural network models, and developing software components that improve inference performance across diverse hardware architectures. You may work on areas such as disaggregated serving, request scheduling, KV cache management, parallel execution, and efficient communication between inference components. By collaborating with researchers and engineers, you will play a key role in advancing Intel's AI capabilities and ensuring industry-leading solutions. Business Group: Intel's Artificial Intelligence Frameworks team is dedicated to empowering transformative AI solutions by developing and optimizing software frameworks for machine learning and deep learning. This group works on enhancing the performance of AI applications across diverse computing hardware backends while contributing to open-source communities. As part of Intel, this team supports the mission to drive technological innovation and deliver impactful AI advancements globally.

Key Responsibilities

Design, develop, and optimize distributed LLM inference systems and related AI framework components. Implement distributed algorithms, including model/data parallel frameworks and asynchronous communication for deep learning. Develop and optimize components such as request schedulers, model workers, communication layers, and KV cache management mechanisms. Profile distributed inference workloads to identify computation, communication, memory, and scheduling bottlenecks. Collaborate with component teams to improve end-to-end latency, throughput, scalability, and resource utilization. Contribute high-quality code, tests, and documentation to internal and open-source projects while following industry engineering standards.

Qualifications

Minimum Qualifications – Master’s degree in computer science, Artificial Intelligence, Software Engineering, or a related field, with 0-1 years of hands-on experience demonstrated through internships, academic projects, coursework, or training. Proficiency in Python and modern C++ programming. Foundational knowledge of deep learning and AI frameworks, such as PyTorch. Experience debugging and optimizing software for performance. Basic understanding of machine learning algorithms and techniques. Strong problem-solving skills and the ability to learn unfamiliar systems quickly.

Preferred Qualifications

Experience or project exposure related to distributed LLM inference and serving. Experience in contributing to open-source projects or collaborating within open-source ecosystems. Understanding of LLM inference concepts such as prefill and decode, KV cache management, continuous batching, parallelism strategies, and disaggregated serving. Familiarity with inference engines or serving frameworks such as vLLM, SGLang, TensorRT-LLM, or similar technologies. Knowledge of large language models and inference optimization techniques. Knowledge of AI Agent architecture and execution workflows, including tool calling, planning, memory, context management, and multi-agent coordination. Effective communication skills, including fluency in written and spoken English. Take the opportunity to be part of Intel's journey in redefining AI frameworks and enabling groundbreaking innovations. Take the opportunity to be part of Intel's journey in redefining AI frameworks and enabling groundbreaking innovations. Your contributions will shape the future of AI software and its real-world impact. Apply now and be a part of advancing transformative AI capabilities.            Job Type: College Grad Shift: Shift 1 (China) Primary Location: PRC, Shanghai

Listing verified 55m ago. Applications go through the company's official careers site.

← Back to Yoinka

Software Engineer — Distributed LLM Inference Systems at Intel, PRC, Shanghai | Yoinka