Infrastructure and Data Engineer
Millennium Management
- Location
- New York, New York, United States of America
- Work model
- On-Site
- Level
- Mid
- Posted
- Aug 19, 2026
Skills
About this role
Infrastructure and Data Engineer About Millennium Millennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution, innovation and focus, Millennium’s mission is to deliver results for our investors. Our people are empowered with both independence and support: the autonomy to pursue ideas with conviction and the backing of a global network committed to collaboration, disciplined risk management and continuous learning. With opportunities to deepen expertise and accelerate development, talent at Millennium is equipped to adapt, evolve and build lasting impact over time. Discover how transformative growth accelerates impact. Meet the Team Technology is core to the health and growth of Millennium’s business. The firm’s active, multi-manager business model demands flexible, scalable technology and advanced proprietary systems, including the development of next-generation analytical and trading capabilities. The Data Science team builds and supports data infrastructure, applications, and AI and machine learning systems used by people and platforms across the business. What You’ll Do Build and maintain Python and SQL data pipelines that support feature stores, embeddings, and model training and inference workflows. Develop infrastructure for extracting, transforming, and loading data from Snowflake, SQL Server, streaming sources, and other systems across on-premises and cloud environments. Design, deploy, and maintain reproducible, scalable environments using infrastructure-as-code tools such as Terraform and CloudFormation. Manage region-redundant Airflow orchestration, Docker and Kubernetes services, and the NGINX, Gunicorn, and Django web-serving stack. Strengthen user-facing applications and APIs by improving authorization, load balancing, containerization, and CI/CD deployment pipelines. Build production infrastructure for LLM-based applications, including retrieval-augmented generation pipelines, vector databases, embeddings, and third-party or self-hosted model APIs. Implement monitoring, logging, observability, and analytics across data pipelines, application infrastructure, and model-serving endpoints to provide insight into system health, usage, and performance. Drive improvements across the technology stack by automating manual processes, strengthening data governance and access controls, improving scalability, and reducing infrastructure and inference costs. What You Bring Two or more years of professional experience with a master’s degree, or three or more years with a bachelor’s degree, in Computer Science, Statistics, Informatics, Information Systems, or another quantitative field. Advanced Python skills, including experience with Pandas, NumPy, and SciPy, as well as familiarity with machine learning and AI libraries such as PyTorch, scikit-learn, or orchestration frameworks such as LangChain. Strong SQL and data engineering experience, including relational databases such as Microsoft SQL Server or PostgreSQL, modern cloud data warehouses such as Snowflake, and pipeline orchestration with Airflow. Hands-on experience with AWS services, Docker, infrastructure-as-code tools such as Terraform or CloudFormation, and preferably Kubernetes. Practical knowledge of production AI and MLOps systems, including vector databases, embedding-based retrieval, LLM APIs, prompt engineering, context management, RAG architecture, model versioning, feature stores, and model monitoring. Strong computer science and infrastructure fundamentals, including distributed systems, data structures, threading, memory management, Unix/Linux environments, CI/CD, security, and observability. Excellent written and verbal communication skills, with the ability to work independently and collaboratively while managing multiple priorities in a fast-paced environment. A proactive, detail-oriented approach to problem-solving, with ownership of outcomes and the ability to deliver quality code as technologies