Lead DevOps/AIOps Engineer
Blend360
- Location
- Columbia, MD, United States
- Employment
- Full Time
- Work model
- Remote
- Level
- Senior
- Salary
- $130k – $155k/yr
- Posted
- 1h ago
Skills
About this role
Compensation
USD 130000 - USD 155000 - yearly Company Description Blend360 is a premier data, AI, and marketing consulting firm that partners with the world's most ambitious organizations to turn complex challenges into competitive advantage. We sit at the intersection of deep analytical rigor and pragmatic business execution—helping Fortune 1000 companies and Private Equity-backed businesses unlock transformational value through data, technology, and human expertise.
Job Description
Blend360 is looking for a Lead DevOps / MLOps Engineer to help architect, automate, and operationalize modern cloud-based data and AI platforms for enterprise clients. This role sits at the intersection of cloud infrastructure, data engineering, machine learning, and software delivery, with a strong emphasis on Google Cloud Platform (GCP). We’re looking for someone who can move comfortably between architecture and hands-on engineering—designing scalable solutions, establishing DevOps and MLOps best practices, and helping engineering teams reliably move data and AI workloads into production.
What you'll do
Lead the design and implementation of cloud-native DevOps and MLOps architectures on GCP. Build and optimize CI/CD pipelines for data, ML, and application workloads. Develop infrastructure-as-code using tools such as Terraform and establish repeatable deployment patterns. Architect and operationalize data platforms leveraging BigQuery, Cloud Storage, Dataflow, Pub/Sub, Dataproc, and Cloud Composer. Build MLOps capabilities supporting the full ML lifecycle, including model development, deployment, monitoring, versioning, and retraining. Establish observability across data and ML platforms, including logging, monitoring, alerting, pipeline health, data quality, and model performance. Implement secure, scalable cloud infrastructure using GCP IAM, networking, secrets management, and appropriate security controls. Partner with Data Engineers, ML Engineers, Architects, and client stakeholders to translate business requirements into production-ready technical solutions. Establish engineering standards around deployment automation, testing, environment management, reliability, and operational excellence. Troubleshoot complex production issues and drive root-cause analysis and long-term remediation. Mentor engineers and serve as a technical leader across DevOps, cloud, data, and MLOps initiatives. Evaluate emerging GCP and AI technologies and determine where they can create meaningful business or engineering value.
Qualifications
7+ years of experience in DevOps, cloud engineering, platform engineering, MLOps, or a related discipline. Strong hands-on experience with Google Cloud Platform, particularly BigQuery and cloud-native data services. Experience designing and implementing end-to-end data platforms on GCP. Strong understanding of BigQuery architecture, performance optimization, data ingestion, partitioning, clustering, and data security. Experience with CI/CD, Git, automated testing, containerization, and Kubernetes/GKE. Strong Infrastructure-as-Code experience, preferably Terraform. Experience with Vertex AI and/or production ML platforms, including model deployment and monitoring. Experience with orchestration and data processing technologies such as Cloud Composer/Airflow, Dataflow, Dataproc/Spark, and Pub/Sub. Strong understanding of observability, reliability engineering, monitoring, logging, and alerting. Proficiency with scripting/programming languages such as Python and/or Bash. Strong understanding of cloud security, IAM, networking, secrets management, and enterprise governance. Ability to operate at both the architectural and hands-on engineering levels. Excellent communication skills and the ability to work effectively with both technical teams and senior client stakeholders.
Nice to Have
Experience with Vertex AI, MLflow, Kubeflow, or other MLOps platforms. Experience implementing GenAI/LLM