Full-Stack Engineer, ML Tooling and Infrastructure
Mecka AI
- Location
- Toronto
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- 5h ago
Skills
About this role
About Mecka AI Mecka AI is building the data infrastructure layer for robotics and embodied AI. We design and operate global systems for data capture, data labeling, and hardware-enabled workflows used by leading AI labs and robotics companies to train and validate humanoid and embodied AI systems. We work closely with frontier robotics teams to bridge real-world data, simulation, learning-based systems, and deployed hardware.
The Role
This role exists as a high-leverage force multiplier for our core Computer Vision Machine Learning (CVML) team. You are not building the foundational models for our robotics systems; rather, you own the critical infrastructure, internal tooling, and targeted ML services that enable the CVML team to move fast. You will own the internal annotation platform used by our labeling team and develop lightweight, high-reliability ML services like our Personal Identifiable Information (PII) blurring pipeline. This is a true hybrid engineering role demanding competence in both deploying practical ML models and writing robust, user-facing full-stack applications.
Responsibilities
Internal Tooling & Infrastructure Own the Annotation Platform (CVA): Build and maintain the web application our labeling team uses daily. You will ship new task types, review workflows, keyboard-driven UI features, and progress dashboards. Data Pipelines: Write the glue code to efficiently move video frames, metadata, and JSON payloads between cloud storage, databases, and the client application. Quality Measurement: Build tools to track inter-annotator agreement, audit sampling, and dataset health. Targeted ML Services PII Blur as a Service: Own the pipeline that detects and blurs faces, screens, and license plates in delivered footage. This is a strict, customer-facing system that must perfectly balance privacy compliance with data preservation. Model-Assisted Labeling: Deploy and optimize lightweight detection and segmentation models (e.g., bounding box assists) so annotators correct rather than create from scratch. Inference Optimization: Take off-the-shelf or provided PyTorch models, optimize them (e.g., TensorRT, ONNX), and wrap them in fast, concurrent APIs to serve the tooling UI.
Who You Are
Required Skills (The Hard Bar) Heavy Frontend Engineering: You can confidently build and ship interactive web applications (React, Vue, or modern JS/TS). You understand how to handle complex state, canvas-based rendering, or video playback in the browser without tanking performance. Production-Grade Python: You know how to build fast, concurrent APIs (FastAPI, gRPC) that don't choke on high-volume media requests. Practical Applied ML: You have experience fine-tuning and deploying detection, segmentation, or tracking models on messy, real-world data. You know how to take a model out of a Jupyter notebook and make it run reliably in a pipeline. Strong Signals You have built or heavily customized an annotation UI (handling bounding boxes, polygons, keypoints) and understand the pain points of human-in-the-loop workflows. Deep familiarity with video processing pipelines (FFmpeg, frame extraction, codec handling, latency optimization). Familiarity with active learning workflows or mining hard negatives to iteratively improve datasets. You care deeply about UI/UX and measure your success by the workflow speed of the end-user. Who Will Not Enjoy This Role Core ML Architects: If your primary goal is designing novel neural network architectures or working on frontier foundation models, this is the wrong fit. Your ML work here uses established architectures to solve practical pipeline problems (like PII detection). Backend-Only Engineers: If you dislike writing TypeScript, debugging CSS, or thinking about UI workflows, you will struggle. You own the frontend annotation product just as much as the ML services. Data-Allergic Engineers: If you expect a perfectly clean, balanced dataset handed to you, this isn't the role. You are building the