Senior Data Scientist
FourKites
- Location
- Chennai or Remote, India
- Work model
- Remote
- Level
- Senior
- H-1B history
- 1 approvals (FY2023)
- Posted
- 3h ago
Skills
About this role
At FourKites we have the opportunity to tackle complex challenges with real-world impacts. Whether it's medical supplies from Cardinal Health or groceries for Walmart, the FourKites platform helps customers operate global supply chains that are efficient, agile and sustainable.
Join a team of curious problem solvers that celebrates differences, leads with empathy and values inclusivity.
As a Senior Data Scientist, you will build and own machine learning models that power core prediction problems across the FourKites platform — including ETA/ATA forecasting and message-based status extraction. You will work end-to-end, from data pipeline to production deployment and monitoring, turning noisy real-world logistics data into models that run at scale and directly move the needle on customer outcomes. You will work closely with product, engineering, and operations teams, hands-on building and shipping models yourself while also guiding the technical direction of other data scientists on the team.
What you'll be doing
• Design, build, and productionize ML models for problems like ETA/ATA prediction, using regression, classification, and time-series forecasting techniques
• Develop NLP/LLM-based extraction pipelines for message-based ETA and status updates (text extraction, entity recognition)
• Own models end-to-end: data pipeline → training → deployment → monitoring → retraining
• Work with noisy, real-world logistics and supply chain data (GPS pings, check calls, carrier data) rather than clean, pre-processed datasets
• Diagnose gaps between offline evaluation performance and live production accuracy, and drive fixes
• Build and maintain automated training/retraining pipelines using orchestration tools such as Airflow
• Set up and maintain model monitoring and observability (e.g., Grafana) to catch drift and degradation proactively
• Replace manual or rule-based processes with ML-driven automation (e.g., automating manual check calls)
• Translate model performance improvements into business impact — operational savings, efficiency gains, and deal-relevant outcomes
• Mentor and guide other data scientists/engineers on technical approach and best practices
• Make build-vs-buy and architecture tradeoff decisions independently
About the team
Our product and engineering teams are dedicated to providing the industry's best-in-class end-to-end supply chain visibility platform. We are committed to building a high-performing, ML-driven team that turns supply chain data into automated, proactive action — and we want you to help lead that effort.
Who you are
• Strong ML fundamentals across regression, classification, and time-series forecasting
• NLP experience — text extraction, entity recognition, or LLM-based extraction
• Production ML experience — you've shipped models serving real traffic, not just built POCs or notebooks
• Strong Python and SQL skills — pandas, scikit-learn, and comfort querying large datasets (Redshift/Snowflake a plus)
• Experience with cloud and data infrastructure — AWS (S3, EC2), and orchestration tools like Airflow for training/retraining pipelines
• Experience setting up or working with model monitoring and observability tooling (Grafana or similar)
• Comfortable working with noisy, real-world data rather than clean, curated datasets
• Experience diagnosing and closing the gap between offline evaluation results and live production performance
• A track record of replacing manual/rule-based processes with ML solutions
• Ability to translate model output into business value and communicate that impact to non-technical stakeholders
• Experience collaborating cross-functionally with product, engineering, and