Content Quality and Evaluation Project Intern (AI Data Service Operations) - 2027 Start
TikTok
- Location
- Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur, Malaysia
- Employment
- Internship
- Work model
- On-Site
- Level
- Intern
- H-1B history
- 148 approvals (FY2023)
Skills
About this role
About the Team Generative AI and large language models are reshaping the core capabilities of global content platforms — and high-quality training and evaluation data is one of the key factors that determines a model's ceiling.
We are the AI Data Service & Operations team behind TikTok's international products, responsible for producing multilingual, multimodal data assets across both safety and non-safety domains. The standards we set and the data we deliver power two critical fronts: TikTok's global content ecosystem governance strategy on one side, and the training and evaluation pipelines for AI / large language models on the other. In short, we produce the fuel that helps models truly understand content, communities, and users around the world.
As an Eco Content Quality and Evaluation intern, you'll sit at the intersection of model performance and ecosystem governance. You'll act as the gatekeeper for AI training and evaluation data — defining quality standards, unifying judgment criteria, and safeguarding data trustworthiness and consistency, so that every data point holds up under the scrutiny of both models and the business.
As a Project Intern, you will contribute to impactful short-term projects and gain hands-on experience in a fast-paced, professional environment. This internship offers the opportunity to develop practical skills, apply your knowledge to real-world challenges, and explore your career interests. Applications are reviewed on a rolling basis, so we encourage you to apply early.
Key Responsibilities 1. Help design and iterate on data quality standards, judgment rules, and acceptance criteria for LLM training and evaluation scenarios; 2. Review and adjudicate data outputs against quality standards, identify issues, and close the loop on corrections to ensure accuracy and consistency; 3. Dig into complex, ambiguous, and edge cases; drive discussion and turn conclusions into reusable judgment rules and knowledge assets; 4. Track quality metrics, root-cause issues, and drive improvements in production processes and execution; 5. Partner cross-functionally with production, product, algorithm, and policy teams to align on quality standards and connect the "standard → production → acceptance → feedback" loop.