yoinka

ML Researcher - Image / Video Diffusion

Krea AI

San FranciscoFull TimeMidVisa sponsorship
Sign in to applyVerified 1h ago
Location
San Francisco
Employment
Full Time
Work model
On-Site
Level
Mid
Sponsorship
Sponsors visa
Posted
2h ago

Skills

LLMPyTorch

About this role

About Krea At Krea, we are building next-generation AI creative tools. We're dedicated to making AI intuitive and controllable for creatives - our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats - text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium. We recently took this a step forward with the launch of Krea 2 , our first foundation model, built completely from scratch for aesthetic diversity and stylistic control. We've raised over $83M and are backed by world-class investors such as a16z, Bain Capital, and Abstract. We work full-time and in-person at our waterfront office in San Francisco. We care about creativity: our team includes musicians, designers, visual artists, and engineers.   We're looking for an experienced Researcher with engineering skills who can work on large-scale image and video models training experiments, with experience training image models at scale.

Our culture

We work full-time and in-person at our North Beach office in San Francisco. We believe that demonstrated interest in the creative space is key: our team includes musicians, designers, visual artists and more. Fast iteration and execution speed. Bias towards action, agency, and independence.

What you'll do

Train diffusion models for image and video generation on large GPU clusters. Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication. Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP. Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design. Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues. Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models.

What we're looking for

Proven track record in working with image or video models at scale (publications or open-source contributions a plus). Strong proficiency in PyTorch and understanding of its inner workings. Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs. Experience in profiling and debugging large distributed training. Being comfortable with analyzing traces to identify bottlenecks and look for improvements. Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8. Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning. Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research. Being comfortable working in a goal-oriented research environment. Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data. Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items. Good research taste — bias towards simplicity and methods that scale well with compute, data, and minimal human supervision. Ability to iterate rapidly, and propose creative research directions. Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality.

What we offer

Team : Work alongside a world-class team building the future of AI creative tooling Impact : Significant scope and company-wide impact Competitive compensation :

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka