Minimum qualifications: Bachelor’s degree or equivalent practical experience. 4 years of experience in designing, training, and evaluating deep generative models (e.g., Diffusion Models, Transformers, GANs) for media synthesis and manipulation (image, video, audio). 4 years of experience with experimentation, dataset curation, eval systems, and building ML pipelines. 4 years of experience with machine learning, computer vision, and deep learning. Preferred qualifications: Master's or PhD degree in Computer Science, Machine Learning, Statistics, or a related field, or equivalent practical experience. Experience with multimodal learning, integrating video, audio, and text. Experience with video generation and editing models. Knowledge of techniques for model optimization and efficient/low-latency inference suitable for real-time applications. Familiarity with real-time systems or media streaming technologies. Experience with Python and deep learning frameworks such as JAX, TensorFlow, or PyTorch.