Currently enrolled in a graduate program (M.Sc. or Ph.D.) in Computer Science, Electrical Engineering, or a related field Publications at top-tier venues (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, EMNLP or similar) Strong programming skills in Python and experience with deep learning frameworks (e.g., PyTorch) Solid foundation in computer vision, natural language processing, or multimodal learning Demonstrated expertise working with Vision-Language Models (VLMs) and/or Large Language Models (LLMs) Experience with human pose estimation, motion modeling, or related body-tracking tasks Familiarity with video understanding tasks and temporal modeling Familiarity with multimodal learning and benchmarks that combine language with visual or spatial data Experience with prompt engineering and optimization techniques