What We Need to See: Bachelor’s, Master’s, or PhD in Computer Science, Electrical Engineering, or equivalent experience. 8+ years of experience building large-scale distributed systems or performance-critical software. Deep understanding of deep learning systems, GPU acceleration, and AI model execution flows and/or high performance networking. Solid software engineering skills in C++ and/or Python, preferably demonstrate strong familiarity with CUDA or similar platforms. Strong system-level thinking across memory, networking, scheduling, and compute orchestration. Excellent communication skills and ability to collaborate across diverse technical domains. Ways to Stand Out from the Crowd: Experience working on LLM - training or inference pipelines, transformer model optimization, or model-parallel deployments. Demonstrated success in profiling and optimizing performance bottlenecks across the LLM training or inference stack. AI Accelerators and distributed communication patterns, congestion control and/or load balancing. Proven optimization process for complex systems, deployed at scale to make impact. Passion for solving tough technical problems and finding high-impact solutions. NVIDIA is widely considered one of the most desirable places to work in tech – we are passionate about what we do and are committed to fostering a culture of excellence, innovation, and collaboration. If you’re excited to help define how the world runs AI at scale, this role is for you.