What We Need to See: Bachelor’s, Master’s, or PhD in Computer Science, Electrical Engineering, or equivalent experience. 5+ years of experience building large-scale distributed systems or performance-critical software. Deep understanding of deep learning systems, GPU acceleration, and AI model execution flows. Solid software engineering skills in C++ and/or Python, with strong familiarity with CUDA or similar platforms. Strong system-level thinking across memory, networking, scheduling, and compute orchestration. Excellent communication skills and ability to collaborate across diverse technical domains. Ways to Stand Out from the Crowd: Experience working on LLM inference pipelines, transformer model optimization, or model-parallel deployments. Demonstrated success in profiling and optimizing performance bottlenecks across the LLM training or inference stack. Familiarity with data center-scale orchestration, cluster schedulers, or AI service deployment pipelines. Passion for solving tough technical problems and shipping high-impact solutions. NVIDIA is widely considered one of the most desirable places to work in tech – we are passionate about what we do and are committed to fostering a culture of excellence, innovation, and collaboration. If you’re excited to help define how the world runs AI at scale, this role is for you.