What we need to see: B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent experience 5+ years of experience in performance analysis, systems engineering, or HPC/AI infrastructure Demonstrated expertise in performance analysis skills and methodologies Hands-on experience with high-performance networking (RDMA, MPI, NCCL, congestion control) Strong understanding of system performance metrics (latency, throughput, resource utilization) Exposure to hardware, firmware, or embedded telemetry environments Strong analytical, problem-solving, and communication skills Ability to work effectively in cross-functional, fast-paced R&D teams Ways to stand out from the crowd: Knowledge of CUDA, NCCL internals, and congestion control algorithms Deep system-level understanding of CPU architectures, GPUs, HCAs, memory, and PCIe Experience with NVIDIA GPUs, CUDA, and deep learning frameworks such as PyTorch or TensorFlow Experience with cloud platforms Proficiency in Python; experience with Bash and C/C++ is a