What we need to see: MSc in Computer Science, Electrical Engineering, or a closely related field; or equivalent experience in an industrial research role. At least 5 years of proven experience in applied research, research engineering, or algorithm engineering. Excellent software engineering skills, particularly in Python and deep learning frameworks like PyTorch. Proven experience with High-Performance Computing (HPC) environments, including training or running inference on large-scale GPU clusters (tens to hundreds of GPUs). Interest in the systems side of deep learning, including inference engines, benchmarking, profiling, GPU efficiency, memory behavior, and deployment constraints. A strong problem-solving mentality and a proactive attitude, driven by the ambition to deliver solutions with real-world impact. Ways to stand out from the crowd: At least one publication in a top-tier AI/ML conference (e.g., NeurIPS, ICLR, ICML). Deep understanding of LLM architectures coupled with hands-on experience in training large-scale models. Hands-on research experience in LLM inference optimization algorithms such as speculative decoding or parallelization strategies. Deep familiarity and experience with popular LLM inference frameworks (e.g., vLLM, TensorRT-LLM). We are an equal-opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.