What we need to see:5+ years of experience in the HPC or AI/ML industry, with deep hands-on technical expertise across the AI compute stack.Deep understanding of inference serving architectures for heterogeneous compute - including serving engines (vLLM, SGLang, or equivalent), support for mixed accelerator environments, and the scheduling and memory challenges they introduce.Solid knowledge of multi-node inference, tensor and pipeline parallelism, and the trade-offs involved in scaling large models across heterogeneous GPU and accelerator clusters.Solid knowledge of KV-cache management and tiering, including disaggregated prefill/decode architectures, CPU/storage offload, and their operational implications at scale.Experience with performance benchmarking of ML workloads - defining methodologies, running experiments, interpreting throughput/latency/cost trade-offs, and communicating results to both technical and business audiences.Familiarity with CCL tuning (NCCL, RCCL) and the impact of collective communication configuration on inference and training efficiency across large GPU clusters.Familiarity with storage systems relevant to ML workloads - including high-throughput distributed file systems (e.g., Lustre, VAST, WekaIO), object storage, and checkpoint/model weight loading strategies under tight latency budgets.Experience engaging technology partners (compute, storage, silicon vendors) to define joint reference architectures and go-to-market proposals.Clear written and oral communication skills with the ability to effectively collaborate with executives, engineering teams, and external partners.Ability to write extensive technical content (white papers, technical briefs, reference architectures) for external audiences with a balance of technical accuracy and clear messaging.Travel as needed.Ways to stand out from the crowd:Hands-on experience with heterogeneous inference deployments - mixing GPU types, accelerators, or memory tiers within a single serving clus