Lead the architectural definition, modeling, and specification of next-generation, high-performance ML compute IP and acceleration blocks for Cloud AI silicon.
Own the ML IP architecture specification throughout the entire product lifecycle: concept exploration, cycle-accurate modeling, implementation, silicon bring-up, and production.
Partner closely with leading AI research and algorithm teams (e.g., Google DeepMind, Gemini research teams) and software compiler teams (XLA, PyTorch) to explore architectural trade-offs and define hardware requirements for emerging model architectures.
Drive comprehensive architecture studies, evaluating compute dataflows, numerical formats, sparsity, and specialized acceleration mechanisms such as key-value (KV) cache optimization.
Drive performance, latency, power efficiency, and silicon area projections across model topologies and workload configurations.
Minimum qualifications:
Bachelor's degree in Computer Engineering, Electrical Engineering, Computer Science, a related field, or equivalent practical experience.
15 years of experience in computer architecture, ML accelerator design, or high-performance processor architecture.
Experience leading architectural definition and authoring architecture specifications for silicon or compute IP blocks.
Experience with performance modeling, workload profiling, and hardware-software co-design.
Preferred qualifications:
Master's degree or PhD in Electrical Engineering, Computer Engineering, or Computer Science with an emphasis on computer architecture or ML hardware systems.
5 years of experience leading the architectural definition and microarchitecture of AI/ML accelerators from concept through production.
Deep knowledge of modern deep learning workloads (Transformers, MoE, Diffusion, Generative AI inference) and their system bottlenecks (memory capacity, KV cache bandwidth, interconnect scaling).
Strong understanding of high-performance memory subsystems (custom SRAM architectures, high-bandwidth memory hierarchies, caching schemes).
Experience working with modern ML frameworks (PyTorch, JAX, TensorFlow) and ML compilers/runtimes (XLA, TVM, Triton).