In this role you will own the design and implementation of CI/CD pipelines, test frameworks, and system-level validation for our next-generation ML inference accelerator platform. You will work across the full stack — from firmware interfaces through data-plane performance benchmarking to production fleet readiness — ensuring every component is validated end-to-end before it reaches customers.
This is a greenfield environment with rapidly growing scope: new silicon, new software stacks (vLLM, NKI, NIXL), and new fleet-scale challenges. We are looking for a senior IC who can independently drive technical decisions, scale our validation infrastructure, and raise the bar on engineering quality across the group.
Key job responsibilities
Own and evolve CI/CD pipelines — from pre-merge gates through continuous deployment to fleet.
Design and implement test frameworks that enable firmware and data-plane developers to write, run, and maintain tests with minimal friction.
Architect system-level test suites that stress control-plane and data-plane components beyond provisioning and vetting flows.
Build and maintain performance benchmarking infrastructure for LLM inference workloads (Prefill + Decode), including dashboarding and regression detection.
Drive integration of third-party vendor code (nightly drops) into CI/CD, ensuring quality gates catch regressions early.
Participate in feature design reviews, contributing test plans and challenging coverage gaps.
Define and own Continuous Testing in production environments (CTS).
Leverage AI-assisted development tools (Kiro, LLM-based code generation) to accelerate team velocity and pioneer new engineering workflows.
A day in the life
You'll start your day reviewing CI pipeline results from overnight runs, triaging failures to determine whether a regression came from a vendor code drop, a firmware change, or an ML serving stack update. Mid-morning you might pair with a hardware engineer to design test cases for a new bus-level reset flow, then pivot to extending the performance benchmarking framework to catch a latency regression. After lunch you'll join a feature design review — challenging test coverage gaps and deciding where system-level validation needs to live. The rest of your afternoon could be spent writing a new pipeline stage that gates deployment on accuracy checks, or building a dashboard that gives the group visibility into fleet-readiness metrics. Throughout the day you'll lean on AI-assisted development tools to accelerate everything from infrastructure code to root-cause analysis.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.