What We're Looking For: 5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer. 3+ years of experience operating production, 24x7 customer-facing systems. Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams. Strong software engineering skills in Python, Go, Java, or a similar language, with a bias toward production-quality code, tests, monitoring, and documentation. Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation. Experience building, tuning, and automating observability systems such as Grafana, Prometheus, New Relic, Datadog, Splunk, or similar tools. Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems. Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools. Ability to troubleshoot AI tools and provider/API issues, including rate limits, quota, auth, permission errors, latency, SDK or API contract changes, content quality issues, and service degradations. Excellent communication skills and the ability to work with stakeholders and domain experts across the company.