You bring at least 3–4 years of hands-on people management experience and a deep technical background in SRE or DevOps disciplines. What will you do? Leadership & Team Management Lead, mentor, and grow a team of SREs, providing technical direction, career development guidance, and day-to-day management. Own the team roadmap for reliability, observability, and automation initiatives — prioritizing work, removing blockers, and driving delivery. Conduct regular 1:1s, performance reviews, and hiring processes to build and sustain a high-performing team. Foster a culture of operational excellence, blameless post-mortems, and continuous improvement. Act as an escalation point for complex incidents and reliability issues, leading post-incident reviews and ensuring follow-through on action items. Automation & Infrastructure Design, develop, and maintain automation tools to support infrastructure and operations teams at scale. Manage pipelines and infrastructure workflows using Jenkins, Ansible, Python, and Bash. Drive the adoption of infrastructure-as-code practices across the organization. Collaborate with system engineers to improve scalability, performance, and fault tolerance of critical systems. Monitoring & Observability Build and extend monitoring and alerting systems using Grafana, the ELK (Elastic) stack, Zabbix, and custom scripts. Implement and enforce observability best practices to ensure full visibility into systems, applications, and infrastructure. Define and track SLIs, SLOs, and error budgets across key services. Partner with development teams to embed observability earlier in the software development lifecycle. Database & Platform Support Support monitoring and infrastructure integration for databases including MongoDB and PostgreSQL. Maintain documentation and champion knowledge sharing around automation, monitoring, and reliability practices. What you need: Experience & Leadership 3–4+ years of experience in a