you have: Experience designing and operating telemetry, metrics collection, and alerting stacks at scale (Prometheus, Grafana, ELK/logging). Practical background in optimizing infrastructure costs and resource efficiency across cloud and on-prem. How you’ll make an impact: Keep our hybrid infrastructure (on-prem, public cloud and AI/ML clusters) highly available, performant and cost-efficient. Build internal software tooling and manage IaC pipelines in Go, Python or Rust to eliminate repetitive operations. Perform deep-dive troubleshooting across the full stack—from CDN edge configurations down to Linux kernel tuning and network layer bottlenecks. Design and maintain monitoring and alerting setups to spot and address system health issues before they impact users. Participate in on-call rotations, lead incident resolution and conduct blameless post-mortems to ensure system resilience. Why Taboola? If you ask Taboolars what they love about working here, they’ll tell you that they’ve been empowered to realize their full potential while growing and learning with smart, talented people: Adam Singolda, Taboola Founder and CEO says: “You can copy anything from another business, but you can’t copy a company’s culture.” Well-being: Enjoy comprehensive benefits (health, etc.), a fully stocked kitchen, and location-specific perks (gym partnerships, parking). Flexibility: We offer a hybrid work schedule (3 days in-office) with flexibility. Work with global leaders: Our publisher partners include Yahoo, Fox Sports, NBCU, ESPN, and CBS, while our clients include major global brands. Ready to realize your potential? Taboola is an equal opportunity employer and we value diversity in all forms. We are committed to creating an inclusive environment for all employees and believe such an environment is critical for success. Employment is decided on the basis of qualifications, merit, and business need.- Learn more about #TaboolaLife on LinkedIn,&nb