Role Overview
This is a senior, hands-on role combining:
70% Production Infrastructure & Kubernetes
30% Backend Development (Node.js)
You will own production and non-production environments end-to-end, lead infrastructure scalability initiatives, and contribute to backend services and internal tooling in Node.js.
We are looking for someone with strong production ownership mindset, deep Kubernetes understanding, and real experience solving incidents in high-scale systems.
What You’ll Do
Own production and staging environments (high availability, scalability, reliability)
Manage and operate Kubernetes clusters end-to-end (deployments, scaling, networking, upgrades, troubleshooting)
Design and maintain CI/CD pipelines
Improve infrastructure automation using Infrastructure as Code (Terraform preferred)
Investigate and mitigate production incidents, perform root cause analysis
Optimize system performance and cost
Develop backend services and internal tooling in Node.js (approx. 30% of the role)
Write automation scripts (Bash / Python when needed)
Requirements
Must Have
4+ years in DevOps / Infrastructure / Platform Engineering roles
1+ year of hands-on experience with Node.js (backend services or internal tools)
Proven experience supporting high-scale production systems
Hands-on Kubernetes experience (cluster operations, scaling, troubleshooting)
Strong understanding of production incident management and SLA-driven environments
Experience building or maintaining CI/CD pipelines
Strong debugging skills – ability to read logs and dive deep into root cause
Fluent English
Nice to Have
Experience working in a startup or high-growth company
Experience with GCP
Experience with TypeScript / React
Must Have
4+ years in DevOps / Infrastructure / Platform Engineering roles
1+ year of hands-on experience with Node.js (backend services or internal tools)
Proven experience supporting high-scale production systems
Hands-on Kubernetes experience (cluster operations, scaling, troubleshooting)
Strong understanding of production incident management and SLA-driven environments
Experience building or maintaining CI/CD pipelines
Experience with Terraform in large-scale environments
Strong debugging skills – ability to read logs and dive deep into root cause
Fluent English
Nice to Have
Experience working in a startup or high-growth company
Experience with GCP
Experience with TypeScript / React