A Defense-Tech company developing advanced AI and real-world technology is looking for a Senior DevOps Engineer (Backend Oriented) to take ownership of its infrastructure and deployment environment. The role combines hands-on DevOps with real backend development, Docker, and production on-premises environments.
What You’ll Do:
Design, build, and own production infrastructure across on-premises and cloud environments, including customer deployments, upgrades, remote management, and recovery.
Build and maintain Dockerized services and multi-container deployments using Docker Compose.
Own CI/CD, Infrastructure-as-Code, release processes, and deployment strategy from development through production.
Build observability across the stack, including monitoring, logging, alerting, and operational practices.
Manage GPU-based workloads and ensure AI workloads run reliably on real hardware.
Contribute to backend services and APIs, improving performance, data processing, database access, and real-time event handling.
Own networking, secrets management, security hardening, and infrastructure reliability.
Work closely with backend, ML, and product teams to improve scalability, deployment, and operational excellence.
Requirements
5+ years of experience in DevOps, SRE, Platform, or Infrastructure Engineering.
Strong hands-on backend development experience, preferably with Python.
Deep hands-on production experience with Docker and Docker Compose.
Real-world experience operating production systems on-premises, ideally on physical servers.
Strong experience with AWS, GCP, or Azure, and with CI/CD and IaC tools such as Terraform or Ansible.
Strong Linux, networking, and system-level troubleshooting skills.
Production experience with PostgreSQL and observability tools such as Prometheus, Grafana, or Loki.
B.Sc. in Computer Science or a related degree, or equivalent academic background.
Proactive, hands-on, and comfortable taking ownership in a fast-moving startup environment.
Bonus Points For:
On-prem / air-gapped environments, edge devices, GPU/NVIDIA.
Real-time video, streaming, distributed systems, or strict performance constraints.
Python/FastAPI, ML infrastructure, LLMs, vector DBs, or AI agents.
Defense-tech, security, robotics, autonomous systems, or physical-world AI.
0→1 / early-stage startup experience.