About the role:
We're looking for a hands-on DevOps engineer to design, automate, and operate the environments our platform runs in- public cloud alongside on-premises, including restricted and air-gapped customer sites, and to build them the way infrastructure gets built in 2026: as code, with AI agents doing much of the heavy lifting and your judgment deciding what actually ships.
Two things sit on your plate. The first is our R&D team: the pipelines, paved roads, and self-service tooling that let engineers ship several times a day without asking permission or holding their breath- increasingly built and maintained alongside AI agents rather than by hand. The second is where our software actually lands: customer environments that are rarely identical, often heavily locked down, and occasionally have no route to the internet at all. Making a modern, AI-powered platform run reliably inside those constraints is the genuinely interesting part of this job.
Why Join Us
Innovation: Modern stack, real cloud-native and GitOps practice, and AI tooling in daily use rather than in a slide.
Autonomy: Senior engineering ownership without management overhead. Your ideas land.
Leverage: Every environment you standardize makes the next one faster- and every workflow you hand to an AI agent stops costing you a day. You'll see your own work compound.
Do you like to fight crime? Our software helps catch financial criminals. That's the actual output of your infrastructure work.
Responsibilities:
Developer enablement: Continuously improve the tools, pipelines, and processes our R&D team depends on- including how AI coding agents plug into them safely. If something is slow, manual, or superstitious, you're empowered to fix it.
Infrastructure and automation: Design, implement, and automate our cloud infrastructure and on-premises OpenShift environments as code, not as runbooks- including the compute and GPU capacity our AI workloads run on.
Deployment and delivery: Build declarative delivery with Argo, Helm/Helmfile, and Kubernetes, so a new environment is a configuration change rather than a bespoke effort. This includes getting our AI-driven platform into restricted and disconnected sites.
On-call and escalation: Join the on-call rotation, support lower tiers through incidents, and push root-cause fixes over repeat mitigations.
Requirements
Experience: 5+ years in DevOps, Platform Engineering, or SRE, including ownership of production systems.
Debugging: Strong, systematic troubleshooting under production pressure — across Linux, networking, and Kubernetes.
Kubernetes: Real experience in both development and production environments (OpenShift an advantage).
CI/CD: Hands-on pipeline authoring (Jenkins an advantage — it's a meaningful part of our current stack).
Infrastructure as Code: IaC pipelines in practice (Terraform an advantage).
Cloud: A major cloud provider (Azure a plus).
Delivery tooling: Argo, Helm, Helmfile, or equivalent declarative deployment tooling.
Linux and scripting: Proficient in Linux, with Python or Bash.
Observability: Prometheus, Grafana, ELK, or similar.
AI-assisted workflow: Comfortable working with AI coding agents day to day- and clear-eyed about where they need supervision.
FinOps: Practical experience keeping cloud spend deliberate.
Nice to Have
Experience running AI/ML or inference workloads (GPU scheduling, vLLM, KServe, Ray, vector databases).
Genuine appetite for new technology and continuous learning.
Strong judgment about production risk- you know which changes deserve caution.
Good with people. You'll work across R&D, support, and occasionally customer-side engineers.