About the Role
We're growing our Data team and looking for a Data Engineer who lives and breathes data. At Appdome, data moves fast and at scale: trillions of live threat events feeding real-time detection, risk scoring, and response. This role is about designing data so it stays fast under heavy volume: the right architecture, models, and storage choices to keep latency low and throughput high. You'll own data systems end to end, from ingest to serving, and have real influence over how we build.
What You’ll Do
Design, build, and maintain scalable, low-latency data pipelines, ETL/ELT, and infrastructure.
Architect data models, schemas, and storage layouts optimized for fast retrieval at large volume.
Build real-time / event-driven ingestion and processing for high-throughput threat and risk data.
Optimize SQL and NoSQL databases - query tuning, indexing, partitioning, cost/performance tradeoffs.
Work with big-data volumes across pipelines, data lakes, and warehouse / lakehouse layers, using Spark for large-scale processing and enrichment.
Own data observability - pipeline health, freshness, and quality.
Collaborate with software engineers, security experts, and data scientists to integrate data into Appdome's products.
Requirements
B.Sc. in Computer Science, Data Engineering, or a related field.
3+ years hands-on with large-scale data infrastructure.
Strong Python (including Pandas).
Deep SQL and NoSQL knowledge, with real performance-optimization experience.
Proven experience with real-time/low-latency systems, and the data architecture and modeling instincts that make them fast.
Comfort with big-data volumes, pipelines, and data lakes.
Experience with streaming / event-driven design (e.g., Kafka).
Spark / PySpark for large-scale processing, or a strong adjacent big-data background.
Strong problem-solving skills and a proactive, independent mindset.
Advantages
Experience with ClickHouse.
Data warehousing and lakehouse / open table formats (e.g., Apache Iceberg).
Vector databases and embeddings - storing and serving them to support AI and data-science features (e.g., pgvector).
Strong system design and data architecture sense - knows how to build the data model the right way.
Experience using AI-assisted development tools (e.g., Cursor, Claude Code) in day-to-day work.
Experience with security, fraud, or other high-volume telemetry data.
Comfort with containerized infrastructure (Docker, basic Kubernetes) and light DevOps (CI/CD, Git).
ElasticSearch, Redis, DynamoDB, or Metabase.
Who you are
Independent and self-driven - comfortable owning systems end to end.
Growth-oriented - eager to develop and take on more over time.
Collaborative - communicates well across engineering, security, and data science.
Adaptable - thrives in a fast-paced environment with a can-do attitude.