You have built production data pipelines in the cloud, setting up data-lake and server-less solutions; you have hands-on experience with schema design and data modeling and working with ML scientists and ML engineers to provide production level ML solutions. You have experience designing systems E2E and knowledge of basic concepts (lb, db, caching, NoSQL, etc) Strong programming skills in languages such as Python and Java. Experience with big data processing frameworks such, Pyspark, Apache Flink, Snowflake or similar frameworks. Demonstrable experience with MySQL, Cassandra, DynamoDB or similar relational/NoSQL database systems. Experience with Data Warehousing and ETL/ELT pipelines Experience in data processing for large-scale language models like GPT, BERT, or similar architectures - an