qualifications: MSc or PhD in Computer Science, Electrical Engineering, Computational Biology, Statistics, Mathematics, or a related quantitative field. Strong background in machine learning, data science, statistics, or computational modeling. Hands-on experience building with LLMs and agentic AI systems. Proven ability to design evaluation methodologies for AI systems, especially LLM-based or agent-based systems. Experience working with LLM APIs such as OpenAI, Anthropic, Google, or open-source LLMs. Experience with agent frameworks or orchestration tools such as LangGraph, LangChain, or similar systems. Experience defining benchmarks, metrics, validation sets, scoring methods, or automated evaluation pipelines. Strong Python skills and ability to write clean, production-aware research code. Ability to work with complex, noisy, high-dimensional data. Strong communication skills and ability to collaborate with experts from different disciplines.