Role Summary
Project A is ALX's AI learning platform – a set of LLM products used by learners. Every one generates a stream of LLM data, and every one has hypotheses baked into it about what "working" means. The LLMOps Engineer owns the analyzer function: turning that stream into an honest answer about whether the products work. Take RAG as one example – documents must be stored accurately, fetched accurately, and fetched in the right mixture: three separate failure modes, each needing its own eval. Every product decomposes like that. This is a junior-to-mid role with a deliberate growth path: you start close to the technical lead's designs and grow into full ownership of the function.
You will work in collaboration with Anthropic Engineers, a cross functional team of AI engineers, product managers and data scientists to design world class learning experiences.
Specific Responsibilities
Evaluation Suites
Build and run eval suites per product, decomposed by failure mode, running on schedule and on every release – regression testing so nothing ships if it broke what worked.
Keep evals cost-effective as the product line grows.
Reporting, Data & Collaboration
Own the reporting loop – findings from evals and platform data in front of the team and stakeholders, including surfacing unintended or problematic model behaviour before learners do.
Steward the core datasets the team depends on, including classified customer-support data.
Partner wit