96 problems · 3 interview rounds
The interview platform built for AI Engineers.
Write real AI/ML code against real tests. Design production ML, LLM and agentic systems on a canvas. Then sit a mock interview that talks, listens and pushes back, all graded against a rubric written for that specific problem.


Coding
Write the code, run the tests
Python in your browser against visible and hidden tests, then a reviewer that reads your method, not just whether it passed.
- Vectorised NumPy, real DataFrame work, layers and gradients from scratch
- Chunkers, retrievers and rerankers written by hand
- Correctness scored from the actual test run, not guessed at
- Hints, a reference solution and a discussion thread beside every problem


System Design
Build the architecture, defend it
Place components, configure them, connect the data flow and say why. Then answer the questions a diagram cannot.
- ML systems, LLM applications, RAG, agents and multi-agent systems
- Typed configuration: index, metric, dimensions, step limits
- Advisory design checks that name what a credible answer is still missing


Interview
Then sit the round, out loud
Solving a problem alone and being asked about it while you solve it are different skills, and only one of them gets tested in an onsite. So the interviewer here is live: it greets you, reads the question out, lights up the line it is asking about, and follows up on what you actually said.
Write it while they watch, and defend the choices. Follow-ups on complexity, edge cases and what you would change at scale.
Build the architecture on the canvas and argue for it. Retrieval, training, serving, evaluation, and what breaks first.
Rapid-fire across a topic: models, data, training, evaluation. No editor, no canvas, just whether you can explain it.
It speaks, and it listens
The interviewer reads the question out and its reply arrives a word at a time, in step with the voice. Hold S and talk, or type. The transcript is written either way.
It follows the thread
Scoping, main round, deep dive, then your own questions. Phases move on the clock and on what you have actually covered, not on a turn count.
It grades like the rest
A round ends as an ordinary submission: the same rubric, the same dimensions, plus a hire recommendation and the follow-ups you fumbled.
Progress
A career ladder, not a completion bar
XP, crowns and a level derived from your own graded attempts, plus the dimension that is quietly holding you back.
- XP and a level from Intern to Distinguished, earned from difficulty and best score
- Three crowns per problem, and a roadmap per track so there is always a next one
- Separate readiness for coding and system design, broken down to the dimension
- Nothing is stored: every number is a view of your scores, so it cannot drift from the work


Realistic Problems
Problems built from real AI engineering scenarios: scale numbers, cost budgets, messy data and the constraints that make the decisions hard.
AI Evaluation
Structured feedback on your method, your architecture and your trade-offs, scored against a rubric written for that specific problem.
Measured Progress
Readiness per arena and per topic, derived from your own graded attempts, including the dimensions that are quietly holding you back.
Coding
Write and run Python against real tests.
Vectorisation, broadcasting and linear algebra without Python loops.
Reshaping, grouping, joining and cleaning real analytical data.
Pipelines, validation and evaluation done without leaking labels.
Layers, gradients and training loops implemented from scratch.
Tokenisation, embeddings, similarity and text classification.
Chunking, retrieval, reranking and context construction in code.
Batching, caching, retries and the plumbing around a model.
System Design
Build production AI architectures on a canvas.
End-to-end machine learning systems at production scale.
Gateways, routing, prompt management, guardrails and cost control.
Ingestion, retrieval, grounding and freshness over large corpora.
Tool calling, planning, memory and failure handling.
Autonomous and multi-agent systems that plan, act and self-correct.
Example problems
Every problem ships with requirements, constraints, a reference solution and a problem-specific grading rubric.
Design a Production RAG System
Design a Real-Time Fraud Detection System
Design an LLM Evaluation Platform
Diagnose Training-Serving Skew
Find out where you stand
One problem today. A round when you are ready.
96 problems across both arenas, an interviewer that will sit a round with you, and a readiness score built from nothing but your own graded attempts.
