How grading works
A score you do not understand is a score you cannot learn from. This page explains exactly what happens between hitting submit and seeing a number.
1. Every case has a rubric
A rubric is written before the case is published. It lists the criteria the answer is judged on and how many points each is worth — for example financial analysis 30, risk assessment 25, recommendation 25, market analysis 20. Each criterion also has a descriptor saying what a strong answer looks like. The weights are fixed, identical for everyone attempting that case, and not chosen by the model.
2. The model sees the case, the rubric, and your answer
Your answer is sent to the configured AI provider along with the scenario, the instructions, the rubric and its descriptors, and the model answer. It is asked to return points per criterion plus specific written feedback: what you did well, what was missing, and what to do differently next time.
The request runs at a low temperature setting, which makes the output more repeatable than a normal chat response — though not perfectly so.
3. The platform, not the model, does the arithmetic
This is the part that protects you from a bad grade:
- Every criterion score is clamped to that criterion's maximum. A model that tries to award 40 out of 30 gets 30.
- Negative scores are floored at zero.
- Any criterion the model invents that is not in the rubric is discarded.
- The total is recomputed by summing the clamped criteria. The model's own claimed total is never used — models are unreliable at arithmetic, and this number decides leaderboard position.
4. What you get back
A per-criterion breakdown, an overall percentage, a one-line verdict, and three lists: strengths, weaknesses, and concrete improvements. The feedback is the point. The number is just a way to track whether the feedback is landing.
How to read your score
- Compare against yourself. A rising trend on a criterion means the feedback is working.
- Read the weakest criterion first. That is where the next marginal point is.
- Do not over-trust a single grade. AI grading has variance. Three attempts across three cases tell you far more than one score on one case.
Where it is weak
Being straight about the limits: an AI grader rewards structure and explicit reasoning, so an answer that shows its working tends to score better than an equally good answer that states conclusions tersely. It can miss genuinely creative arguments that do not match the rubric's expectations. And it is not a substitute for a human interviewer pushing back on you.
If a grade looks wrong, it may well be. Use the report link on the case and it will be reviewed.