โ ๏ธ Runs call real paid APIs. All sampled questions are sent to each model in
one call (so its "Total time" is the real round-trip). Web search is OFF โ pure reasoning,
graded objectively against the known answer. Each model shows full step-by-step working for
every question (in the drill-down). If a model can't fit all the questions and drops some
answers, those score wrong โ a fair capability strike.