A practical field guide to AI systems

Find the fit.
Know the limits.

Understand the layers, ground the work in evidence, choose the simplest reliable system, and count the cost of results you can actually use.

Open the reference cabinet
Task What does success require?Evidence How was the claim tested?Limits Where does it stop working?

Open the reference cabinet

Useful methods.
Explicit assumptions.

Browse all nine entries →
TOOL / ECONOMICS

Accepted-result calculator

Use your own costs, review time, retries, and acceptance rate. Calculation stays in your browser.

Interactive tool

From the original research / 2022

See the method
behind the measurement.

A source diagram worth reading alongside the evaluation checklist.

Scroll sideways to inspect the full diagram

Stanford HELM diagram comparing single-metric evaluation with a grid of scenarios assessed across accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency.
Multiple questions for the same scenario.

HELM’s original diagram makes the evaluation dimensions visible. The checks describe evaluation coverage in this historical illustration; they are not model scores.

Source: Stanford CRFM · 17 November 2022 ↗
Open full-size diagram ↗
Put the diagram beside your own test →

What counts as evidence?

A published claim.
An observed result.
Keep them distinct.

The starting collection explains how to evaluate models. We do not imply that an illustrative task is a benchmark run, or that a vendor’s score predicts your workflow.

Read our evidence policy →