How models, context, and memory fit together
A practical map of generation, working context, retrieval, tools, and systems of record.
Evergreen explainerA practical field guide to AI systems
Understand the layers, ground the work in evidence, choose the simplest reliable system, and count the cost of results you can actually use.
Open the reference cabinetOpen the reference cabinet
A practical map of generation, working context, retrieval, tools, and systems of record.
Evergreen explainerMove from discovery to claim ledger, synthesis, counterevidence, and publication checks.
Step-by-step workflowMatch system complexity to path variability, external evidence, and action risk.
Architecture guideSeparate model, retrieval, OCR, schema, tool, permission, and interface failures.
Diagnostic orderSet the privacy, permission, confirmation, recovery, and human-review boundary.
Risk guideUse your own costs, review time, retries, and acceptance rate. Calculation stays in your browser.
Interactive toolFrom the original research / 2022
A source diagram worth reading alongside the evaluation checklist.
Scroll sideways to inspect the full diagram

HELM’s original diagram makes the evaluation dimensions visible. The checks describe evaluation coverage in this historical illustration; they are not model scores.
Source: Stanford CRFM · 17 November 2022 ↗What counts as evidence?
The starting collection explains how to evaluate models. We do not imply that an illustrative task is a benchmark run, or that a vendor’s score predicts your workflow.
Read our evidence policy →