OutFigure Get in touch

baseline-v1

The model card for baseline-v1: what it predicts, what it reads, when it declines to predict and what it must not be used for, published as it is written beside the model's code.

At a glance

Version
baseline-v1
Kind
smoothed_learner_accuracy
Status
shadow, unvalidated. Never promoted, never shown to a learner.
Written
2026-09-16, with the prediction routes (backlog B-40, SCI-01 to SCI-03)
Code
packages/engine/src/baseline-v1.ts; tests packages/engine/src/baseline-v1.test.ts
Artifact
SHA-256 of the definition returned by baselineArtifact(), stored with every prediction

What it does

For one learner, one concept and one item version, it estimates the chance that the learner's next eligible response on that concept is correct, inside a declared window.

The rule is the learner's own eligible accuracy, smoothed:

p_correct = (correct_attempts + 1) / (eligible_attempts + 2)

Eligible means what eligibility-v2 means everywhere else in this system: a binary response, explicitly unassisted, the first for that item in its session, with no revealed answer, collaborator, hint or feedback attached to the attempt (docs/concepts/state-and-evidence.md).

The same rule is one of the baselines the offline audit evaluates, as learner_accuracy, so a prospective result here and a retrospective result there measure the same thing.

Target

next_eligible_response_v1: the first eligible response by this learner on this concept and item version, occurring after the prediction was issued and within horizon_seconds (one hour to ninety days; seven days by default).

This is a next-response target inside a window. It is not retention, not mastery, and not a statement about what the learner knows. A window that closes with no eligible response is censored, which is missing evidence and never a wrong answer.

Inputs

Exactly these, read from the learner-concept state computed at the feature cutoff and frozen with the prediction:

baseline-v1: inputs
FeatureUse
eligible_attemptsDenominator, and the minimum-evidence check
correct_attemptsNumerator
distinct_sessionsMinimum-evidence check only
last_eligible_at, hours_since_last_eligibleRecorded, not used by this rule
state_revision, eligibility_version, windowProvenance

No item difficulty, no other learner's data, no content, no timing signal, no assistance inference. Nothing outside the named learner and concept enters the number.

Minimum evidence

Below 5 eligible attempts or 2 distinct sessions the model refuses: the API returns 422 insufficient_evidence and no prediction is stored. These are the same disclosure thresholds the observed summary uses. With no evidence at all, the smoothing alone would return 0.5, and a made-up 0.5 is worse than saying nothing.

Uncertainty and calibration

uncertainty.kind is not_estimated, interval is null, and calibration_status is unvalidated. The smoothing is a Beta(1, 1) prior in count form; that is a modeling assumption, not a measured property of these learners, and a posterior interval from it would not be a calibrated prediction interval. Nothing here has been checked against outcomes.

How it would be evaluated

Not yet evaluated. When enough matched and censored outcomes exist, the same protocol the offline audit uses applies (docs/guides/audit-evaluation.md): chronological split fixed before results are examined, Brier score and log loss as primary metrics, calibration curves with bin counts, coverage of matched outcomes, and learner-clustered bootstrap intervals. Comparison against the partner's own heuristic where one is supplied.

Excluded uses

  • Anything a learner sees or that changes what a learner is given. Predictions are shadow mode; the M4 gates (A29) and the owner's authorization would be required first.
  • Grading, placement, certification, admission, employment or any decision about a person.
  • A claim that the learner has or has not learned something.
  • Live projects: the routes refuse to issue a prediction outside a sandbox project while the model is unvalidated.

Failure modes

  • Sparse or lopsided history. Five attempts is a disclosure threshold, not a sample size. Early predictions will be noisy.
  • Selection. Only learners who return produce outcomes; censored windows are not missing at random, and any evaluation must report coverage.
  • Item mix. The rule ignores which item version is being predicted, so an unusually hard or easy item is not accounted for.
  • Stale evidence. Elapsed time is recorded but unused: a prediction from month-old evidence looks the same as one from yesterday.
  • Assistance reporting. If a partner does not report assistance, those responses never count as eligible, and the model sees less than the learner actually did.

Rollback

There is nothing to roll back to: this is the first model, it is not promoted, and no decision depends on it. Removing it means not calling POST /v1/predictions; stored predictions and outcomes stay as the record of what was predicted, under the artifact hash that produced them. A change to any parameter in baselineArtifact() is a different artifact hash and therefore a different model row, so old predictions can never be re-attributed to a new rule.