OutFigure Get in touch

recent-v1

The model card for recent-v1: what it predicts, what it reads, when it declines to predict and what it must not be used for, published as it is written beside the model's code.

At a glance

Version
recent-v1
Kind
smoothed_recent_accuracy
Status
candidate, shadow, unvalidated. Never promoted, never shown to a learner.
Written
2026-10-02, with paired predictions and declared comparisons (backlog B-45, ADR 0035)
Code
packages/engine/src/recent-v1.ts; tests packages/engine/src/recent-v1.test.ts
Artifact
SHA-256 of the definition returned by recentArtifact(), stored with every prediction

What it does

For one learner, one concept and one item version, it estimates the chance that the learner's next eligible response on that concept is correct, inside a declared window: the same target as baseline-v1.

The rule is the learner's accuracy over their five most recent eligible responses on the concept, smoothed the same way as the baseline:

p_correct = (recent_correct + 1) / (recent_responses + 2)

It asks one question: does what a learner did lately say more about their next response than everything they did? The baseline reads the whole window of up to 200 eligible responses; this rule reads the last five of them. With five or fewer eligible responses the two give the same number.

Eligible means what eligibility-v2 means everywhere else (docs/concepts/state-and-evidence.md).

How it is issued

Never on its own. Every POST /v1/predictions request freezes a prediction from each model that supports the target, on the same state, issue time and window (ADR 0035). The requested model, baseline-v1 unless the request names this one, is returned; the others are stored as shadows. One response answers the whole set, and a window that closes in silence censors it together, so evaluation runs compare the two rules on exactly the same predictions.

Inputs

Read from the learner-concept state the prediction was issued from, and frozen with it:

recent-v1: inputs
FeatureUse
recent_correct, recent_attemptsNumerator and denominator
recent_windowlatest_5_eligible_responses
eligible_attempts, distinct_sessionsMinimum-evidence check, the same thresholds as the baseline
correct_attempts, last_eligible_at, hours_since_last_eligibleRecorded, not used by this rule
state_revision, eligibility_version, windowProvenance

The recent outcomes are stored with each state when it is computed, from that revision's own evidence (recent_eligible_outcomes), so a prediction can never read a response the state did not count. A state computed before this model existed has none: the model is left out of sets issued from it, and asking for it directly returns 422 insufficient_evidence with the code state_predates_model until the learner's next response recomputes the state.

Minimum evidence

The baseline's: below 5 eligible attempts or 2 distinct sessions it refuses, and no number is stored.

Uncertainty and calibration

uncertainty.kind is not_estimated, interval is null, and calibration_status is unvalidated. Five responses carry little information: the smoothed value moves in steps of one seventh, and any interval drawn from them would not be a calibrated prediction interval.

How it would be evaluated

In evaluation runs (ADR 0032, ADR 0035), against baseline-v1, on the requests where both predictions were matched to the same response: Brier score and log loss for each, the difference, and a learner-clustered bootstrap interval for it. A verdict is given only when the comparison was declared with the run, with its primary metric and the smallest improvement that counts. Below 30 pairs or 10 learners nothing is reported. A verdict, whichever way it goes, never promotes the model; that is a reviewed decision (SCI-10). No comparison on real data exists.

Excluded uses

The baseline's, without exception: nothing a learner sees, no decision about a person, no claim about what a learner knows, and no live project while the model is unvalidated.

Failure modes

  • Noise. Five responses are few. One answer moves the estimate by up to a seventh, so it will often look worse than the baseline on a small sample even where recency matters.
  • Recency without time. The last five responses may span a day or a year; elapsed time is recorded and unused.
  • The baseline's failure modes apply too: selection by who returns, item mix ignored, and assistance that goes unreported.

Rollback

Nothing depends on it. Removing it from the build stops new shadows; stored predictions and outcomes stay as the record of what was predicted, under the artifact hash that produced them. A change to any parameter in recentArtifact() is a different artifact and a different model row.