recent-v1
The model card for recent-v1: what it predicts, what it reads, when it declines to predict and what it must not be used for, published as it is written beside the model's code.
At a glance
- Version
recent-v1- Kind
smoothed_recent_accuracy- Status
- candidate, shadow, unvalidated. Never promoted, never shown to a learner.
- Written
- 2026-10-02, with paired predictions and declared comparisons (backlog B-45, ADR 0035)
- Code
packages/engine/src/recent-v1.ts; testspackages/engine/src/recent-v1.test.ts- Artifact
- SHA-256 of the definition returned by
recentArtifact(), stored with every prediction
What it does
For one learner, one concept and one item version, it estimates the chance that the learner's next eligible response on that concept is correct, inside a declared window: the same target as baseline-v1.
The rule is the learner's accuracy over their five most recent eligible responses on the concept, smoothed the same way as the baseline:
p_correct = (recent_correct + 1) / (recent_responses + 2)It asks one question: does what a learner did lately say more about their next response than everything they did? The baseline reads the whole window of up to 200 eligible responses; this rule reads the last five of them. With five or fewer eligible responses the two give the same number.
Eligible means what eligibility-v2 means everywhere else (docs/concepts/state-and-evidence.md).
How it is issued
Never on its own. Every POST /v1/predictions request freezes a prediction from each model that supports the target, on the same state, issue time and window (ADR 0035). The requested model, baseline-v1 unless the request names this one, is returned; the others are stored as shadows. One response answers the whole set, and a window that closes in silence censors it together, so evaluation runs compare the two rules on exactly the same predictions.
Inputs
Read from the learner-concept state the prediction was issued from, and frozen with it:
| Feature | Use |
|---|---|
recent_correct, recent_attempts | Numerator and denominator |
recent_window | latest_5_eligible_responses |
eligible_attempts, distinct_sessions | Minimum-evidence check, the same thresholds as the baseline |
correct_attempts, last_eligible_at, hours_since_last_eligible | Recorded, not used by this rule |
state_revision, eligibility_version, window | Provenance |
The recent outcomes are stored with each state when it is computed, from that revision's own evidence (recent_eligible_outcomes), so a prediction can never read a response the state did not count. A state computed before this model existed has none: the model is left out of sets issued from it, and asking for it directly returns 422 insufficient_evidence with the code state_predates_model until the learner's next response recomputes the state.
Minimum evidence
The baseline's: below 5 eligible attempts or 2 distinct sessions it refuses, and no number is stored.
Uncertainty and calibration
uncertainty.kind is not_estimated, interval is null, and calibration_status is unvalidated. Five responses carry little information: the smoothed value moves in steps of one seventh, and any interval drawn from them would not be a calibrated prediction interval.
How it would be evaluated
In evaluation runs (ADR 0032, ADR 0035), against baseline-v1, on the requests where both predictions were matched to the same response: Brier score and log loss for each, the difference, and a learner-clustered bootstrap interval for it. A verdict is given only when the comparison was declared with the run, with its primary metric and the smallest improvement that counts. Below 30 pairs or 10 learners nothing is reported. A verdict, whichever way it goes, never promotes the model; that is a reviewed decision (SCI-10). No comparison on real data exists.
Excluded uses
The baseline's, without exception: nothing a learner sees, no decision about a person, no claim about what a learner knows, and no live project while the model is unvalidated.
Failure modes
- Noise. Five responses are few. One answer moves the estimate by up to a seventh, so it will often look worse than the baseline on a small sample even where recency matters.
- Recency without time. The last five responses may span a day or a year; elapsed time is recorded and unused.
- The baseline's failure modes apply too: selection by who returns, item mix ignored, and assistance that goes unreported.
Rollback
Nothing depends on it. Removing it from the build stops new shadows; stored predictions and outcomes stay as the record of what was predicted, under the artifact hash that produced them. A change to any parameter in recentArtifact() is a different artifact and a different model row.
