Skip to content

Migration, adoption and rollback plan

Temporary planning document. Planning only: no migration, backfill, inventory of production data or cutover is authorised by this page. MIG1 authorises planning outcome migration; execution, activation and live migration each need separate approval. This page covers adoption for every domain in the integrated plan, not just outcomes. The outcome specifics stay in the outcome-migration proposal, which this plan adopts as the proposal for Q-05.

Revised after round-2 review (3 October 2026). Ownership is now enforced by markers on the documents legacy writers write, cutover follows ADR-020's lock, verify, stamp and release protocol, and rollback and restore follow the consistency model (§6, §13 and §17.6). The allocation and target rows come from programme integration §3.

1. Principles

  1. Additive and evidence-preserving. Never delete or overwrite reviewer data. Originals, manifests and adapters are retained after any cutover.
  2. No fabricated history. Legacy records become labelled current snapshots (historyCoverage), never invented versions, votes, adjudications, gold history or original timestamps (EX2, research §5). Values that may be untouched defaults are labelled "value or default (unknown)", never read as answers.
  3. Per-project adoption. New projects start on the canonical path once admitted. Existing projects adopt one complete scope at a time after a reviewed manifest. A scope is never split between legacy and canonical writers (C16).
  4. Every writer and reader is accounted for. Writers must either use canonical commands or be refused for canonical scopes through R0's ownership markers and write guard (research A27):
  5. interactive saves, screening and reconciliation;
  6. reviewer session removal (a hard delete today);
  7. question edit and the question-delete cascade through the legacy API, which the new editor and the #3934 import both use;
  8. reference-file screening import, and question-template import (#2781/#3934, through the legacy question API above; this is not annotation import). FEAT-004 annotation import from other tools, if D4-14 approves it, is a canonical lane after R2a with Imported provenance (C3), not a legacy writer;
  9. bulk study update v2 (with study locks) and batch risk of bias;
  10. FEAT-024 writers, including the fold worker that bumps Study's audit version;
  11. the inclusion recalculation and every other UpdateMany on pmStudy;
  12. the tracking writers (hub, PM consumers, claim pipelines, typed admission);
  13. the notification stack's study-issue (#3945) and checked-PDF (#3947) Study writers;
  14. the reversible-deletion scheduler once built, and import-failure compensation that deletes studies;
  15. bulk PDF finalisation; preview seeding (its delete is exempt from bulk locks);
  16. admin tools, Quartz jobs, seed and fixture loaders;
  17. support-impersonation writes, which must record the real actor.

Readers need the same treatment (route, refuse or adapt): pool filters, capacity guards, reconciliation readiness, StudyStats, exports, the AF2 reconcile source, FEAT-024 classifiers, the allocation and eligibility facts, the presence snapshot and FEAT-024's availability calculators. #3944's candidate lookup reads legacy sessions and stores their IDs, so conversations refuse canonical scopes until R4a binds them to the task. The Study canonical summary (C1), merged by R0's floor, keeps legacy readers correct during coexistence. The merged bulk-update study locks (#3909) must be honoured by canonical writes. The inventory requirements are in the consistency model §6.5. 5. Compatibility before canonical writes (R0). Old binaries must capture and write back new fields, and R0 also ships the reader merge that keeps Study's computed legacy fields correct. Legacy writers refuse canonical scopes by a marker on the documents they write, checked in aggregate methods and by a composite write guard: data, not configuration. Flags gate only new admission. Each release ADR records its minimum rollback image. 6. Canonical-aware rollback. Before canonical writes exist, roll back by routing reads back to intact legacy data. After canonical writes exist, roll back by stopping new writes, keeping new data authoritative for what it holds, and using read-only containment or a verified forward-recovery adapter. Never down-migrate destructively, never let an old writer strip fields, and never let a configuration change hand a canonical scope back to legacy writers. 7. Rehearse before running. Every cutover is rehearsed on synthetic fixtures, then on an authorised non-production copy, including induced failure at each cutover step (copy, verification, lock, delta, stamp and release), and an image rollback with canonical data present. 8. Erasure and retention are designed, not assumed (E32). Account deletion's existing behaviour (the irreversible Deactivated record) is extended to immutable revisions, exposure events, receipts and inbox items; drafts, exposure events and notifications get retention rules. Answers stay attributed to an anonymised identity, and as-of exports are identical except erased identities, which manifests record (D2-14). If today's behaviour can't extend, Chris decides.

2. Adoption modes

Mode Who What happens When available
Greenfield New projects admitted through R0's admission service Everything created on the canonical path from the start Forms and sessions from R2a (one stage) and R2b (several stages); steps and decisions from R3a; profiles from R3b; reconciliation from R4a; schemas from O1; classification from C1. Default for all new projects only from the GA milestone.
Reviewed adoption Existing projects chosen by Chris Inventory → manifest → shadow → verify → fenced cutover → monitor, per complete scope R6 waves, each separately authorised
Legacy All other existing projects Unchanged legacy behaviour, legacy reads/exports, no partial new semantics Indefinitely, until adopted or retirement is separately approved

Scope completeness (a G-ADOPT check): a stage is adoptable for a domain only when every reader and writer of its data is canonical. In practice:

  • annotation-only stages without reconciliation: after R2a/R2b;
  • Combined stages (screening and annotation together need atomic Complete-and-Include): after R3a;
  • anything that is reconciled: after R4a, and after R4p for profile reconciliation;
  • extraction: after O1 and O2.

A "sessions only" wave that leaves screening or reconciliation of the same stage on legacy writers would split ownership and is never offered.

3. Domain mapping (what legacy data can and cannot become)

Legacy evidence Canonical target Never inferred
Project annotation questions Question identity plus a v1 question version (Scope = project), keeping original IDs. Adopted questions count as published, so they can never be permanently deleted afterwards (QD1). Earlier wording or options that were overwritten
System questions Immutable snapshots per question and SystemQuestionVersion (v0 and v1 projects differ structurally), pinned by adopted form versions That a code change since authoring left the definition unchanged
Stage question assignments One initial form version per stage assignment set, bound to that stage. Merging identical sets across stages into one shared form is an admin-reviewed choice, never automatic. That two stages historically shared evidence
Stage targets (Stage.SessionCountTarget, an override or inherited from the project screening threshold) The form's minimum target (an operational setting, D2-05), materialised from the stage's effective value at adoption, so later changes to the agreement threshold no longer move it. When stages merged into one shared form have different targets, an admin chooses (AP-13; criterion AC-R6-14) That the screening threshold was meant as an annotation target; any target that tracks the threshold after adoption
Annotation sessions (Incomplete/Completed, Reconciliation flag) Legacy session snapshot: current explicit state only, labelled current-only; nullable timestamps kept. Per E10: an answer whose stage differs from the session's is marked "membership-uncertain" (a stage-A session's view drops answers later re-saved from stage B); merging stages into one shared form leaves two sessions whose statuses may conflict, so an admin chooses; legacy "Completed" sessions were never validated on the server, so they adopt as "legacy-completed, unvalidated" and an admin decides whether they count Prior Save/Complete versions; an Include vote; a completed contribution under a new form version; validity
Annotation answers and parent/child trees Annotation heads with legacy IDs (kept through LegacyIdAlias) and a single legacy revision; ambiguous edges preserved and flagged; adopted revisions carry AuthoredUnder: Verified(v1) when the stored wording equals the adopted v1 wording, otherwise Unknown, which is excluded from same-version agreement and exact-match prefill; conflicting legacy duplicates adopt as one Conflicted head (C2, C3) Context sharing from text similarity; owned-decision edges from answer text; the wording an answer was authored under
Cross-stage answer replacement effects Current values only, with stage provenance where stored Values that a later save overwrote
Screening records (project + screener) Decisions under an admin-labelled "legacy project screening" compatibility profile; current value only. For projects admitted before adoption, the R2a capture log adds the screening history recorded since admission. Earlier decisions, stages, criteria, reasons or adjudications; profile split by StageId
Project agreement threshold, InclusionInfo[] The legacy compatibility profile's collective rule, reproducing the characterised legacy maths; recalculated outcomes. Decoupled from targets: after adoption the threshold supplies no stage or form target (AP-13) Separate historical profiles per array entry
Reconciled answers and reconciliation sessions Legacy authority snapshot marked LegacyAuthorityUnknown, treated per Q-35 (recommended: exportable as "legacy reconciled (authority unknown)", never a gold snapshot, not queryable, and a canonical reconciliation is needed for gold) Source candidate vectors; resolver history; an accepted snapshot timeline
Single-reviewer studies No automatic promotion to gold (Q-29); exports label them "single reviewer, unreconciled" Gold
Outcome data and time points Per the outcome-migration proposal: legacy-compatible schema series with stable observation paths. Each outcome measure gets a single canonical direction (ODIR1). Per-source or per-row values are evidence only: legacy GreaterIsWorse is a non-nullable bool defaulting to false, and the client defaults error type SD, average type mean and zero animals, so those values count as "value or default (unknown)" unless the dry-run proves an explicitly stored answer. The outcome-level system answer is preferred over row copies, and stale copies are flagged. Values are consolidated only within one author's (or one reconciled authority's) series for one outcome, never across reviewers. Conflicting values block binding until a reviewed mapping or separate measures. Missing values from zero/defaults; a majority direction; a direction from untouched defaults; cohort enrolment from result N
Permissions and groups Existing project/stage grants preserved exactly; owner-only actions enforced (R1b) Any broadened default (ChangeOwner, AssignPermissions and Delete stay owner-only)
Materialized statistics Rebuilt projections with the legacy basis labelled Historical checkpoints for semantics that didn't exist
Reservations and claims Drained or translated under a coordinated lease transition at cutover; enabling review eligibility requires migrating legacy reservations first Reviews, drafts or votes
Proportional allocation regime (if enabled on the stage) A frozen legacy regime record kept for provenance; allocation is disabled on adoption unless AL1 is live (D3-13a, AP-13; criterion AC-R6-14) Shares, bucket assignments or plans on the canonical form; claims from historical allocation
Progressive batch plans (#3939, if merged) Existing immutable membership kept; completion stream rebound to canonical requirements Re-ordered membership; any PRISMA count from batch release that amendment A doesn't define
Search/import records and PRISMA source data Citations and source columns only where evidence exists; otherwise unknown/unclassified with coverage labels. Backfilled Citations come from re-parsed retained reference files, or are labelled as derived from current Study metadata (which bulk updates may have rewritten), never presented as "raw, as imported". Source type is inferred only where FEAT-011's table allows (for example PubMed XML → Database). A guessed "Database" source; deduplication as a side effect of any other migration
Notification-stack aggregates (StudyConversation, StudyIssue, inbox SourceIds) Listed in each adoption manifest; references to legacy sessions remapped through LegacyIdAlias or marked unresolved Silent dangling references

4. Per-project adoption protocol

  1. Inventory (read-only; production inventory needs its own authorisation). Per-project counts and checksums, duplicate natural keys, invalid reviewer IDs, dangling questions/edges, custom thresholds, missing timestamps, oversized documents, writer activity. Without an authorised aggregate-only survey, storage benchmarks rest on synthetic sizes.
  2. Manifest. Deterministic source → target identities, compatibility profile labels, form-binding and session-merge choices, target and threshold mapping and any legacy allocation regime's disposition (AP-13), unresolved records with dispositions. An admin approves it; source changes invalidate it.
  3. Shadow. Idempotent, checkpointed backfill into non-authoritative storage with no serving change; reruns create nothing new. FEAT-024's staged operation fence is raised for every family for the shadow and cutover window, so nothing is served Fresh over a half-migrated population (MS-17).
  4. Verify. Semantic parity (decisions, outcomes, pools, exports, permission-filtered API output, statistics), not byte equality. Statistics parity is automated for ProjectScreening and labelled manual for other families until per-family audits exist (#3845). Pinned sessions and snapshots still read their original versions, and the integrity checker passes (consistency model §13.2).
  5. Cutover through ADR-020's protocol. An operation record with a lease and a generation; lock batches that refuse busy studies (reservations, canonical claims, open tasks, active drafts); refresh and re-verify the delta under the lock; stamp CanonicalScopes on the Project and on every Study of the scope and write the pmCanonicalOwnership registry; release with an Audit.Version bump; record the cutover revision and operator. Before the commit write a failure unstamps and releases; after it every interruption leads forward to release. Cutover is all-or-nothing through locks, not one atomic switch: legacy writes during the window are refused with a retry message that keeps drafts. Claims, presence records, connections and scheduled messages for the scope are drained or converted.
  6. Monitor and recover. Checkpointed resume, quarantine of mismatches, originals retained; legacy writers refused for the adopted scope by the markers; FEAT-024 families reset and rebuilt under the new family source version; the checker runs nightly.

5. Release-by-release adoption and rollback

Release New writes introduced Adoption scope Rollback before canonical writes Rollback after canonical writes
R0 Compatibility floor, admission Extra-element capture on extended types; reader merge logic (inert until a canonical writer exists); CanonicalScopes markers and the composite write guard; enrolment (admission) records; the writer floor Platform-wide, behaviour-neutral Image rollback to the previous release (no canonical data yet) Not applicable: R0 precedes canonical writes. Its image becomes the minimum rollback image for R2a; once canonical data exists, binaries below it are never redeployed, and canonical commands refuse while any instance runs below it
R1a Question templates and import Template imports through the existing import contract All projects, flag per capability Flag off Imported questions stay; legacy semantics unchanged
R1b Members & groups visibility Owner-only enforcement (security fix); no new data All projects Flag off for the UI; enforcement stays Not applicable
R1c/R1d Groups and delegation Configurable groups, grants, delegation envelope All projects, behind authorization gates Flag off Groups and grants stay; flag-off hides management UI but never removes or broadens grants
R2a Versioned forms, immutable sessions Question and form versions, form sessions, revisions, drafts, session versions, the Study canonical summary, the capture log Greenfield and admitted pilots Flag off; legacy path untouched Stop new canonical writes; read-only containment for pilot data; canonical exports remain; markers keep legacy writers out; image rollback no earlier than R0, whose floor keeps the summary merged; if a statistics writer changed, FEAT-024's rollback order applies
R2b Shared sessions Multi-stage bindings, re-keyed claims, form-unique tallies Same Flag off Shared sessions stay readable from every bound stage; claims drained
R2c Publication with impact Policy records, publication operation records, impact manifests (audit only), projection rewrites Same Flag off Published versions and policy records stay; a running phase-2 operation stops at its next item and resumes on roll-forward, while the per-study fail-closed rule keeps gates safe; new publications stop
R2d Overlap, Fix, FV4 Outdated flags, Fix transitions, policy revisions Same Flag off Flags/fixes stay readable; new Fix disabled
R3a Steps, routing, decisions Canonical screening decisions, stage settings versions, steps, admission records, pool-entry events Greenfield and pilots, after the screening floor step Flag off Decision- and step-aware read-only containment; legacy single-ScreeningInfo can't represent canonical decisions, so no flattening
R3b Profiles Profile versions, eligibility answers, derived decisions, reasons Same Flag off Profile-aware read-only containment
R3c Lifecycle Status events, change requests and approvals Pilots Flag off Status history stays; manual lifecycle continues
R3d Guided setup Setup drafts New projects Flag off; old wizard remains until GA Created projects stay canonical
R4a Form reconciliation and gold Tasks, matches, gold snapshots, assignments, extra-review requests Pilots Flag off Snapshots remain effective and readable; new reconciliation writes stop; candidates are never rewritten
R4p Profile reconciliation Adjudications, final screening outcomes Pilots Flag off Outcomes stay authoritative; new adjudication stops
R4b Queries Query items, concerns, resolutions, notices Pilots Flag off Open queries frozen readable; gold stays effective
R4c Outcome reconciliation Outcome series gold Pilots Flag off As R4a
R5a History and as-of export Export manifests All admitted projects Flag off Manifests are append-only; disable generation
R5c Agreement statistics None persisted beyond derived, rebuildable results All admitted projects Flag off Results are derived from canonical revisions; disable the view
R5b PRISMA reporting Frozen report snapshots All admitted projects Flag off Snapshots are append-only records; disable generation
C1/C2 Classification and inference Entity/population records, assertions, rules; inference is derived and rebuildable Pilots Flag off Explicit records stay readable/exportable; inference can be switched off at any time
O1 Outcome schemas Schema versions, bindings, schema-reference answers, measures, new-shape observations Greenfield and pilots Flag off New-shape data stays authoritative; legacy export returns a typed "unsupported shape" result rather than coercing
O2 Outcome migration Staged copies, then cutover per project Chosen projects after Q-05 and separate execution approval Route back to intact legacy data Forward recovery or verified reverse adapter only
AL1 Shared-form allocation Shared allocation plans Pilots Flag off (back to refusal) Plans stay readable; allocation falls back to refusal for shared forms
P1 Identification provenance Immutable Citation/source capture for new imports; retrieval and lifecycle events New imports in admitted projects, after the Study floor step Flag off Captured provenance retained; reports label coverage
P2 Identification and dedup Publications, Citation links, lifecycle status, dedup audit, reviewed merges Admitted projects Flag off Dedup decisions reversible through the audit; no destructive unmerge
GA milestone Canonical default for new projects New projects Revert the default; created projects stay canonical Not applicable
R6 Legacy adoption waves Adopted legacy snapshots per complete scope; markers and registry entries Projects Chris selects Per protocol step 5 (before the commit write: unstamp and release) Per protocol step 6 (forward recovery only)
R7 Retirement Removal of adapters and legacy writers Platform-wide, separately approved — Earlier images stop being valid rollback targets here; restore rehearsal is required first

6. Rollback rehearsal checklist (every release with canonical writes)

  • An older binary at the recorded minimum reads documents containing canonical fields: no exception and no stripped field, and in a mixed fleet every persisted computed field and tally equals an authoritative recount (AC-R0-09, AC-R0-06).
  • A legacy writer retrying after canonical data exists is refused by the marker and changes nothing, including writers that open no transaction (research A19; AC-R0-11; C16-T07).
  • Removing a project from enrolment, or reverting configuration, doesn't hand its canonical scope to legacy writers.
  • Image rollback to the recorded minimum image, with canonical data present; canonical commands refuse while any instance runs below the writer floor (AC-R0-15).
  • Mixed-version API and web clients during deployment: stale clients get typed retry messages that keep drafts.
  • Interrupted backfill and rerun: idempotent, no duplicate heads or revisions.
  • Operations in flight (publication phase 2, sweeps, capture moves) survive the rollback: older binaries leave their records alone, gates fail closed on stale projections, and the operations resume on roll-forward.
  • Restore per D2-13: a point-in-time restore into an isolated database; the integrity checker across collections (history, snapshots, pins, drafts, command-bearing records, the canonical summary, markers and the registry); the history-discontinuity record; scheduled-state reconciliation. Never a selective per-project overwrite of canonical data (E31, E55; AC-R2a-29, AC-R7-02).
  • Statistics fall back to authoritative queries when a projection kind is incompatible. When the release changed a statistics writer, family or protocol, the rehearsal follows FEAT-024's rollback order and records the fold mode and stamp before and after (MS-16).

FEAT-024 rollback order (MS-16). When a release changed a statistics writer, family or protocol, rollback follows FEAT-024's order: (1) fold-disable for every fold project and wait for foldMode: "Disabled"; (2) close the project narrow gate, which takes two calls around the quarantine, and the fleet gate for a full stop; (3) turn the flags off in one cluster-gitops change for both hosts, never through runtime overrides, which the PM host does not see; (4) only for an image rollback past the fold, apply the allowlist guard alone and then change the images; rolling forward again takes two resets around guard removal. The rehearsal records the fold mode and stamp before and after (docs/features/materialized-project-statistics/phase2c-staging-proof-runbook.md, Step 10 and "Disable and rollback", read on main eb93caffa; consistency model §17.6).

Restore policy (D2-13, PROPOSAL until Chris answers). There is no selective per-project restore of canonical data. A point-in-time restore goes into an isolated database, and recovery into production is manifest-driven forward recovery: records are recreated as new commands with provenance, never overwritten. A whole-database restore writes a history-discontinuity record per affected project (restore point, stamps lost, reason), which manifests and as-of requests report. Writes reopen only after the integrity checker passes. Scheduled state in SQL Server (MassTransit scheduled messages, Quartz) does not rewind with Mongo, so it is reconciled from Mongo state. Commands committed after the restore point are lost; a client retrying one either succeeds against the restored bases or gets StaleBase. FEAT-024 families are rebuilt (consistency model §13.1).

7. PRISMA-specific adoption

PRISMA adoption needs amendments A–O (Q-06a, Q-06b, Q-37). Adopted projects can record steps done outside SyRF, such as deduplication before import, as reported counts (amendment K), and can run ASySD's retroactive deduplication inside SyRF after adoption (amendment L). Legacy projects keep current-state screening counts labelled with their basis. Report snapshots for adopted projects record coverage for import, deduplication, retrieval and pool-entry history that is missing, and use evidence-based lower bounds where possible (a study with a recorded legacy decision certainly entered screening). Deduplication of reviewed records is admin-reviewed and never a side effect of outcome or session migration (amendment D). Amendment G replaces FEAT-011's platform-wide backfills (MIG-11, MIG-12) with this per-project adoption, and amendment I replaces its $unset rollbacks with canonical-aware rollback.

FEAT-011 release checklists mapped to this plan (FEAT-011's "Release ½/3" are unrelated to this plan's R-numbers):

FEAT-011 checklist item Must pass in
Release 1: nullable sourceType and sourceName on SystematicSearch P1
Release 1: no Study or SystematicSearch fields that clash with planned PRISMA names (lifecycleStatus, screeningOutcomes[], duplicateGroupId, publicationId, citations[], fullTextStatus, status, state, type, category) Every release that adds persisted fields, checked at F1a (storage ADR) and in R0 floor steps
Release 2: classification fields don't use lifecycle enum names C1
Release 2: reconciliation never overwrites screeningOutcomes; no single-outcome field on Study; consistent authority patterns; Reconcile covers annotation and screening R3a (outcome shape), R4a and R4p (reconciliation)
Release 3: lifecycle status enum, pool exclusion of Duplicate and Merged P2
Release 3: backfill of lifecycle status on all studies; migration of all screening to screeningOutcomes[] Per-project adoption in R6 (amendment G), not platform-wide
Release 3: screeningOutcomes[] schema; structured exclusion reasons groupable by primary reason R3a and R3b, with the per-profile shape of amendment H and the reason shape of Q-22
Release 3: pmPublication with DOI/PMID indexes; immutable Citations with all raw fields; duplicate count derivable P1 (Citations), P2 (Publications, duplicates)
Release 3: sourceType on new imports; backfill where determinable P1 (new imports, admin classification tool); R6 (inference per project)
Release 3: flow diagram generation; all 34 fields derivable; source columns; PRISMA and dedup exports (EXP-05/06) R5b (and P2 for the dedup export)