Unified annotations, screening and cohort classification: viability and delivery proposal¶
Current DP6 correction: personal Include governs within-stage steps; cross-stage routing is configurable. DP7 confirms Collective Include required as cross-stage default; advanced own-Include option allowed. Strict within-stage details remain proposed. Older blanket cross-stage DP1 wording below is historical/superseded.
Recommendation and scope¶
3 October ownership correction (DP4): screening eligibility questions/configuration belong to their screening profile, distinct from ordinary study facts. Templates copy into profile-local definitions; template edits do not silently update existing profiles. Same-profile stages share profile-scoped answers; importing into separate profiles does not share answers. Reuse question-tree infrastructure without conflating ownership. This supersedes older generic project-asset wording for screening definitions; ordinary study-fact sharing remains unchanged. See the current owner ledger.
Later reconciliation acceptance (RE2): final reconciliation form submission accepts valid displayed answers including matching prefill, with obvious populated indicators and an unseen-control warning offering Complete anyway and source provenance. No per-field confirmation or gold promotion before submission; Save/autosave unfinished. Query child validity and profile screening rules remain unchanged.
Later export/statistics-access decisions: EX1 defaults downloads to current answers and requires previous versions/date-specific review state for reproducibility. AG1 is a separate agreement-statistics view capability without admin control, candidate-answer exposure or identity-blinding bypass. Reuse suitable existing grants and recovered #2461/#2574 groundwork; historical mode implementation is not verified. See owner ledger and read-only investigation.
Later explanation correction (RE1): reconciled annotation-answer explanation is optional even when differing from every candidate; reminders are non-blocking, never a completion/ publication gate. This supersedes older rationale-requiredness recommendations below. EX1 confirms current/historical/as-of downloads; Reconcile is separate from extra-review permission.
Later 2 October workflow/permission decisions: EW1 is stage default Allow with advanced per-step override; BL1 identity blinding is stage-owned. RA1–RA5 keep pool default, permit eligible explicit admin assignment, scope expiry to unstarted explicit work, and require release/reacquire for started work. Additional independent review needs its own capability. See the owner ledger and permission inventory; older placement/expiry/ role assumptions below are historical where they conflict. No runtime changes are authorised.
Later 2 October update/query decisions: VU1–VU3 and QY1–QY7 in the linked owner ledger supersede earlier proposal-only query treatment. Needs-updating answers remain visible; version reason/guidance are optional. Pending accepted gold stays effective, per-version grouped concerns retain individual resolution, affected children are resolved before replacement gold, authorised query self-review is allowed/audited, rejection explanation optional, existing viewers may query, and resolution notifications preserve isolation. Stale-version/concurrent handling remains open. Publish-pause retry/notification UX is only a recommendation. No runtime/contact authorization.
Later 2 October lifecycle correction (SL3): the latest explicit Save OR Complete is the current form-session version. Incomplete Save replaces completed status and no longer counts as completed; autosaved changes alone remain unconfirmed edits over the explicit version. Retain immutable history, separate draft indicators and publication treatment for completed, saved incomplete and draft-only sessions. This supersedes any earlier effective-completed- until-Complete assumption below. It does not silently change gold or screening decisions. Existing admin-confirmed requireReanswer/autoUpdate/doNothing choices are recovered from Annotation Versioning §3.1–3.3, not a new owner decision. Their application to shared saved incomplete/draft work and confirmation UX still need design; see the linked decision ledger. Older added-question “no impact” wording is superseded by FV1–FV3.
2 October owner-decision overlay: the shared-form decision record and updated review-step handoff supersede older per-step submission/target and blanket accepted-gold visibility assumptions. Forms own shared evidence, reviewer sessions and study targets across stages; compatible answers can be shared across overlapping forms. Exact revision provenance, immutable Save/Complete, configurable accepted visibility and independent/informed agreement are confirmed. Publishing checks sessions from any earlier form version and asks the admin how they should qualify; materialized form/question usage must be current at the protected publication point. Exact transition/statistics protocols and query lifecycle remain proposals, not new runtime or statistics-programme authorization.
The combined model is viable as a shared annotation infrastructure with explicit domain capabilities. It is not viable as unrestricted generic questions, metadata and links from which SyRF guesses scientific meaning. Share identity, revisions, typed schemas, owned answers, references, evidence, submission and reconciliation infrastructure. Keep screening decisions, population classification, quantitative observations and their derived results semantically distinct. No single inheritance hierarchy, unrestricted rule language or universal CRUD endpoint should control all those behaviours.
This document expands the screening-as-annotation investigation. That investigation remains the detailed source for screening lifecycle, permissions, allocation, reservations, APIs, statistics, migration and its M0–M8 convergence plan. This companion adds the recovered classification design, metadata and response modes, configurable definitions and references, quantitative-data implications and a combined delivery sequence. It does not approve implementation, migrate data, enable a feature, merge or deploy anything. The separate materialised-project-statistics programme is not resumed or modified.
The 24 September discussion additionally requires outcome-data structure to be configurable at project design, with stage-level binding/configuration: today's fixed time/average/error rows do not cover all review types. Section 4.5 makes this a first-class part of the target, not merely extra metadata attached to an unchanged continuous-outcome table.
Scope correction from the follow-up discussion: SyRF records, reconciles and exports evidence; this proposal does not add statistical analysis, cohort pooling or combined-effect calculation. Relationship inference and count-consistency checks describe recorded evidence. Downstream analysis tools decide how to use it. Uncertain overlap is exported as uncertain; there is no proposed combined-analysis approval gate or acknowledgement requirement in SyRF.
The smallest useful combined-model release is one explicit screening profile, one reviewer decision identity, preserved revisions, optional decision-owned reasons, and an opted-in Complete-and-Include form submission, using the shared annotation contract. It must preserve drafts and the existing allocation/permission boundaries. Classification reasoning and category migration do not block that release. Its critical path is identity/schema compatibility → specialised atomic submit → compatible UI/readers/exports → acceptance and rollback rehearsal. The first classification release is separately useful: record scoped definitions, coverage and explicit relationships, reconcile them and export them accurately. Inference follows that slice.
1. Evidence recovered, and what it does not establish¶
1.1 Source ledger¶
| Source | Verified evidence and status | Use in this proposal |
|---|---|---|
| Local Codex task Workspace Coordinator, discussion and document-creation report on 13 August 2026 | Read through the Codex app's local/archived history. Discussion explicitly covers Farm A/Farm B, pregnancy/female containment, classification expressions, complete/partial/unknown cohort coverage, double counting, provenance and general Definition Types. | Historical requirements and design context, not implementation evidence. Later corrections take precedence over earlier exploratory answers. |
SyRF_Cohort_Classification_V1_Requirements_Specification and SyRF_Cohort_Classification_Feature_Explainer_and_Prototype_Brief |
The task reports producing both in Markdown and Word. Their file contents/current existence were not directly retrieved from Windows in this investigation. | The recovered conversation is evidence; this document is not represented as a byte-for-byte incorporation of the original files. |
| Response modes and metadata context | Draft, captured 10–11 August from the earlier Excel Template and Parser discussion. | Typed question-side field definitions, response-side values, alternative response modes, repeatable groups and metadata boundaries. |
| Extensibility architecture | Document marked Approved; individual delivery items have their own status. Its custom-reference section expressly remains exploratory. | Additive compatibility, AF2-first delivery, import/export and writer gates; not proof that the entire architecture is implemented. |
| Annotation schema/profile proposal | Inspected the local PR document docs/decisions/ADR-016-annotation-schema-profile-boundary.md, dated 30 August, In-Review. |
Reusable schema versus contextual profile, separate value/control/mode/metadata/reference semantics. This is not an adopted ADR merely because it has a number. |
| Earlier validation/category proposal | Draft May scoping spike, identified as superseded by the August proposal. | Motivation to remove hard-coded biomedical categories; do not revive its assumed QM-v2 or ProjectTemplate implementation. |
| Screening investigation | Code and existing test evidence pinned to c30ec96e625d18f642151e2fd2c6ca7e6caa6cd0. |
Current behaviour and detailed screening target. Additional source links below use the same code baseline. |
Historical artifact locator, for a local Windows task to verify without publishing a share link:
Task: Workspace Coordinator
Task ID: 019fcf96-c562-7043-9667-a5ccf68936a1
Document-creation turn: 019ffc69-0286-7b92-ac4c-10153504f86a
Folder: C:/Users/chris/Documents/Codex/2026-08-05/realtime-voice-chat-3/outputs/
Files, each with .md and .docx extensions:
SyRF_Cohort_Classification_V1_Requirements_Specification
SyRF_Cohort_Classification_Feature_Explainer_and_Prototype_Brief
The conversation continued after document creation, including the SEBI comparison and generalisation of Disease Model/Treatment into Definition Types. Those later ideas may not be in the files. Before treating either historical artifact as a normative implementation input, compare it with the corrected discussion and this proposal. No new approval is inferred from an old assistant saying “agreed.” The current user has authorised integrated research only.
1.2 Recovered requirements and corrections¶
The recovered discussion supports these requirements as historical design context:
- Reuse ordinary questions, repeatable branches and annotation lookups to describe locally named entities. A project-defined type selects the schema once; each study-local instance owns its own property answers. “Other” supports unexpected reported concepts without treating similarly named concepts in different papers as identical.
- General Definitions can optionally have Classification capability. Classification adds containment, exclusivity, exhaustiveness and cohort reasoning. An ordinary Treatment link need not imply any set relationship. Outcome retains specialised quantitative behaviour.
- Cohorts reference classification options. Normally derive their relationships from option rules instead of asking reviewers to compare every pair. Allow explicitly reported cohort relationships/compositions where classification evidence is insufficient.
- Corrected inference rule: matching predicates alone do not establish containment in a particular reported cohort. The broader cohort must have complete coverage of its matching population in the applicable scope. Earlier conversation examples omitted this guard; the later complete/partial/unknown correction is essential.
- Exhaustiveness of a set of options, completeness of a particular cohort and completion of an annotation form mean different things. A missing classification is not a negative fact.
- Use a study/experiment context or a reported cohort as scope. Do not invent an empty “all animals” cohort solely to complete a hierarchy. Do not add a second Population entity duplicating Cohort. A free-text “Total” label does not establish full coverage.
- Record reported versus reviewer-interpreted evidence and derived proof chains. Default exports contain explicitly recorded relationships; inferred relationships are opt-in and visibly labelled Inferred by SyRF. Preserve uncertainty and contradictory evidence.
- Separate reported cohort size from N contributing to a particular quantitative result. The later discussion settled on analysed/result N, with a series default and timepoint values, rather than mandating a second assessed-N field. Defaults need provenance and must not overwrite edited values or masquerade as reported facts. Numerator/denominator requirements emerged in the subsequent SEBI comparison.
- Version-one temporal membership reasoning was explicitly deferred. Later crossover and “same members as” examples exposed a future requirement, not proof that it entered V1.
1.3 Additional verified current-code seams¶
| Current implementation | Consequence for the combined proposal |
|---|---|
AnnotationQuestion.Category is a string; AnnotationLookup is a Boolean; system labels, categories and lookup questions are supplied by fixed definitions. Question model · System definitions |
Renaming tabs cannot safely convert categories into configurable domain types. Category presentation, record type and question ownership need separate identities. |
CreateCohortNumberOfAnimalsQuestion creates a system integer question parented to the Cohort label. Cohort size question |
Cohort size already is an annotation answer to a system question. Preserve that source of truth; the proposed Cohort count is a semantic role/read projection, not a second independently editable field. |
| Relationship validation checks OutcomeData experiment/cohort/outcome references against specific label-question GUIDs, and has a fixed source-question → target-question lookup map. Validator · Lookup map | General references need declared target contracts and equivalent server checks; adding a selector alone bypasses today's assumptions. |
Current QuestionOptionDto supplies a string Value, Description and parent filter, not a separate stable semantic option ID. Option DTO |
Preserve legacy value keys, introduce explicit versioned option identity for new types, and distinguish rename from meaning change. Do not assume the older QM-v2 proposal is active. |
OutcomeData carries stage, reviewer, reconciliation, experiment, cohort and outcome identities, with series-wide NumberOfAnimals; TimePoint stores time, average and error. Outcome · Timepoint |
There is no per-timepoint N in these inspected types. Preserve existing fields and old exports before adding result-specific count semantics. |
| AF2 keys outcome cells by experiment/cohort/outcome and copies the enclosing cohort extraction count into submission data. Topology · Submission | The historical concern about copied cohort counts still has a concrete code basis. Do not reinterpret historical copies as independently reported analysed N. |
| AF2 refuses a null cohort count at the non-nullable wire boundary. Persistence | New nullable/missing count semantics require a versioned API/writer contract, not just a UI field. |
| Quantitative export emits experiment/cohort/outcome IDs, cohort size, Disease Model/Treatment control answers and labels, outcome units and values. Export row | These are real compatibility dependencies even where a category is mostly descriptive. A generic exporter must preserve legacy columns and stable mappings. |
| Existing export tests check fixed columns, category order, treatment/model groups and missing-group sentinels; API validation tests check fixed reference types. Export tests · Validation tests | Existing regression evidence covers today's graph, not configurable types, population inference or count correctness. |
No claim that Control has no use anywhere in SyRF is needed for this recommendation. The inspected exporter demonstrably uses it. Generalisation must preserve that meaning and audit all consumers before retiring the legacy fields. The baseline has no implementation evidence here for the proposed cohort set-reasoning engine.
2. A coherent shared model¶
2.1 Shared infrastructure, explicit capabilities¶
Extend the screening investigation's shared annotation envelope, not today's stage/question-bound
Annotation class unchanged. Its initial Answer and ScreeningDecision kinds can be followed by
typed named-record and relationship kinds using the same revision, evidence and submission
contracts. The exact internal C# class layout remains a conformance-tested engineering choice.
| Concept | Owns | Does not imply |
|---|---|---|
AnnotationSchemaVersion |
Immutable question identities, value kinds, option keys, metadata definitions, reference contracts and capability requirements | A project workflow or a screening vote |
AnnotationProfile / schema binding |
Contextual labels, enabled response modes, metadata requirements and permitted reference scopes, bound to a schema version | The ScreeningProfile; these have distinct jobs and IDs |
DefinitionTypeVersion |
Project-defined name, instance question schema, presentation and permitted system capabilities | Study-local factual instances or permission to execute arbitrary code |
| Named annotation record | Study-local Definition or Cohort identity, immutable revisions and owned property answers | Equivalence with another record having the same label |
| Classification capability | Typed option/rule contracts, scoped expressions and admissible population reasoning | Every reference being a subset relationship |
| Cohort capability | Reported group identity, membership definition, scope, coverage and cohort count | Individual-level membership data, full coverage or independence |
| Outcome capability | Measure definition and rules for quantitative observations | Storage of every cohort's result inside one reusable definition |
| ScreeningDecision | Reviewer/study/screening-profile identity, Include/Exclude revision, owned reasons and submission provenance | Collective eligibility, automatic voting from a Boolean answer, or cohort membership |
| Relation assertion | Typed, scoped facts linking named records/options, evidence and author/reconciliation authority | An owning parent/child edge or an automatically inferred truth |
| Outcome series | Cohort + outcome + measurement context, result/timepoint values and result-specific counts | A new animal population each time a result is measured |
| Outcome-data schema capability | Versioned series/observation fields, dimensions, roles and registered validation/export compatibility | Every review reporting time, mean and error, or statistical analysis being performed in SyRF |
| Derived result | Screening agreement, cohort relation proof or count bound, each with its own policy and input revision vector | An independent reviewer assertion, a vote or a reconciled annotation |
The capability set is closed and implemented by SyRF. Administrators configure known capabilities, schemas and labels; they cannot invent new privilege-granting behaviours through metadata. A generic Boolean/choice “decision question” remains an ordinary answer unless an explicit authorised screening binding makes its committed response a ScreeningDecision. That binding invokes the screening command, natural key and permission checks; it cannot bypass them through an ordinary answer endpoint.
Stable named-record IDs identify targets of references. Their property answers use the same question/context/revision mechanics as other annotations. Screening is another specialised annotated record, not a fake cohort or a selection from the classification catalogue. Avoid a single mutable catch-all JSON payload: each kind has a closed schema, validator and command policy. Common services handle revisions, evidence and referential integrity once.
flowchart TD
Schema[Versioned schemas and definition types] --> Binding[Project and stage bindings]
Binding --> Session[Stage review and submission snapshot]
Session --> Answers[Ordinary answers and metadata]
Session --> Decisions[Profile-scoped screening decisions]
Session --> Entities[Study-local definitions and cohorts]
Entities --> Relations[Typed reference and classification assertions]
Entities --> Series[Outcome series and timepoints]
Decisions --> Agreement[Collective screening outcome]
Relations --> Inference[Scoped relationship proofs and count bounds]
Agreement --> Eligibility[Eligibility policy]
Inference --> Export[Export with relationship evidence and uncertainty]
The two derived engines share revision/invalidation mechanics, not scientific rules. A cohort classification conflict does not itself cast an Exclude vote. An eligibility gate may require current resolved evidence, but that is an explicit screening-profile policy.
2.2 Three different graphs¶
- Ownership tree: a parent owns child answers, such as a decision's exclusion reasons or a Definition instance's dose properties. Revision/deletion rules preserve past trees.
- Reference graph: a cohort references a treatment/assay/option defined elsewhere. References do not move, clone or delete their targets. Target retirement preserves old snapshots and prevents inappropriate new selection; hard deletion of referenced history is forbidden. Validate project, study, actor/authority, type and version scope server-side.
- Semantic relationship graph: assertions such as subset, disjoint, same-members or exhaustive composition. Their meaning comes from typed payloads, scope and evidence, not from graphical nesting or ownership. Subset cycles may express equivalence; reject strict-subset cycles as inconsistent. Ownership cycles are always structurally invalid.
A rule can be edited through familiar question controls, but at submission it must compile to a typed validated assertion. A free-text note saying “subset of” remains a note. A container question groups fields without pretending to be a scalar answer or a vote. Reuse layout and tree traversal; do not require fabricated label/value answers just to render a container.
2.3 System annotations for cohort size and quantitative values¶
Yes: retain cohort size as a system annotation, as it already is today. The cohort owns an
answer to a system-defined integer question with the semantic role CohortSize. The domain
reads that answer through the role contract; it does not depend on the displayed wording or
maintain a competing editable count on a separate entity. Existing copies on OutcomeData
remain compatibility projections during migration, with their source meaning retained.
Use the same pattern for the proposed AnalysedN system question, owned by an outcome-series
or timepoint context. These are different roles and owners, not different values of one global
count. Cohort size, analysed N and proportion numerator/denominator may all use typed numeric
answers, but their role validators and calculation rules differ. System ownership supplies the
meaning and invariant; projects can configure compatible presentation and evidence requirements.
For example, cohort C's system size answer is 20; C/body-weight/week-3's analysed-N answer is 18. Both carry ordinary annotation identity, author, revision, evidence, missing-state handling and reconciliation. A calculation resolves the exact role-bound answer revisions and returns its provenance. Derived values are labelled projections, not fabricated reviewer answers. Zero remains a reported number; missing, not reported and not applicable are separate states. Validate count units and nonnegative integral values where the role requires them.
A series default can propose a timepoint answer, but preserve whether it was copied, inherited or confirmed and the source revision. Pin inherited values in submitted snapshots; changing a cohort or default later must not silently rewrite historical timepoints. No second assessed-N field is required by the recovered final requirement. An ambiguous legacy count remains legacy or unknown in meaning until reviewed, rather than being relabelled as a confirmed analysed N.
The same system-role approach can eventually cover quantitative average/error/time values. That is a further storage/API migration, not something today's TimePoint already implements. Whether those typed payloads are stored in one series document or separately is an engineering choice; there must be one canonical logical value with shared revision semantics, not two independently writable scalar and annotation representations. Container/value separation keeps large outcome tables usable without manufacturing extra reviewer work sessions.
2.4 Metadata, alternatives and substantive answers¶
Question-side metadata definitions declare stable keys, labels, types, units/options and requirements. Response-side metadata stores values against that exact definition version. Evidence location, confidence and measurement unit are qualifiers. Cohort coverage, set relationships, screening decisions and eligibility criteria are substantive typed assertions with domain validation, not arbitrary metadata tags.
Ordinary value, reference value and alternative response mode remain distinct. “Not applicable” is a recorded alternative, not an empty value; “not reported” is not a negative classification; hidden by a condition is a derived presentation state. The existing metadata proposal permits profile-controlled metadata-only responses. They must remain distinguishable from missing responses and from complete substantive answers. Requiredness is evaluated for the selected mode/value and its own metadata, never by counting nonempty JSON objects.
Keep the approved architecture's mode-specific reason contract and export/legacy-writer gates. Mode change clears or supersedes the active reason under that contract while retaining the prior revision for history. An owned substantive reason annotation can additionally use evidence metadata; that does not make it a response-mode reason or merge the identities.
2.5 Context, sharing and reconciliation¶
An answer's semantic context includes its subject record, question identity/version compatibility, repeatable-instance path and relevant decision/profile or measurement scope. Stage is provenance and workflow context, not an automatic duplicate answer identity. Two stages can share a study answer only if those semantic dimensions and the permitted authority agree. An answer about Farm A in one study is not the Farm A answer in another. A decision-owned answer keeps its decision/profile context even if its question wording matches an ordinary answer elsewhere.
Reviewer-local entity creation introduces a correspondence problem: two reviewers may assign different IDs to the same reported cohort, or the same label to different cohorts. Reconciliation must explicitly match, split or distinguish instances and record an immutable correspondence map. It must then translate references into reconciled targets, preserving original reviewer graphs. Never merge solely by label, count or array position. Different-stage shared answers reconcile once per compatible context, as specified in the screening investigation.
Compute candidate inferences only within one reviewer's authorised snapshot. Compute collective inferences only from a pinned reconciled snapshot or an explicitly permitted authority policy. Do not combine reviewer A's coverage claim with reviewer B's option relationship before reconciliation. Revealing another reviewer's inferred structure can break blinding even if the underlying answers are hidden. Matching/reconciliation UI and inference visibility need the same permissions as their source facts.
3. Sound cohort reasoning with incomplete evidence¶
3.1 Scope, expressions and coverage¶
Let U be an explicitly identified, compatible study/experiment/cohort scope, and P a
structured predicate over study-local classification options. A recorded cohort C with those
criteria establishes C ⊆ {x in U | P(x)}. Complete coverage additionally establishes
equality. Partial/sample or unknown coverage does not. A scope reference is context, not a
new Population aggregate or an assertion that every study participant belongs to a fabricated
root cohort. Resolve omitted scope only from an explicit enclosing binding; otherwise unknown.
To derive A ⊆ B, require compatible membership scopes, a sound proof that A's conditions
imply B's conditions, and complete coverage by B of those matching members. A need not itself
be complete for this implication. Compatible nested scopes may also suffice if their own
containment is proven; merely sharing a study ID does not.
Example: project-configured evidence in this study states Pregnant ⊆ Female. Both cohorts use
the enrolled-cattle scope. All members of A satisfy Pregnant; B contains all matching females.
Then A is a subset of B. If B is a sample of females, A may contain females omitted by B and
the conclusion is unknown. These option semantics are recorded for the review's scientific
context, never inferred universally from a label or from an assumption equating sex and gender.
Store an expression AST with stable references and a versioned grammar. The first reasoning
slice supports conjunction, explicit containment, disjointness and scoped option-set rules.
Preserve later extension points for any, none, exactly k, at least k, nested expressions
and typed numeric intervals. An unsupported expression can be saved/exported only through a
reader/writer contract that preserves it; the reasoner reports Unsupported rather than replacing
it with a weaker expression. Do not accept arbitrary executable expressions.
none(A,B) means the complement of A ∪ B within U, not missing answers for A and B.
exactly one(A,B) counts satisfied predicates, not numbers of animals. If options overlap,
selecting both is not necessarily a contradiction under an all operator. Numeric interval
proofs require comparable units, explicit endpoints/inclusivity and the same measurement
meaning; age labels or omitted maxima do not supply those facts.
3.2 Separate relationship questions and uncertainty¶
Do not use one enum implying that subset, overlap and equality are mutually exclusive. Return separate propositions with proof state: proven, disproven where supported, unknown, conflicting, unsupported or not evaluated. Equality proves containment both ways, but a nonempty overlap still requires nonemptiness evidence. “Possible overlap” is a caution that disjointness has not been established, not proof of shared animals.
An option-set rule names its members, the scope or enclosing option it covers, and independent exhaustive/mutually-exclusive claims. Farm A and Farm B may partition the same total as Young and Adult. That does not make Farm A and Young disjoint. A direct composition assertion can state that selected cohorts cover a total even when classifications are incomplete, with independent coverage/exclusivity evidence. It remains attributed explicit evidence; reconcile conflicts with inferred relationships rather than silently giving either source priority.
Contradictory facts must not cause arbitrary conclusions. Retain the evidence, identify a conflicting component and withhold authoritative deductions dependent on it. Structural errors (dangling cross-project references, invalid AST, privilege violations) block commit. Scientific contradictions can be recorded and flagged in drafts/submissions; a specific authoritative calculation may refuse them. “Complete review” means the reviewer completed the work, not that the source paper is logically consistent. Policy may require reconciliation before use.
3.3 Counts and double counting¶
The feature can detect redundancy and unsafe sums; it cannot reconstruct missing individuals or joint distributions from aggregate labels and marginals.
| Evidence | Safe result | Unsafe assumption |
|---|---|---|
| A subset B, compatible cohort sizes 20 and 70 | Union of A and B has 70 members | Sum to 90 |
| A and B disjoint, sizes 70 and 50 | Union has 120 members | Conclude they exhaust a separately defined total without coverage evidence |
| A and B sizes 70 and 50, common scope size 100, unknown overlap | Union lies between 70 and 100; intersection between 20 and 50 | Assert exact overlap or use 120 unique participants |
| Age partition totals 120; farm partition totals 120, each exhaustive over the same total | Both describe the same 120 participants | Sum the partitions to 240 |
| Equal labels or equal sample sizes | No identity/membership conclusion | Automatically merge cohorts or count them once |
| Two outcomes on one cohort, analysed N 18 and 16 | Preserve two result-specific Ns | Treat them as 34 distinct animals or infer exactly which animals overlap |
| Shared control observation used by two comparisons | Store one observation ID and two comparison links | Duplicate the observation and count its animals twice |
For multiple overlapping cohorts, pairwise overlaps alone may still leave higher intersections unknown. Compute exact totals only when enough compatible constraints determine them; otherwise return justified bounds or Unknown. Counts must match subject unit (animals/people/farms), scope, time and count basis. A proportion denominator is not automatically cohort size, and proportions are not additive. A subset-size contradiction becomes a flagged conflict, not an automatic data fix.
Expose recorded and inferred relationships, redundancy and possible-overlap explanations in the reviewer interface and export. Downstream consumers can use that evidence to avoid double counting; SyRF does not pool cohorts or calculate a combined effect. Do not alter primary recorded data to “deduplicate” it. Reviewer vote counts, distinct study counts, unique subject counts and quantitative observation counts remain different measures. Optional count bounds are descriptive data-quality results, not a pooled-analysis result. This is a semantic integration contract only; no materialised-statistics task is changed by this research.
3.4 Efficient calculation and proof history¶
Start with bounded per-study/per-authority graphs, adjacency indexes and cached containment closure. Validate option rules once per immutable revision. Request pairwise cohort results on demand; do not eagerly materialise every possible cohort intersection. Pairwise display is O(n²) in the number of cohorts even before expression reasoning, so page queries and set budgets.
General Boolean/cardinality implication can become expensive. For a later bounded solver,
checking A AND NOT B for impossibility is a possible proof strategy only after scope,
coverage and consistency checks. Exhausted budgets and unsupported grammar return Unknown or
Not evaluated, not a scientific conclusion. Select a solver only after fixture-driven benchmarks;
this proposal does not assume a library, unrestricted ontology engine or production scale result.
A derived view is keyed by project/study, authority, schema/rule versions, input annotation revision vector and engine version. It stores the conclusion, reason codes and proof dependencies. An edit invalidates affected results through a reverse dependency index. Read-your-write preview can operate on the current draft but must be labelled provisional. Committed results use exact snapshots; a stale calculation cannot overwrite a new one. On failure, source facts remain accessible, while consumers see Stale/Unknown instead of silently using yesterday's proof.
This is a proposed relationship cache, not an instruction to extend the separate project statistics implementation. Design its event contract now; negotiate implementation ownership when that phase is authorised.
4. Screening, extraction and classification working together¶
4.1 Full-text review with metadata and reasons¶
A reviewer opens an eligible full-text stage and records ordinary study answers, typed evidence metadata and reusable Treatment definitions. Draft saves create no Include vote. Explicit Complete-and-Include validates the stage form and profile evidence, commits the submission snapshot and upserts the one effective reviewer/study/profile decision with preserved history. Explicit Exclude opens decision-owned reason questions; draft exclusion notes alone do not vote. A study filtered out before assessment remains unassessed for that profile.
A profile shared by a risk-of-bias stage and an extraction stage still has one vote per reviewer. The second stage references the effective decision revision or performs an explicit permitted correction; submitting it does not create a second vote. If classification capability is not enabled, those stages can still use Treatment and other Definitions as ordinary references. Work assignment and reservations remain stage-specific for annotation and profile-specific for screening. Neither a reused decision nor an inferred cohort relation creates a fake session.
4.2 Farm, age and reproductive-state example¶
A paper reports 120 cattle. Farm A has 70 and Farm B has 50; a scoped explicit rule says the farms are disjoint and exhaustive. Age groups form another complete partition. The reviewer records each local farm definition once and references it from cohorts. Repeated sex/age questions can share a compatible question schema, but answers remain owned by the particular local instance.
The Female cohort is recorded as complete within the same 120 cattle, with N=70; Pregnant is contained in Female under the paper's applicable classification rule, with cohort N=20. Their union has 70 animals. Replacing Female's coverage with Partial invalidates that subset proof; saved annotations and old proof snapshots remain. A cohort described by Farm A AND Juvenile can be placed beneath a complete Farm A cohort, but no Farm A/sex intersection count is invented. Evidence, including coverage assertions, participates in reconciliation.
4.3 Diabetes sub-study¶
The parent project and diabetes sub-study use different ScreeningProfiles because their eligibility criteria differ. Include in the parent does not imply Include in the diabetes profile. The same reviewer may legitimately have one distinct decision per profile. An ordinary shared answer about whether a paper reports diabetic participants can support both assessments without becoming two copies; the decisions remain independently attributed.
A Diabetes classification describes subjects within a cohort. It is not itself a study-level Include. A study containing both diabetic and non-diabetic groups may qualify for a diabetes sub-study, but which cohorts/results enter analysis is a separate explicit selection policy. Conversely, a missing Diabetes answer means unknown, not non-diabetic, excluded or not applicable. Screening authority does not silently rewrite population membership or discard excluded work.
4.4 Outcome counts, shared controls and crossover¶
An Outcome Definition identifies body weight and its measurement semantics. One cohort's series has analysed N=20 at week 1 and N=18 at week 3; cohort size remains 20. A series default can suggest 20, but export distinguishes inherited/unconfirmed from reported/confirmed N. Counts used for a proportion have distinct numerator/denominator semantics where enabled.
Two comparisons can reference the same control series. A future normalised observation identity therefore separates observation from comparison grouping. Do not remove today's Experiment key during the category rename: current topology, validation and exports depend on it. Introduce an observation ID plus explicit comparison links, retain legacy triple aliases, and require an evidence-based mapping of existing records; similar values are not proof of a duplicate measurement.
Treatment A-period and Treatment B-period can be distinct cohort contexts involving the same animals. A reported same-members relationship can be retained as explicit evidence, but temporal membership inference remains a later capability. Different treatments do not prove disjoint animals. Sequence belongs to use/context, not the reusable Treatment B definition. A disabled temporal reasoner must not apply timeless classification rules across incompatible periods.
4.5 Configurable outcome-data structures¶
OC1 owner update, 3 October (supersedes catalogue-only scope of 27 September): first release includes legacy-compatible and event-count supplied schemas plus project schema creation/customization. Retain project-owned versioned schema definitions, availability and question bindings; stages select validated subsets/versions. Supplied system definitions stay system-owned, while customization is represented in project-owned versioned schemas, not structural edits hidden in stage bindings.
The dedicated schema selection question can reference enabled supplied or project-owned schema versions. Validate its reference and permission, not an annotation-instance lookup. Existing results remain pinned when versions change. Numerator/denominator, range, confidence interval, standard error and variation are requested field examples, not a mandatory single fixed schema; variation has no approved variance/SD meaning. Exact field/event-count specifications remain engineering/domain work; P9 automatic migration/adoption is not settled by legacy compatibility. The project-authoring design below is first-release capability; advanced semantics and optional breadth remain phased rather than deferring all customization.
Requirement: projects, and their authorised stage designs, must be able to configure what
outcome data is collected. Adding optional fields to TimePoint(time, average, error) alone
does not meet it. A non-continuous review must not fill fake time, average or error values to
fit the old shape. Time itself is an optional measurement dimension, not a compulsory property
of every scientific observation.
Use an OutcomeDataSchemaVersion, implemented as an outcome-capability projection of the same immutable annotation-schema machinery, not a second incompatible schema registry. It defines:
- series-level questions and optional defaults;
- repeatable observation/row structure, stable field/question identities, typed values and reference targets;
- dimensions such as time, assay, anatomical site or measurement condition, with their units and reference semantics where applicable;
- system semantic roles such as analysed N, mean, SD, event count, denominator, person-time or reported estimate, plus ordinary project-specific questions;
- requiredness, allowed response modes, evidence metadata and conditional structure;
- permitted registered validators, semantic role bindings and export formats;
- table/form presentation, column order and labels, separately from scientific identity.
Clarified domain requirement (owner discussion, 24 September): outcome-data schemas are project-owned domain definitions, distinct from annotation answers. Administrators add the actual schema definitions at project level. Reviewers select existing definitions; they do not create those definitions while extracting a paper. Reuse the annotation-schema machinery for their questions, rather than maintaining a second copy of the same question structure.
Clarification from owner discussion, 27 September: use a dedicated system outcome-data-schema selection question. This is distinct domain/UI behaviour, not an ordinary administrator-authored annotation lookup. The system controls its question text, purpose, validation and presentation. Its allowed options are generated from the schema versions configured as available in the project; they are not a separately editable list of answer labels. Published question/set versions pin the allowed references, so editing the project catalogue does not silently change an existing submitted form.
Its answer is an annotation whose typed value references a project-owned schema version. An ordinary treatment/disease-model lookup targets a study annotation branch rooted in a label annotation; this question targets a configuration definition, not a reviewer-created entity branch. Schema references require their own resolution and validation for target kind, project ownership, allowed version and historical readability. Never interpret a schema ID as an annotation ID or manufacture a label annotation to represent the schema.
Reuse common annotation identity, ownership, revisions, draft/save/submit and history machinery, and UI primitives where useful. That does not mean the existing lookup data model already supports this target unchanged. The precise C# class/discriminator and physical storage design remain engineering choices; a general-purpose reference framework is not a prerequisite for implementing this distinct system question.
For newly configured extraction, keep the question visible in the administrator's actual question tree. With one available schema, show reviewers a fixed schema label and its applicable metadata child questions; with several, show a selector and the selected branch. An automatically supplied single-option answer records configuration provenance, not a deliberate reviewer choice. Branch answers belong to the reviewer's outcome context, not to the project schema definition. Common outcome questions remain outside schema-specific branches. The generated/configured questions have normal versioned identities; do not maintain competing copies of a schema's question definitions and generated form structure.
The project defines the ordinary study-fact question library and available outcome-data structures; screening eligibility questions/configuration instead belong to their profile (DP4). A stage selects a subset of those questions, preserving the dependencies needed to interpret selected structures. The reviewer creates the outcomes actually reported in each paper and answers their outcome questions. In Experiments, those outcomes are assigned to cohorts and their result series use the selected structure to collect observations. Project/stage configuration does not enumerate the paper-specific measures reviewers will encounter. Earlier illustrative examples requiring administrators to predefine Mortality or Body Weight as outcome types must not be read as a requirement or prerequisite for extraction.
A stage/review step selects a permitted subset of the project-defined question structure, with required ancestors and schema metadata dependencies validated. It does not independently author outcome branches, schema options or structural overrides. New structures and structural or semantic changes are authored and versioned at project level, then explicitly adopted by question sets and their stage bindings. A shared observation can be reused/reconciled across stages only when its schema, role meanings and measurement context are compatible. Hiding an existing field in a stage view does not delete it or overwrite it with missing.
Low-friction authoring recommendation: project-level outcome-extraction setup selects schemas and prepares their required metadata questions at the appropriate outcome, cohort, series or observation scope. Reuse existing compatible questions by explicit identity and meaning, not text similarity; prepare missing questions and required ancestors for designer review before publication. Stages assign subsets of this resulting project structure. Outcome measures themselves are still created by reviewers from the paper.
Moving from one schema to several is a published form-configuration change. In the proposed new-project structure the system selector already exists; a new option/branch need not reparent its existing children. Use the earlier Question Management version-transition flow to adopt question/set versions and decide whether prior submissions remain pinned or receive an update draft. Never mutate submitted trees or mark missing required metadata satisfied. Conditional visibility and parent ownership remain distinct concepts, but non-owning conditions are not an excuse for showing administrators a false hierarchy. Existing legacy trees may require a deliberate redesign rather than an in-place parent change.
An Outcome series pins its schema version, cohort, outcome definition and measurement context. Each observation has stable identity and owns ordinary/system annotation answers to the schema's questions. Time, mean, error and N are role-bound answers when that schema uses them. The series and row records organise those answers; they do not maintain a second independent copy of their values. Rows may be efficiently stored together physically, provided logical answer identity/revisions and atomic submission remain consistent with the common infrastructure.
Linked outcome annotations: units and average/error types remain answers about the reviewer's outcome, not fixed labels embedded in a project result schema. The existing system questions define outcome units and average type; the error question's parent and option filter depend on the system-question version. System question definitions The result schema declares references to the relevant outcome-answer roles. A continuous observation's numeric fields use those linked answers for their labels and interpretation. Submission pins the relevant revisions; changes to units or statistic/error type mark affected results for review, preserving their prior interpretation and never silently converting values. Question-tree parentage, conditional display and ownership of an answer are separate: the numeric answers belong to observations within an outcome–cohort series, not directly to the outcome merely because its selection question activates that branch.
The schema defines the fields to collect, not the numeric values. Outcome-level answers provide field meanings such as unit, average type and error type. For example, “grams”, “mean” and “SD” label and interpret a result whose time/average/error values are entered separately from the paper in Experiments. Cohort size remains distinct from result-specific number measured/analysed. Metadata that varies by series or observation belongs at that scope rather than being forced into one outcome-wide answer.
System-recognised semantic roles enable specific validation, labelling and export behaviour; custom fields/metadata can still be collected without special system interpretation. The role catalogue and extension mechanism remain unresolved. One selected schema per outcome is the initial design boundary. Multiple representations under one outcome, versus separate outcome measures, is deferred pending concrete use cases; repeated schema selection is not an established requirement.
Changing an outcome's selection before results exist requires answering the new branch's required metadata. After results exist, a selector change cannot reinterpret existing numbers. Preserve the original series and schema/metadata references, identify every affected series, and require explicit replacement data or a separately specified validated conversion before submitting a consistent replacement. A mean/error pair must never turn into events/denominator by field position. Old branches and results remain in history; generic automatic conversion is not part of the initial scope.
Concrete storage candidate, not a final collection decision: store published versions in a
shared MongoDB AnnotationSchemaVersions collection. Each document has an ID, project ID,
stable schema ID, version, an outcome-data discriminator, and its question/group definitions
or immutable question-version references. Stage configuration, schema-selection answers and
result series reference the version ID. This describes a separate configuration entity, not a
reviewer annotation or a structure recreated within every study. Embedded versus referenced
question definitions, draft publication, indexes and transactional boundaries still require the
U0 storage spike; no migration or physical schema change is authorised by this discussion.
| Project-designed schema | Example collection | Validation and exported meaning |
|---|---|---|
| Continuous summary (legacy-compatible built-in) | Optional time; reported mean or median; declared dispersion; result N; unit | Existing extraction can project to its current shape where representable; preserve declared statistic and error roles |
| Event count / proportion | Events; denominator; optional time, assay and evidence | No mean/error field required. A binomial proportion capability may check events ≤ denominator when their defined meanings justify it |
| Incidence rate | Events; person/animal-time and unit; follow-up context | Events may exceed number of subjects; never apply the binomial constraint merely because both fields are numbers |
| Reported estimate | Reported estimate; estimate type; confidence bounds/level; comparison reference | Preserve what the source reported and its meaning; SyRF does not calculate a combined estimate |
| Custom outcome | Project-defined typed numeric/categorical/reference fields and repeatable records | Collect, reconcile and export the declared structure and semantics for downstream consumers |
Administrators configure known roles and compose supported schemas. A role contract validates owner, type, unit, cardinality, missing-state policy and necessary companion roles; a field's label alone cannot assign those semantics. System-owned validators and export adapters consume these contracts. Custom structures remain valid collected evidence when their schema allows them, even without a specialised export adapter. The normalised export carries their schema manifest for downstream interpretation. No statistical-calculation plugin architecture is required by this proposal. A derived relationship or count check retains its input revisions and is never presented as a reviewer-transcribed result.
Schema design and repeated table entry should share ordinary question controls, metadata, draft/save/complete rules and evidence handling. Reconciliation matches corresponding series and observation rows explicitly before comparing their fields. Time value, row index and equal measurements are not safe row identities: a paper can report several assays at one time, and two reviewers can enter those rows in different orders. Default propagation must preserve edited answers and pin the source version in submissions.
Introduce a versioned observation API/payload with schema identity and typed answers, not a widened legacy DTO that accepts arbitrary unvalidated fields. Imports preflight the same schema, roles, references and writer capabilities. Exports offer normalised series/observations/answers with a schema manifest, and schema-specific flat views. A legacy continuous export is available only for a representable compatible schema; otherwise return a clear unsupported-format result. Never flatten categorical/reference values into zero or silently drop non-legacy columns.
Migration initially supplies a legacy continuous schema preserving every current triple, timepoint and value. Mint a durable row ID for each legacy row without pretending the historic row had cross-reviewer identity. Preserve raw values/order and record the migration alias. Keep legacy NumberOfAnimals semantics; do not label it confirmed analysed N. Existing projects remain bound to the legacy schema until explicitly migrated. The first alternate-schema pilot should exercise a genuinely different shape, such as events/denominator without mandatory time/average/error, all the way through authoring, review, reconciliation and export.
27 September migration refinement: existing extraction-enabled projects need not receive a new selector question or synthetic selection answers merely to remain readable. Record an explicit legacy compatibility binding to the fixed schema and existing metadata mappings, preserving the actual question hierarchy, stored type fields, answer values and result values. Legacy compatibility is not a selectable authoring mode for newly created stages. Absence of a selector alone must not classify a record as legacy; provenance/configuration must establish that fact. Upgrading an existing project to selectable schemas is a separate administrator-led, versioned redesign. The earlier suggested automatic selector-answer backfill is superseded. Do not automatically apply the new-project conditional branches to legacy submitted answers.
Current repository export inspection on 27 September confirmed fixed columns: OutcomeUnit
reads the outcome units annotation, OutcomeAverageType and ErrorType read stored outcome-data
fields, and NumberOfAnimals reads the cohort annotation, not the similarly named outcome-data
field. TimeInMinute is a fixed header; time/average/error values come from each timepoint.
Preserve these distinct legacy sources rather than claiming all export metadata is already
read directly from outcome annotations. This was repository inspection, not production access.
Removing Experiment from observation identity and advanced temporal membership remain independently sequenced work. They are not prerequisites to prove configurable outcome data. Statistical analysis is outside this programme, rather than an assumed later implementation phase.
5. Implementation boundary and impact map¶
| Area | Required integration | Safety/compatibility boundary |
|---|---|---|
| Backend domain | Shared annotation envelope/revisions; typed capability policies; schema and definition-type versions; scoped reference validation; cohort expressions/coverage; correspondence-aware reconciliation | Never allow generic answer writes to bypass screening/authority checks or reinterpret scientific uncertainty as validation success |
| Storage | Stable logical IDs plus immutable revisions; typed owned trees and reference edges; relation assertions/proof dependencies; explicit legacy aliases | Additive readers first. Preserve embedded annotations, BSON discriminators/GUID conventions and original values. No destructive generic collection rewrite |
| API | Typed commands for screening, named-record/property updates, rule assertions and reconciliation; expected revisions and submission manifests; typed relationship query results | Server resolves authorised targets and whole change-set references atomically. Imports use the same policies. Define bounded batch/idempotency contracts |
| Frontend | AF2 schema-driven definition/instance forms; metadata/mode renderer; custom reference picker; coverage controls; relationship/evidence inspection; resolved entity matching | Existing AF2 is on the current backend. Do not make its live UI depend on dormant QM-v2. Prototype role/type UI without changing storage meaning |
| Questions/configuration | Stable option keys; type-specific schemas; shared Definitions tab versus dedicated tabs; ordinary and container questions | A tab is presentation, not identity. Semantic type/option changes require a new version and reviewed mapping; wording edits preserve keys |
| Permissions | Existing project/stage design and review/reconcile permissions extended by capability; access to source and target; blind candidate graphs | Schema design is not authority to adjudicate reviews. Hidden facts must not leak through inferred labels, counts, pickers or events |
| Allocation/reservation | Screening profile claims independent of stage annotation claims; edits/corrections retain saved-work access | Relation recomputation is background work, never a reviewer allocation or an incomplete annotation session |
| SignalR | Revision-bearing annotation/schema/decision/rule invalidation; authorised target groups; stale derived-view notifications | Duplicates/out-of-order events cannot regress a view. Durable reload after reconnect; never broadcast another reviewer's evidence to blinded peers |
| Queries/statistics | Separate vote, session, study, entity, observation and subject-count metrics; freshness/provenance returned with derived relations | Old stage totals remain comparable. No fake completion or Include counts; no summing overlapping groups under a “total animals” label |
| Import | Two-pass stable ID resolution for named records then edges; validate schema/context/capabilities/authority; preview collisions and legacy mappings | No inference from labels, empty fields or completed sessions. Unresolved targets fail or quarantine explicitly; preserve source rows |
| Export | Versioned normalised tables plus convenience flat views; stable IDs/revisions; metadata and response modes; explicit relations by default, inferred opt-in with proof | Keep current columns until consumer migration. Repeating a referenced observation in a flat view does not make it a new observation |
| Quantitative paths | Outcome model/DTO, TimePoint replacement/projection, versioned outcome-data schema binding, AF2 topology/submission/persistence, schema-driven matrix and export writers | Collect alternate shapes without fake mean/time/error values; sequence observation/comparison identity separately; validators/exporters use declared roles |
A possible target physical split is definitions/configuration, annotation heads/revisions, submission manifests and rebuildable derived views. Use the screening investigation's commit and concurrency contract, not a second independent revision system. Bound transactions; a submission manifest must reference a complete committed snapshot before it is visible. Never publish only half a rule/target rename or half a combined decision/form submission. Specific Mongo collections/indexes remain an M0 design spike, not an authorised schema change.
6. Migration: what is known and what must remain unknown¶
| Legacy evidence | Safe additive mapping | Not inferable |
|---|---|---|
| Fixed Disease Model/Treatment labels and owned questions | Preserve annotation IDs, original question GUIDs, stage/reviewer context and label/property revisions; alias to corresponding Definition Types | Classification capability, set semantics or cross-paper concept identity |
| Cohort label/count and existing lookup answers | Cohort identity, recorded properties and explicitly existing target links | Complete coverage, disjointness, exhaustive options, overlap, common members or unique subject count |
| Experiment grouping and OutcomeData triples | Preserve grouping, triple aliases, source IDs and all quantitative values | Same observation across experiments, statistical independence or temporal membership |
| Series NumberOfAnimals populated from cohort | Preserve as legacy count with source semantics | Independently reported analysed N at each timepoint, numerator or denominator |
| Existing scalar/array option values | Preserve original value keys and explicit schema-scoped aliases | Label equality as stable semantic identity across questions/projects |
| Screening rows | Map explicit reviewer decisions into a legacy compatibility profile as assessed in the screening investigation | Unrecorded historical revisions, votes from completed forms or auto-Include from eligibility |
| Notes, missing answers, excluded/hidden work | Retain text and original state without semantic promotion | Typed metadata meaning, negative predicates, not-applicable status or permission to delete work |
Backfill should generate a reviewable report of source IDs, mappings, unmapped fields and ambiguous cases. Duplicate-looking records stay separate pending reconciliation. Administrator mapping of a legacy screening question is an explicit, auditable scientific attribution, not an automated heuristic. Existing annotation progress survives every migration and flag change.
Before any future migration, inventory real data only under separate authorisation. This research uses source code, historical planning and synthetic examples; it does not estimate production prevalence. Round-trip fixture tests and a dry-run mapping report precede any data-writing phase. Rollback retains new records and submitted snapshots rather than deleting them to make an older application start.
7. Phased programme, validation and rollback¶
The sequence below extends, rather than replaces, screening milestones M0–M8. Each release has its own independently usable outcome and no requirement to finish all future capabilities first.
| Phase | Usable outcome and acceptance boundary | Dependencies / rollback |
|---|---|---|
| U0 — integrated contracts | Reviewable schemas, command/read examples and synthetic mapping/solver fixtures; compare recovered Windows files with history when available; choose inheritance versus composed payload via conformance tests | Research/prototype only. No rollout. Establish performance/error budgets and reader/writer capability negotiation |
| U1 — screening vertical slice | One explicit profile, one effective reviewer vote, preserved revisions/reasons, safe Complete-and-Include and explicit Exclude; old/new readers agree, allocation and permissions preserved | Screening M0–M3/compatibility gates. Disable new admission/writes, retain readable submitted data and old projection; never revert to a writer that drops new fields |
| U2 — metadata and custom references | One existing kind of question supports typed qualifiers and a declared reference target end-to-end: author → review → reconcile → import/export | Compatible schema IDs and U0 contracts; can proceed independently of screening UI once shared contracts are fixed. Export/readers before writers; stage writer floor after first extended response |
| U3 — definitions and explicit classification | One opted-in project creates a custom Definition Type, study-local instances, cohort predicates/coverage and explicit scoped rules; counterpart matching and export work | U2 references and immutable context. Initially no automatic inference; existing categories stay usable. This is the minimum useful classification release |
| U4 — conservative inference | Explainable containment/disjointness and safe count/bounds checks with provenance; unknown/conflict/stale states and revision invalidation; synthetic farm and pregnancy cases pass | U3. Feature can be disabled leaving explicit records fully readable/exportable. No automatic votes or analysis data deletion |
| U5 — legacy category adoption | Disease Model/Treatment become compatibility-mapped Definition Types, tab configuration independent; old exports and question GUID aliases retained | U3 and full consumer inventory. Per-project opt-in; no classification semantics backfilled. Rollback uses compatibility renderer/projection, not a destructive reverse migration |
| U6a — configurable outcome data | Legacy continuous schema plus one alternate schema, project design and pinned stage binding, system-annotation result fields/counts, schema-driven review/reconciliation and lossless export | U2/U3; versioned API and reader/writer gates. No fake time/average/error. Preserve legacy records and limit old export to representable schemas |
| U6b — observation/comparison identity | Shared-control reuse, explicit comparison links and evidence-reviewed legacy observation mapping | U6a's schema foundation; independently gated consumer migration. Preserve ambiguous old observations; no equality inferred from values or labels |
| U7 — advanced expressions and contexts | Bounded nested/cardinality/numeric reasoning; optional temporal/same-member reasoning and richer relationship export | Measured U4 use and separately approved scientific scope. Unsupported/budget-exceeded results remain unknown; no blocker for earlier useful releases |
| U8 — convergence and retirement | Remove superseded adapters only after migrated consumer parity and authorised stability gates; one canonical annotation/revision engine remains | Screening M7–M8 plus adoption evidence. Rollback floor documented; do not promise indefinite old-binary compatibility after new semantics exist |
Feature flags: this documentation needs no runtime flag. Proposed implementation needs separate capabilities for extended response writes, configurable record types, classification assertions, inference reads, configurable outcome-schema writes and quantitative identity/count changes, alongside screening's flags. Names are illustrative, not additions to the generated catalogue. Server capability grants and schema/version prerequisites must be authoritative; client-only flags are insufficient. Per-project activation requires support to be verified rather than assumed from an existing Boolean flag.
Disabling a flag stops new writes/admission or inference use; it must not hide saved work or make current snapshots unreadable. Once an extended response exists, route to a compatible renderer, possibly read-only during rollback. A kill switch for inference is easy because proofs are rebuildable; rolling back the schema writer is harder and requires a compatible minimum reader. Never silently flatten references, metadata, coverage or modes into text/blank values.
Acceptance gates supplement the screening investigation's existing matrix:
| Test | Required outcome |
|---|---|
| Same reviewer, same profile, two stages | One vote; two legitimate work submissions; shared answers only where semantic context matches |
| Draft/reopen/Complete/Exclude retries | No draft votes; atomic validated submit; idempotent effects; preserved history and saved work |
| Metadata-only, Not reported, Not applicable, hidden, missing | Distinct persistence, validation and export; none implies Exclude or a negative classification |
| Rename type/option and retire referenced instance | IDs/references/history survive; incompatible meaning change is versioned; invalid new selection blocked |
| Reference to wrong type/study/project or blinded reviewer graph | Server rejects or withholds access; no information leaked in search, proofs or SignalR |
| Reviewer-local IDs with identical labels | Explicit correspondence required; split/match reference translation retains both input histories |
| Pregnant implies Female, B complete then partial | Subset proof valid only with sufficient scope/coverage; edit invalidates current proof without erasing the old snapshot |
| Farm and age partitions of same total | Separate valid exhaustiveness claims; cross-axis overlap unknown unless proved; no 240-animal total from two 120 partitions |
| No classification / Other / none-of | Unknown / locally named concept / explicit scoped complement remain distinct |
| Conflicting disjointness, containment or size evidence | Evidence saved with conflict; dependent authoritative conclusions withheld; no arbitrary logical explosion |
| Equal labels/counts in different studies or periods | No inferred identity or membership relationship |
| N=20 cohort, result N=18, other outcome N=16 | Preserve each basis and provenance; no sum to unique participants, no automatic new cohorts |
| CohortSize and AnalysedN system answers | Separate owner contexts and role bindings; calculation pins answer revisions; compatibility fields cannot override canonical annotations |
| Events/denominator schema without time/mean/error | Author, enter, save, reconcile, import and export the alternate shape with no dummy legacy values |
| Reviewer records a measure not anticipated by the project designer | Create a study-local outcome, answer the stage-selected questions, select an existing project schema through the reference question, and collect cohort-linked data in Experiments without creating a project outcome type |
| Schema reference targets another project or an unavailable version | Reject the invalid reference; keep existing saved answers and pinned versions readable under the applicable permissions |
| Linked outcome units or average/error answers change | Preserve the revisions used by submitted results; mark affected context for review, without silently relabelling or converting historical numeric values |
| Same numeric fields with different statistical roles | Binomial and rate validators differ; exports preserve their meanings; field labels cannot select semantics |
| Stage binding changes schema or hides a shared field | New semantic version/context where needed; old submissions readable; hidden fields and other-stage answers preserved |
| Same-time observations entered in different orders | Stable row identity and explicit reconciliation correspondence; no index/time-only matching |
| Custom schema requested through legacy continuous export | Typed unsupported-format result unless exactly representable; no silent data loss or coercion |
| Same control used in two comparisons | One verified observation reference, two uses; existing ambiguous legacy duplicates retained for review |
| Rule edit races with inference/reconcile/export | Expected-revision conflict or new exact snapshot; no mixed-version proof or partially translated graph |
| Reconnect, duplicate or out-of-order events | Reload current revision; no stale proof restored or lost saved draft |
| Import/export round trip of all enabled kinds | Stable IDs, units, missing states, authority and explicit/derived distinction preserved; unsupported writer fails closed |
| Time/solver-budget unsupported | Explicit Unknown/Unsupported/Not evaluated; no false subset or exact-count claim |
| Flag off and old writer retries after extended data | Saved work remains accessible; unsafe writer blocked by capability floor, not allowed to strip fields |
Validation requires fixture-level domain tests, API permission/concurrency tests, Mongo round-trip and transaction tests in isolated containers, AF2 browser journeys and independent export consumer fixtures. Characterisation of current code is not acceptance of the new design. Measure worst-case graph/expression size, proof invalidation fan-out and payload size before enabling inference broadly; no performance claim follows from small examples alone.
8. Feasibility verdict and unresolved decisions¶
| Design choice | Assessment |
|---|---|
| Extend current Annotation subclasses and add more GUID special cases | Low initial code volume, poor fit for profile-scoped votes, independent entity identity and evidence revision; not the target |
| One new inheritance tree for every semantic concept | Possible shared mechanics, but subtype combinations and serializers become coupled; does not itself solve identity, scope or authority |
| Shared annotation envelope + typed payloads + composed capabilities | Recommended. Reuses common infrastructure while preserving independently testable scientific/permission rules; allows incremental migration |
| Fully generic question/metadata graph with inferred semantics | Reject. Cannot safely derive eligibility, completeness, counts or reference authority from arbitrary labels and fields |
| Separate permanently bespoke engines for screening, cohorts and metadata | Useful transitional adapters, but repeats history/reference/reconciliation code and obstructs cross-stage sharing; converge on common infrastructure |
The principal risks are semantic migration, entity correspondence across independent reviewers, legacy writer/export compatibility, and reasoning over incomplete evidence. None is solved by renaming categories. The combined approach is feasible if releases preserve those boundaries and unknown states. A proof of production-scale performance and a complete Windows artifact comparison are still outstanding evidence, not reasons to withhold the present design proposal.
Genuinely unresolved product decisions, in addition to the screening investigation's five:
- First non-screening pilot: which concrete project/type should exercise U2/U3? The SEBI examples are useful fixtures, but this investigation does not select or activate a live project.
- Authority for shared scientific rules: who may publish project defaults and who confirms their applicability to each paper? Recommend project-owned schema/defaults plus explicitly scoped study evidence and normal reconciliation, not universal truths from option labels.
- First classification expression scope: whether the first released reasoner needs operators beyond conjunction/containment/disjointness/exhaustiveness. Recommend the bounded subset above, preserving the historical broader expression model for ordered follow-up work.
- Quantitative transition priority and pilot schema: which alternate result shape should exercise configurable outcome data first, and when to decouple comparison grouping from observation identity. Configurability itself is now a user requirement, not an open decision. Preserve today's Experiment contract until a separate scientific/consumer migration is accepted. Temporal inference stays deferred by default.
Not reopened: the need to distinguish complete coverage from form completion, no label-based identity, no fabricated historical Include votes, one effective vote per reviewer/study/profile, preservation of saved work, inference provenance, or the separation of stage allocations. The proposed combined-analysis export gate was withdrawn following the user's clarification: SyRF exports recorded evidence and relationship uncertainty; downstream analysis is outside its scope. Technical naming, cache layout and solver choice are engineering validation work, not reasons to ask the user to redesign the model before research can finish.
9. Validation of this investigation¶
On Juniper, 24 September, in the isolated documentation worktree:
- 34 existing quantitative-export tests passed (row and header writers; no failures/skips).
- 28 existing AnnotationRelationshipValidator tests passed (no failures/skips).
- A temporary Python set-enumeration experiment over a four-element universe checked 256
combinations of
Pregnant ⊆ Femaleand cohort subsets with complete Female coverage. All satisfied the guarded containment rule. Enumerating proper partial Female cohorts produced 1,105 counterexamples to the unguarded rule. All 256 pairs of subsets satisfied the stated union/intersection bounds. Two orthogonal exhaustive partitions demonstrated that summing their counts gives eight while their common universe contains four.
The experiment enumerates possible sets as a research check; it is not a proposed individual- animal storage model, an implemented reasoner, a production benchmark or a proof of every future expression operator. No production data was accessed. Existing tests characterise the legacy contract, including its limitations; they do not certify the proposed capabilities.
dotnet test src/libs/project-management/SyRF.ProjectManagement.Core.Tests/SyRF.ProjectManagement.Core.Tests.csproj --no-restore --filter 'FullyQualifiedName~OutcomeDataFormatRowWriterTests|FullyQualifiedName~OutcomeDataFormatHeaderWriterTests' --verbosity quiet
dotnet test src/services/api/SyRF.API.Endpoint.Tests/SyRF.API.Endpoint.Tests.csproj --no-build --filter 'FullyQualifiedName~AnnotationRelationshipValidatorTests' --verbosity quiet
The API binary had already been built for the initial investigation and no runtime source changed in this documentation worktree. Existing build warnings were emitted by the Core run. The baseline/source-range check validates paths and line bounds; the code was also read at the cited seams. Documentation validation and final whitespace checks are required before commit.
Source references¶
References identify the inspected code baseline, not deployment state. Current-test execution and the small reasoning experiment are recorded with the parent investigation's validation.
Latest confirmed direction and concrete approval proposals — 3 October 2026¶
This section supersedes earlier deferred/open wording for P9–P14 where stated; original source prototype and historical discussion are retained, not rewritten as implemented behavior.
- LC1 / P11: automatic completion requires no unresolved applicable work, including drafts and outstanding corrections. Alert project admin and obtain confirmation before admitting a change that would reopen a Completed stage. Commit approved change then auto transition in automatic mode; manual mode requires explicit reopen/switch-off. Preserve new-study auto reopening under this gate, current Complete counting until actual incomplete-version action, autosave/draft distinction and history. Precise approval/concurrency/draft routing is proposed, not a newly approved permission or spontaneous reopening.
- UA1 / P14: required applicable questions cannot be omitted in completed candidates; optional unanswered questions can be decided by the reconciler. Blank is not N/A. Configurable all-applicable-gold completeness was suggested; project/form scope/default and statistical missing/Unknown comparisons/denominators remain approval proposals. AG3 N/A/version and EX2 available-history decisions stay settled; Complete anyway never waives answer validation.
- PM2 / P12: project owner can assign permission administration to a membership group; authorized members administer/delegate within approved scope. This deliberately extends currently owner-only AssignPermissions. Ownership transfer remains owner-only. Proposed owner-reserved delegation-envelope administration/non-recursive delegation boundaries are explicit recommendations. Possessing a grant never means administering it. Group template names/bundles are implementer discretion after complete action inventory, not owner blockers.
- ODIR1 / P10: one versioned outcome-measure direction across cohorts in a paper/population; no context override. Genuinely different meanings require separate measures. Earlier open override proposal is superseded; retained conflicting legacy values require reviewed mapping.
- MIG1 / P9: drafting outcome migration/adoption plan is now authorized, superseding deferred planning status. Execution, activation and live migration remain unauthorized.
- IP1: concrete implementation planning authorized, not runtime code. Major product choices are covered; review remaining exact proposals before implementation instead of reopening them.
Reviewable standalone documents:
- Permission matrix proposal
- RBAC primary-source research
- Outcome-data migration plan
- Lifecycle and gold-answer settings proposal
- Implementation sequence
All are planning deliverables. Approval choices are marked; no default/grant, data conversion, statistical formula, production lifecycle behavior or source-prototype update is claimed.
DP6 — current cross-stage correction (supersedes DP1 cross-stage extension)¶
Chris, 3 October 2026, during PRISMA integration: personal Include with collective Exclude veto applies to steps within the same stage. Between stages, route availability should be configurable. Proposed choices are Collective Include required versus own Include sufficient; collective Exclude veto, permissions/allocation and independent gating remain. DP7 confirms default Collective Include required, with advanced own-Include option; collective Exclude veto retained. Record this as a change of direction, not a denial of earlier DP1 confirmation. Older blanket personal-Include cross-stage language in historical sections is superseded. Keep within-stage personal work, cross-stage routing and PRISMA collective report authority separate. Cross-stage default is settled under DP7. Review strict within-stage proposal and PRISMA amendments before implementation approval.