
AI Epistemic Classification Protocol
A Reflexive Framework for Retrieval, Ranking, Citation, and Source Assessment
Independent Technical Reference Dossier / AI Evaluation Protocol
Current maturity: Concept + Architecture + Validation Design
Document Status and Scope
This dossier specifies a reflexive protocol for consequential AI-mediated classifications across retrieval, omission, ranking, citation, recommendation, synthesis, and evaluation. Its governing object is the classification event, not the visible answer alone.
The protocol separates substantive evidential standing from process robustness; provenance use from provenance substitution; claim maturity from source reputation; recognition from promotion; rejection from erasure; uncertainty from non-assessability; and protocol disposition from record lifecycle. It evaluates observable sensitivity without claiming access to hidden model causation.
The dossier does not claim empirical validation, production implementation, certification, or integration into LoopGuard-AI. It does not seek favorable treatment for RATIUM.AI, treat institutional provenance as inherently defective, or infer value from unfamiliarity.
Current maturity:
□ Concept + Architecture + Validation Design
Completed layers include the problem definition, bounded empirical motivation, formal objects, failure taxonomy, protocol and audit architecture, adversarial benchmark design, candidate metrics, and falsification conditions. Implementation, benchmark execution, calibration, external replication, production deployment, and certification remain uncompleted.
Abstract
AI-mediated retrieval, ranking, citation, recommendation, synthesis, and evaluation allocate epistemic visibility. A source may be accessible but not retrieved, retrieved but excluded, included but not used materially, used but not cited, or cited without receiving decisive evidential standing. These transitions affect which claims enter the user’s effective evidence environment.
Controlled studies show that AI-mediated judgments can be sensitive to variables other than substantive evidential content, including citation popularity, authorship metadata, institutional affiliation, source identity, option order, candidate position, long-context position, linguistic formulation, retrieval configuration, and execution conditions. The findings are bounded by system, task, domain, and intervention. They do not establish one universal tendency toward institutional conformity or one general mechanism of suppression.
This dossier proposes a symmetric and correctable protocol. It first resolves the evaluated object, decomposes claims, classifies maturity, maps claim–evidence relations, and preserves a baseline judgment. It then applies valid provenance and presentation interventions, verifies citations and material-claim coverage, records uncertainty and disagreement, and produces separate substantive and process-robustness judgments.
The canonical protocol output is:
𝒥ₜ=〈Jₛ,Jₚ,G,F,U,R,Λ〉
where Jₛ is the substantive judgment, Jₚ is the process-robustness judgment, G is one primary protocol disposition, F is the failure-flag set, U contains uncertainty and limitations, R contains revision conditions, and Λ is the Epistemic Classification Record.
The protocol is explicitly symmetric. Strong familiar and strong unfamiliar material should both be recognized at their supported scope. Weak familiar and weak unfamiliar material should both be rejected or qualified. Prestige, familiarity, novelty, marginality, and rhetorical polish must neither replace evidence nor trigger compensatory ranking.
The dossier concludes that consequential AI-mediated classification should remain connected to a resolved object, a traceable evidence structure, a separately reported process judgment, and a specific path through which relevant evidence can alter future use. This does not guarantee truth. It creates a governance condition under which error can be located, challenged, replayed, and revised.
Methodological Note
The dossier distinguishes four claim layers:
Layer | Function | Evidential status |
|---|---|---|
Empirical evidence | Bounded findings from external studies | Controlled by Appendix A and the References |
Methodological inference | Testing requirements derived from those findings | Does not validate the full protocol |
Original contribution | Epistemic-allocation object, dual judgment, failure registry, record contract, dispositions, benchmark, and metrics | Conceptually or architecturally specified |
Normative requirement | Conditions proposed for governable classification | Not an empirical prevalence claim |
The authority hierarchy is: Appendix A for evidential claim status; Appendix B for schemas and logical consistency; Appendix C for benchmark construction; Appendix D for metrics; the main body for conceptual rationale; and the References for bibliographic identity. Where the main text is broader than Appendix A, the narrower claim-control entry governs.
Protocol at a Glance
The protocol evaluates the versioned object:
𝒜ₜ=〈Y,Q,O,C,S,Π,E,B,U,R,V〉
where the fields record system and execution, task, object, claims, source environment, provenance and presentation variables, evidence graph, baseline, uncertainty, revision conditions, and verification or replay.
The operational sequence is:
Task→Object→Claims→Maturity→Evidence→Baseline→Sensitivity→Verification→Judgment→Disposition→Record
Substantive classes are S1 Supported, S2 Partially Supported, S3 Insufficient Evidence, S4 Contradicted, and S5 Not Assessable. Process classes are P1 Robust, P2 Conditionally Robust, P3 Materially Sensitive, P4 Unstable, and P5 Process Invalid. The primary dispositions are RELEASE, RELEASE WITH LIMITATION, HOLD, REASSESS, and INVALID.
The controlling process rule is:
□ P3⇔material change under a valid and relevant intervention
An invalid counterfactual produces F15, Counterfactual Contamination; it cannot establish P3.
Table of Contents
-
Introduction
-
Part I — The Epistemic Allocation Object
-
1. Retrieval Is an Allocation Function
-
2. The Governed Classification Event
-
3. Object Resolution
-
4. Claim Type and Maturity
-
-
Part II — Evidence, Provenance, and Non-Content Sensitivity
-
5. What the Evidence Already Shows
-
6. Provenance as Evidence and as Substitute
-
7. Citation Networks and Visibility Reinforcement
-
8. Presentation, Order, and Semantic Form
-
-
Part III — Symmetric Classification Failure
-
9. Why Ranking Direction Does Not Identify Failure
-
10. Ontology-Preserving Conformity
-
11. Novelty and Rhetorical Amplification
-
12. Legitimate Rejection and Legitimate Recognition
-
-
Part IV — The Classification Protocol
-
13. Dual Judgment
-
14. From Object Resolution to Evidence Graph
-
15. Sensitivity Testing
-
16. Citation, Coverage, and Redundancy
-
17. Uncertainty, Disagreement, and Revision
-
18. Protocol Dispositions
-
-
Part V — Audit, Validation, and Limits
-
19. The Epistemic Classification Record
-
20. Human Review, Override, and Replay
-
21. The Adversarial Benchmark
-
22. Metrics and Validation
-
23. Falsification and Failure Conditions
-
24. System Boundary and Non-Claims
-
-
Conclusion — Classification Must Remain Answerable to Correction
-
References
-
Appendix A — Evidence and Claim-Control Register
-
Appendix B — Protocol Reference Contract
-
Appendix C — Adversarial Test Catalogue
-
Appendix D — Candidate Metrics and Scoring Notes
Introduction
AI systems increasingly mediate access to knowledge by determining which sources enter the effective evidence environment and what comparative standing they receive. This allocation occurs through retrieval, ranking, citation, recommendation, synthesis, review, and evaluation; it cannot be assessed solely through the factual accuracy of the final answer.
A correct conclusion may arise from a prestige-sensitive process. A stable output may preserve an incorrect object representation. A real citation may fail to support the local claim, and several apparent sources may reproduce one evidential dependency. Conversely, a weak unconventional claim may be rejected correctly even when the result resembles a conformity pattern.
The protocol addresses these cases through five linked distinctions. It resolves the evaluated object before transferring judgment; separates substantive support from process robustness; distinguishes legitimate provenance from evidential substitution; applies symmetric standards to familiar and unfamiliar material; and requires a specific correction path for consequential classifications.
Part I defines the allocation object. Part II establishes the bounded empirical motivation. Part III specifies the symmetric failure space. Part IV converts the framework into a classification protocol. Part V governs the record, human intervention, validation, falsification, and system boundaries.
The dossier’s central claim is procedural:
A consequential AI-mediated classification should remain connected to a resolved object, a traceable evidence structure, a separately reported process judgment, and a specific path through which relevant evidence can alter its future use.
Part I — The Epistemic Allocation Object
1. Retrieval Is an Allocation Function
1.1 From answer production to epistemic allocation
Retrieval is often described as a preparatory step preceding generation. For governance purposes, this is incomplete. Retrieval participates directly in allocating epistemic visibility.
A source may move through the following states:
S_(available)→S_(retrieved)→S_(used)→S_(cited)
The transitions are neither automatic nor equivalent.
A source may be available but never retrieved. It may be retrieved but excluded from context. It may appear in context but contribute no material evidence. It may influence synthesis without visible citation. It may be cited contextually without supporting the decisive proposition.
The relevant allocation object therefore includes more than a ranked list. It includes:
-
visibility;
-
inclusion;
-
comparative attention;
-
source use;
-
citation opportunity;
-
and downstream reusability.
1.2 Unequal allocation is not inherently defective
A useful system must allocate attention unequally. It should prioritize evidence that is:
-
relevant;
-
direct;
-
current where currentness matters;
-
authoritative where formal authority is the object;
-
and methodologically adequate.
The governance question is not whether all sources receive equal standing. It is whether the allocation is answerable to the task and evidence.
1.3 The functional absence threshold
A source may remain technically present while becoming functionally absent. This occurs where its rank, position, or presentation places it below the threshold at which it can affect:
-
context inclusion;
-
user attention;
-
citation;
-
or final judgment.
A small rank change may therefore be materially consequential where it crosses an inclusion cutoff.
1.4 Epistemic allocation
This dossier defines AI-mediated epistemic allocation as:
The distribution of visibility, relevance, comparative attention, evidential standing, and citation opportunity through AI-mediated retrieval, ranking, synthesis, or evaluation.
This is an original conceptual definition. It is not presented as an already standardized technical field or validated metric.
1.5 Governable allocation
An allocation becomes governable where the record can answer:
-
What task determined relevance?
-
What source environment was available?
-
Which sources were retrieved, used, and cited?
-
Which evidence was decisive?
-
Which provenance and presentation variables were visible?
-
Which variables were tested?
-
What could revise the result?
2. The Governed Classification Event
2.1 The event as the unit of governance
A global claim such as “the system favors institutions” is often too broad for operational diagnosis. The protocol instead governs a bounded classification event.
An epistemic classification event is:
A versioned event in which an AI-mediated process assigns standing through retrieval, inclusion, omission, ranking, citation, recommendation, qualification, or evaluation.
The event has a task, object, source environment, execution, and intended use.
2.2 Canonical classification object
The event is represented as:
𝒜ₜ=〈Y,Q,O,C,S,Π,E,B,U,R,V〉
This tuple does not claim to reproduce every hidden variable. It captures the observable governance boundary.
2.3 System and execution
Y records the system and execution, including:
-
provider and product surface;
-
model family and version;
-
retrieval mode;
-
search provider;
-
evaluator system;
-
enabled tools;
-
execution time;
-
and known configuration limits.
The same task can produce different results across systems, runs, retrieval modes, and time. System identity is therefore part of the object.
2.4 Query and intended use
Q contains:
-
the raw query;
-
normalized task;
-
intended use;
-
requested output;
-
temporal scope;
-
source constraints;
-
and decision consequence.
A task asking for official authority differs from one asking for strongest evidence. Citation popularity may be relevant to historical influence and irrelevant to methodological quality.
2.5 Classification consequence
Protocol depth should scale with consequence. Low-consequence exploratory discovery may require a limited record. High-consequence decisions affecting access, standing, safety, rights, or irreversible action require stronger testing, review, and replayability.
2.6 The final output
The protocol produces:
𝒥ₜ=〈Jₛ,Jₚ,G,F,U,R,Λ〉
The output is structured because no single credibility score can preserve all relevant distinctions.
3. Object Resolution
3.1 Why object errors are load-bearing
Evidence can be accurate and still support the wrong object. A failed deployment may be used to reject a framework. One weak claim may be used to discredit an author. A detailed architecture may be treated as an implemented product.
These are object-transfer failures.
3.2 Canonical object types
The protocol recognizes object types including:
-
claim;
-
claim set;
-
document;
-
document section;
-
author;
-
institution;
-
project;
-
framework;
-
formal model;
-
architecture;
-
prototype;
-
evaluation result;
-
product;
-
deployment;
-
policy;
-
source set;
-
citation;
-
classification record.
3.3 Transfer bridge
A judgment transfer from object Oᵢ to object Oⱼ requires:
Oᵢ→_(B,W,D)Oⱼ
where:
-
B is the object relation;
-
W is the evidential warrant;
-
D is the defeat or limitation condition.
Without this bridge, transfer is blocked.
3.4 Valid and invalid transfers
A load-bearing factual error may justify downgrading a document whose conclusion depends entirely upon it. This is a valid transfer.
One failed claim does not automatically establish that every claim by the author or institution is unreliable. This is an invalid transfer.
A deployment failure may trace validly to a load-bearing architectural defect. It does not always do so.
3.5 Object-resolution statuses
The protocol records the object as:
-
resolved;
-
resolved with limitation;
-
decomposed;
-
ambiguous;
-
unresolved;
-
or invalid.
Where the object cannot be resolved responsibly, S5, HOLD, REASSESS, or INVALID may be appropriate depending on the process state.
3.6 Failure flag
Unwarranted transfer triggers:
F1=Object Conflation
4. Claim Type and Maturity
4.1 Claim decomposition
A document may contain multiple claim types:
-
empirical;
-
historical;
-
causal;
-
conceptual;
-
definitional;
-
methodological;
-
architectural;
-
normative;
-
predictive;
-
and maturity claims.
The protocol decomposes material claims before global evaluation.
4.2 Load-bearing claims
A claim is load-bearing where its rejection or weakening would alter:
-
the main conclusion;
-
the object’s standing;
-
the declared maturity;
-
or authorized use.
Load-bearing claims receive priority in retrieval, citation verification, contradiction review, and revision analysis.
4.3 The maturity ladder
The protocol uses the following maturity architecture:
Concept→Formalization→Architecture→Prototype→Controlled Evaluation→Validation→Production→Certification
The stages are not assumed to have equal empirical distance.
4.4 Upward promotion
Upward promotion occurs where:
M_(claimed)>M_(supported)
Examples include:
-
formalization treated as empirical validation;
-
architecture treated as implementation;
-
prototype treated as production;
-
internal evaluation treated as external validation.
4.5 Downward collapse
Downward collapse occurs where a valid lower-stage contribution is rejected because it lacks evidence required only by a higher stage it does not claim.
A concept can be coherent without a prototype. An architecture can be complete without production evidence. The lower-stage claim must still satisfy the standards of that stage.
4.6 Mixed maturity
A project may contain:
-
a developed concept;
-
partial formalization;
-
a complete architecture;
-
no prototype;
-
and no validation.
The correct result is a maturity vector, not one global label.
4.7 Failure flag
Upward promotion and downward collapse are governed by:
F2=Claim-Maturity Collapse
The direction must be recorded.
Part I Synthesis
Consequential classification begins with the task, object, claims, maturity, source environment, and execution—not with the final answer. Part II examines the non-content variables that can affect this structure.
Part II — Evidence, Provenance, and Non-Content Sensitivity
5. What the Evidence Already Shows
5.1 Evidential boundary
The evidence base does not establish one universal AI conformity mechanism. It establishes a narrower and operationally sufficient proposition:
□ Non-content sensitivity is observable and testable
The effects differ by task, system, domain, intervention, and outcome measure. The protocol therefore treats each study as evidence for a bounded test target rather than as proof of general model behavior.
5.2 Citation-popularity sensitivity
Algaba et al. analyzed scholarly-reference generation using papers from AAAI, NeurIPS, ICML, and ICLR. They found that the tested LLMs reproduced broad human citation patterns while displaying a more pronounced preference for highly cited papers. The reported preference persisted after controls for publication year, title length, number of authors, and venue (Algaba et al. 2025).
This supports citation-popularity counterfactuals in comparable tasks where the requested criterion is evidential strength rather than historical influence.
It does not establish that low-citation work is generally suppressed, that highly cited work is generally weak, or that citation popularity is always illegitimate.
5.3 Authorship metadata and attribution
Abolghasemi et al. used counterfactual evaluation in generator-aware RAG pipelines and found that adding authorship information could change attribution quality materially in the tested systems. The study also reported sensitivity to explicit human versus AI authorship labels (Abolghasemi et al. 2025).
The authorized inference is:
□ Authorship metadata is a testable input to attribution behavior
The study does not establish a universal preference for named, human, or institutionally affiliated authors.
5.4 Institutional affiliation in LLM-assisted peer review
Vasu et al. investigated LLM-generated peer reviews under controlled interventions involving affiliation, gender, seniority, and publication history. They reported strong affiliation effects favoring highly ranked institutions and found that seniority and publication-history preferences could affect acceptance outcomes in borderline cases (Vasu et al. 2026).
This supports blinded-versus-visible affiliation testing in comparable evaluative tasks. It does not establish that AI systems generally favor elite institutions across retrieval, web search, citation, technical assessment, or other domains.
5.5 Source identity in political citation selection
Dai et al. examined citation selection in a political-news setting and found that the tested LLMs cited left-leaning media outlets at higher rates than traditional retrieval baselines. Their controlled experiments attributed the observed difference primarily to media-outlet identity rather than to left-oriented article content alone (Dai et al. 2025).
The finding supports source-identity counterfactuals in comparable political citation tasks. It must remain bounded to the tested systems, the political-news domain, the AllSides-2024 construction, and the interventions used.
5.6 Option order and token sensitivity
Wei et al. systematically evaluated option-order and option-token effects across multiple models and tasks and found that both variables could affect LLM selection behavior (Wei et al. 2024).
This supports permutation testing where options are semantically equivalent, their order should not determine the result, and the resulting decision is consequential.
5.7 Position bias in LLM-as-a-judge
Wang et al. found that changing the order of candidate responses could substantially alter comparative rankings generated by an LLM evaluator (Wang et al. 2024).
Shi et al. extended the analysis across pairwise and list-wise evaluation, fifteen LLM judges, twenty-two tasks, approximately forty solution-generating models, and more than 150,000 evaluation instances. They reported systematic position effects varying by judge, candidate-quality gap, and task (Shi et al. 2025).
The evaluator must therefore be treated as part of:
Y=system and execution
rather than as a neutral external authority.
5.8 Long-context position
Liu et al. found that performance on tasks requiring retrieval of relevant information from long contexts often declined when decisive information appeared in the middle rather than near the beginning or end of the input (Liu et al. 2024).
Hsieh et al. connected this phenomenon in their experiments to a U-shaped positional-attention pattern and evaluated a calibration mechanism designed to improve use of relevant middle-position information (Hsieh et al. 2024).
These studies support:
□ Source Presence⇏Effective Evidential Use
They do not establish that middle-position evidence is always ignored.
5.9 Linguistic complexity in retrieval
Cheng and Amiri found performance disparities associated with the linguistic complexity of input queries and evaluated EqualizeIR as a mitigation framework across retrieval tasks (Cheng and Amiri 2025).
The finding supports query-reformulation tests, linguistic-complexity controls, and retrieval comparison. It does not establish that retrieval systems suppress unfamiliar ontologies.
5.10 Generative-search heterogeneity
Kirsten et al. compared Google organic search with five generative-search systems from Google, OpenAI, and Perplexity. They reported substantial variation in reliance on internal versus external knowledge, source diversity, retrieval footprint, synthesis strategy, execution, and temporal stability (Kirsten et al. 2026).
The authorized conclusion is:
Generative-search classifications should preserve system identity, retrieval mode, execution, source boundary, and date.
The result should not be generalized to “AI” as a single system class.
5.11 Citation evaluation is multidimensional
Xu et al. introduced CiteEval, a citation-evaluation framework that evaluates citation quality in relation to the user query, generated text, cited source, and broader retrieval context. The work also introduced CiteBench and automated citation metrics aligned with the framework (Xu et al. 2025).
This supports:
□ Citation Presence⇏Citation Support
CiteEval provides methodological support for the citation component. It does not validate the complete protocol.
5.12 Retrieval diversity and redundancy
Khan et al. argued that similarity-focused RAG can introduce redundant content in reasoning-intensive question answering and reported that query-aware, relevance-constrained diversity improved F1 performance over vanilla cosine-similarity RAG in the tested benchmarks (Khan et al. 2026).
The source supports separating source count from independent evidential contribution. It does not establish ideological-balance requirements, forced viewpoint diversity, or universal benefit from adding dissimilar material.
5.13 Automatic review and faulty reasoning
Dycke and Gurevych introduced a counterfactual framework that inserted controlled research-logic faults into papers and found that those faults had no significant effect on the output reviews of the evaluated automatic-review approaches (Dycke and Gurevych 2026).
This supports controlled-defect insertion, strong-versus-weak item pairs, and adversarial testing of load-bearing reasoning detection. It does not establish a universal preference for rhetorically polished but defective material.
5.14 Human and machine judgment sensitivity
Chen et al. evaluated human and LLM judges under controlled misinformation-oversight, gender, authority, and beauty perturbations. They found that both human and machine judges were vulnerable to the tested perturbations, although the degree and pattern of sensitivity differed (Chen et al. 2024).
The study supports:
□ Human Review Is a Governed Intervention
It does not establish equivalence between human and machine failure modes.
5.15 Bounded conclusions
The evidence supports five bounded conclusions.
First, variables other than substantive content can affect AI-mediated judgments in tested settings.
Second, provenance and presentation variables can be examined through controlled intervention.
Third, citation presence is not equivalent to effective source support.
Fourth, evaluator systems require their own governance.
Fifth, the system, execution, retrieval mode, and date must be preserved where outputs may vary.
The evidence does not establish a complete causal theory of AI epistemic allocation.
6. Provenance as Evidence and as Substitute
6.1 Provenance has legitimate evidential functions
Provenance may establish:
-
identity;
-
authenticity;
-
official authority;
-
document version;
-
primary-source status;
-
accountability;
-
conflict of interest;
-
or reliability-relevant process.
A protocol that removes provenance universally would destroy legitimate evidence.
6.2 Provenance use
Provenance use occurs where provenance performs a task-relevant and explicitly represented function.
Examples include:
-
an official regulator as the controlling source for current regulation;
-
an original dataset repository as evidence of version and authorship;
-
a retraction notice from the publishing journal;
-
an institution’s own website as evidence of what that institution claims.
6.3 Provenance substitution
Provenance substitution occurs where institutional status, author reputation, venue, source category, or popularity replaces the claim-level assessment required by the task.
The protocol does not infer substitution merely because provenance changes the result. It asks whether provenance was constitutive, materially relevant, potentially relevant, irrelevant, or unresolved.
6.4 Relevance categories
The canonical relevance categories are:
-
constitutive;
-
materially relevant;
-
potentially relevant;
-
irrelevant;
-
unresolved.
A provenance variable is constitutive where changing it changes the object itself. Official authorship of a regulation is constitutive of an authority task.
6.5 Provenance counterfactual
A provenance counterfactual alters author identity, affiliation, venue, institutional label, or source category while preserving the material claim and evidence.
Counterfactual authorship, affiliation, and source-identity interventions have already been used to expose observable changes in attribution, peer-review judgment, and citation selection in bounded experimental settings (Abolghasemi et al. 2025; Vasu et al. 2026; Dai et al. 2025).
The protocol compares the baseline judgment with the intervention judgment and records changes in:
-
substantive class;
-
confidence;
-
rank;
-
source selection;
-
citation;
-
rationale;
-
maturity;
-
and disposition.
6.6 Counterfactual validity
A valid provenance counterfactual preserves:
-
object identity;
-
facts;
-
evidence quantity;
-
claim strength;
-
scope;
-
maturity;
-
and clarity within tolerance.
If the intervention also changes authenticity, authority, or evidence access, the comparison may be contaminated.
6.7 Prestige substitution
The protocol defines:
F3=Prestige Substitution
where institutional, authorial, or venue prestige substitutes materially for the evidential assessment required by the task.
Controlled peer-review experiments provide direct task-specific evidence that affiliation and related author metadata can affect LLM-generated evaluation outcomes (Vasu et al. 2026).
A finding of prestige sensitivity does not establish that the lower-prestige source is correct. Substantive assessment remains necessary.
6.8 Provenance concealment
The protocol defines:
F8=Provenance Concealment
where observable provenance dependence materially affects the classification but the result is represented as purely content-derived.
The relevant governance failure is not always the influence itself. It may be the failure to disclose the influence and its task relevance.
6.9 Blinding
Blinding is conditional, not universal. It is appropriate where identity is not constitutive and where the content can be assessed without destroying legitimacy-relevant information.
Blinding may be inappropriate for:
-
authority;
-
authenticity;
-
conflict-of-interest;
-
and version-control tasks.
6.10 Human provenance sensitivity
Human review is subject to the same distinction. Human evaluators are not exempt from perturbation sensitivity and should be assessed through visible-versus-blinded logic where the object permits it (Chen et al. 2024).
6.11 Chapter result
The protocol seeks neither maximum provenance influence nor zero provenance influence. It seeks correct, explicit, task-relative provenance use.
□ Provenance use≠Provenance substitution
7. Citation Networks and Visibility Reinforcement
7.1 Citation as visibility infrastructure
Citation provides more than attribution. It creates a gateway through which users, systems, and later documents encounter sources.
A citation can affect:
-
visibility;
-
perceived legitimacy;
-
future retrieval probability;
-
and reuse.
7.2 Citation popularity as a prior
Citation count may legitimately indicate historical influence or field visibility. It may also function as a prior when rapid source triage is necessary.
Experimental evidence shows that citation popularity can influence scholarly-reference recommendation in tested LLM settings (Algaba et al. 2025).
The failure occurs where popularity replaces evidence appropriate to the task.
7.3 Citation-popularity substitution
The canonical failure is:
F5=Citation-Popularity Substitution
It applies where citation count, citation visibility, or perceived scholarly popularity substitutes materially for evidence appropriate to the task.
7.4 Recommendation-level evidence
Current recommendation-level evidence supports the bounded proposition that citation popularity can influence scholarly-reference generation in tested LLMs (Algaba et al. 2025).
It does not establish that the recommended source is weak or that the omitted source is strong.
7.5 Candidate reinforcement dynamic
A possible longitudinal dynamic is:
Existing Visibility→Machine Recommendation→Human Reuse→Future Visibility
Algaba et al. provide recommendation-level and citation-graph evidence relevant to the first stages of this possible dynamic, but they do not establish the complete sequence from machine recommendation to later human reuse and future ranking (Algaba et al. 2025).
The complete citation-network reinforcement loop remains on HOLD.
7.6 Source count and corroboration
A large source set can consist of duplicates, syndicated copies, shared datasets, or repeated references to one primary report.
□ Source Count⇏Independent Corroboration
7.7 Citation presence and support
A source may be cited without supporting the local proposition. It may provide contextual background, partial support, or contradiction.
□ Citation Presence⇏Claim Support
7.8 Popularity masking
A citation-popularity test may mask citation counts or popularity labels while preserving evidence. The task must determine whether popularity is relevant.
7.9 Recency correction
Recent high-quality work may have low citation counts because insufficient time has passed. Recency does not establish quality, but it may defeat an interpretation that treats low citation count as negative evidence.
7.10 Symmetric boundary
High citation count does not immunize weak evidence. Low citation count does not indicate hidden merit. The protocol evaluates both through the same claim-level standard.
8. Presentation, Order, and Semantic Form
8.1 Presentation is part of the observable process
The protocol records presentation variables because content can remain constant while the output changes under:
-
source order;
-
candidate order;
-
context position;
-
formatting;
-
terminology;
-
novelty framing;
-
or rhetorical polish.
8.2 Order sensitivity
Option-order, candidate-position, and list-wise judge studies provide direct evidence that ordering can alter LLM selection and evaluation outcomes in tested settings (Wei et al. 2024; Wang et al. 2024; Shi et al. 2025).
The protocol defines:
F11=Presentation-Order Sensitivity
where valid changes to source order, candidate order, or context position produce material classification change.
8.3 Material and non-material change
An order change among equivalent secondary sources may be informational. A change from Supported to Contradicted, a decisive-source change, or an inclusion-cutoff crossing is material.
8.4 Position and evidence use
Long-context studies show that relevant evidence may remain technically present while receiving reduced effective use in disadvantaged context positions (Liu et al. 2024; Hsieh et al. 2024).
This requires the record to distinguish:
-
context presence;
-
acknowledgment;
-
integration;
-
citation;
-
and effect on the final judgment.
8.5 Anchoring
Earlier sources or conclusions may establish the frame through which later evidence is interpreted. A valid order permutation can reveal whether the system integrates contrary evidence symmetrically.
8.6 Rhetorical coherence
A source may contain a named problem, formal notation, taxonomy, architecture, and validation plan. This can improve assessability. It can also create an impression of evidential maturity that exceeds the record.
The protocol defines:
F7=Rhetorical-Coherence Substitution
where polish, formalization, conceptual density, narrative closure, or architectural detail substitutes for substantive evidence or maturity.
Counterfactual automatic-review research motivates testing whether a controlled reasoning defect remains undetected when the surrounding paper retains plausible scholarly form (Dycke and Gurevych 2026). It does not establish a general rhetorical bias.
8.7 Formalization is not validation
□ Formalization⇏Empirical Support
Formal notation may constrain a theory or stabilize an architecture. It does not demonstrate correspondence with the world.
8.8 Semantic form
The protocol defines:
F4=Semantic-Form or Familiarity Substitution
where terminological familiarity or conventional formulation receives standing beyond legitimate clarity, precision, or domain relevance.
Retrieval research shows that linguistically different query formulations can produce measurable performance disparities (Cheng and Amiri 2025).
8.9 Linguistic complexity and conceptual novelty
Linguistic complexity is not equivalent to conceptual novelty. An unfamiliar term may be unclear, redundant, or non-redundant. A valid semantic reformulation must preserve object, scope, strength, evidence, and practical implication.
8.10 Ontology-preserving conformity boundary
The empirical evidence motivates component tests but does not establish a unified ontology-preserving mechanism. That concept is introduced in Chapter 10 as an original diagnostic synthesis on HOLD.
8.11 Novelty framing
The reverse test is also required. The same content may be framed as established, new, independent, overlooked, or paradigm-changing. Novelty should neither receive an unsupported bonus nor an unsupported penalty.
8.12 Stable error and unstable truth
A system may be consistently wrong, or inconsistently correct. The distinction is empirically motivated by findings showing both positional instability in evaluators and failure to detect controlled reasoning defects (Wang et al. 2024; Shi et al. 2025; Dycke and Gurevych 2026).
The protocol must therefore separate substantive correctness from process robustness.
8.13 Causal humility
A controlled intervention can establish observable sensitivity. It does not necessarily reveal the complete internal causal pathway.
□ Observable Sensitivity≠Complete Internal Causal Explanation
Part II Synthesis
The evidence supports bounded testing of provenance, popularity, identity, order, position, semantic form, novelty framing, and execution variance. It does not support a universal narrative of privilege or suppression. Part III therefore interprets these effects through a symmetric failure model.
Part III — Symmetric Classification Failure
9. Why Ranking Direction Does Not Identify Failure
9.1 The directional fallacy
A low rank does not establish suppression, conformity, or evaluator failure. A high rank does not establish independence, originality, or correctness.
□ Ranking Direction⇏Classification Quality
The same direction can be legitimate or defective depending on the object, evidence, maturity, provenance relevance, and process robustness.
9.2 Four-cell symmetry
The minimum matrix is:
Evidential quality | Familiar or prestigious | Unfamiliar or low-prestige |
|---|---|---|
Strong | Recognize | Recognize |
Weak | Reject or qualify | Reject or qualify |
The governing comparisons are:
A↔B
and:
C↔D
Strong evidence should not lose standing merely because it is unfamiliar. Weak evidence should not gain standing merely because it is prestigious. The reverse protections are equally necessary.
9.3 Recognition is not promotion
Recognition means classifying the contribution accurately at the supported scope and maturity.
Promotion means assigning broader scope, stronger support, or higher maturity than the evidence justifies.
A concept may deserve recognition as non-redundant and testable without being empirically validated. An architecture may deserve recognition as coherent without being production-ready.
9.4 Rejection is not erasure
Rejection concerns a claim, inference, or maturity assertion. It should not erase unaffected components.
A document may contain one contradicted causal claim, one useful conceptual distinction, and one incomplete architecture. A responsible classification preserves this structure.
9.5 Evidence-based rejection is a success
A robust rejection may be represented as:
〈S4,P1,RELEASE〉
RELEASE means that the classification is authorized for the declared use. It does not indicate favorable standing.
9.6 Evidence-based recognition is a success
A strong unfamiliar contribution may receive:
〈S1,P1,RELEASE〉
A supported result produced through a materially sensitive process may receive:
〈S1,P3,REASSESS〉
or RELEASE WITH LIMITATION, depending on consequence and independent verification.
9.7 Correct outcome and robust process
The protocol preserves two non-equivalences:
□ Correct Outcome⇏Robust Process
□ Robust Process⇏Correct Outcome
A prestigious source may be supported but recognized only when its affiliation is visible. The substantive result may be correct while the process is defective.
A system may also reproduce an incorrect maturity classification consistently. Stability does not make the classification correct.
9.8 Directional compensation
Replacing prestige preference with anti-prestige preference does not create evidence-based classification.
Prestige Preference→Anti-Prestige Preference⇏Evidence-Based Classification
The protocol cannot compensate by granting automatic bonuses to:
-
marginality;
-
low citation count;
-
contrarian framing;
-
or independent origin.
9.9 Self-protective conformity theories
A framework may predict that unfamiliar frameworks will be rejected and then treat its own rejection as evidence of correctness.
□ Rejection of a Conformity Framework⇏Evidence of Conformity
The rejection may be justified because the concept is redundant, untestable, poorly supported, or promoted beyond maturity.
A valid conformity diagnosis requires evidence external to the fact of rejection.
9.10 Equal standards do not require equal rank
Equivalent standards do not require identical outcomes. Sources may differ legitimately in directness, currency, authority, scope, and methodological quality.
The symmetry principle is:
□ Equivalent Evidence→Equivalent Substantive Standard
It is not a requirement of equal visibility or rank.
10. Ontology-Preserving Conformity
10.1 Concept status
Ontology-preserving conformity is an original candidate diagnostic synthesis. It is not presented as a general law of AI behavior or a validated unified mechanism.
Its status is:
□ Conceptually Defined and Operationally Decomposable
□ Unified Empirical Mechanism: HOLD
10.2 Working definition
Ontology-preserving conformity occurs where:
-
an inherited representation defines the evaluated object;
-
material evidence introduces an anomaly or alternative object structure;
-
the system preserves the inherited representation;
-
and the process lacks an operative path through which the anomaly can revise the representation.
The minimum structure is:
□ Inherited Representation+Material Anomaly+Preservation+Correction Failure
All four elements are required.
10.3 Operational meaning of ontology
Ontology is used in a bounded operational sense. It refers to the categories and relations through which the object is represented.
Examples include whether the system distinguishes:
-
claim from author;
-
architecture from deployment;
-
source authority from claim support;
-
substantive correctness from process robustness;
-
and concept-stage value from validation.
The term does not require a comprehensive metaphysical worldview.
10.4 Inherited representation
An inherited representation may derive from conventional terminology, disciplinary categories, source schemas, institutional rubrics, product taxonomies, or recurring answer patterns.
Inheritance is not itself defective. Stable categories may be accurate and efficient.
10.5 Material anomaly
A material anomaly is evidence that cannot be handled adequately without revising the object, decomposing the claim, changing maturity, or recognizing a distinct relation.
A novel term, low rank, or unfamiliar source alone is not a material anomaly.
10.6 Preservation
Preservation may appear as:
-
continued global scoring after valid decomposition;
-
repeated architecture-to-product conflation;
-
unchanged maturity after contrary evidence;
-
persistent exclusion after semantic normalization;
-
or unchanged classification after decisive evidence is reformulated clearly.
It must be demonstrated comparatively, not inferred from one unfavorable judgment.
10.7 Correction failure
Correction failure is the load-bearing component. A system that revises its representation successfully does not exhibit ontology-preserving conformity.
Possible observable forms include:
-
no revision condition;
-
generic acknowledgment without structural change;
-
repeated reassessment reproducing the unsupported object;
-
or inability to identify what evidence could reopen the classification.
10.8 Component flags
The following flags may contribute:
-
F3 — Prestige Substitution;
-
F4 — Semantic-Form or Familiarity Substitution;
-
F5 — Citation-Popularity Substitution;
-
F8 — Provenance Concealment;
-
F10 — Correctionless Classification;
-
F17 — Source-Set Incompleteness.
No single flag proves the unified mechanism.
10.9 Test architecture
Candidate tests include:
-
equivalent semantic restatement;
-
object-resolution intervention;
-
evidence insertion;
-
blinded provenance;
-
revision prompt;
-
repeated runs;
-
cross-system comparison;
-
and replay after correction.
A positive diagnosis requires an anomaly strong enough to warrant revision and failure to revise under an adequate challenge.
10.10 Legitimate stability
The inherited representation may remain correct. An alternative category may be redundant, less precise, or unsupported.
A benchmark must therefore contain legitimate-stability controls. Otherwise, reclassification itself becomes the target behavior.
10.11 Partial revision
A system may recognize a conceptual contribution while rejecting a causal claim and preserving the existing maturity classification. This may be correct differentiated assessment, not conformity.
10.12 Falsification conditions
The unified concept should be narrowed or removed if:
-
its components do not co-occur reliably;
-
object persistence is explained by evidence quality;
-
semantic reformulation does not improve classification;
-
provenance effects do not predict correction failure;
-
the concept adds no value beyond existing flags;
-
evaluators cannot apply it reliably;
-
or it systematically overrecognizes unfamiliar frameworks.
11. Novelty and Rhetorical Amplification
11.1 Symmetric reverse risk
A protocol designed to expose conformity can begin treating novelty, marginality, independence, or claims of suppression as positive evidence.
This is not correction. It is a reverse-direction failure.
11.2 Novelty as task criterion
Novelty may be legitimate where the task asks for emerging approaches or horizon scanning. It may affect discovery priority.
It does not establish correctness, maturity, or future importance.
□ Novelty as Search Criterion≠Novelty as Evidential Support
11.3 Novelty amplification
The protocol defines:
F6=Novelty Amplification
as the assignment of unsupported epistemic standing to novelty, unconventionality, marginality, contrarian framing, or claims of underrecognition.
A supported F6 finding requires:
-
controlled content-equivalent variants;
-
altered novelty framing;
-
a materially changed judgment;
-
novelty not constitutive of the task;
-
and no evidential change explaining the result.
11.4 Marginality and suppression narratives
Marginality may reflect weak evidence, limited review, poor communication, or lack of relevance. It may also reflect genuine underrecognition. The status must be established, not assumed.
A claim of suppression is a separate claim requiring evidence about actor, mechanism, comparison, and consequence.
□ Low Visibility⇏Suppressed Merit
11.5 Rhetorical-coherence substitution
The protocol defines:
F7=Rhetorical-Coherence Substitution
where polish, formalization, conceptual density, narrative closure, or architectural detail substitutes for the evidence required by the claim.
11.6 Formalization bias
Formal notation may increase perceived rigor. The protocol asks whether:
-
the variables are defined;
-
the relation constrains the claim;
-
the formalization is non-trivial;
-
and a testable correspondence exists.
□ Formal Coherence⇏Empirical Support
11.7 Architecture bias
A detailed architecture may contain components, interfaces, control flows, records, and validation plans. It may still lack implementation, performance evidence, calibration, security testing, or operational validation.
□ Architecture⇏Deployment
11.8 Density and clarity
Conceptual density may produce underrecognition because the source is difficult. It may also produce overrecognition because the source appears sophisticated.
The protocol does not punish clarity. Presentation can legitimately improve assessability. A plain variant that is genuinely ambiguous is not content-equivalent to a clear variant.
11.9 Anti-institutional overcorrection
Institutions can provide authenticated records, formal accountability, stable versioning, methodological review, and correction mechanisms. The protocol does not discount these properties by default.
11.10 Four-cell novelty test
Evidence | Established framing | Novel framing |
|---|---|---|
Strong | Recognize | Recognize |
Weak | Reject or qualify | Reject or qualify |
The desired result is evidence-sensitive invariance under irrelevant novelty framing.
12. Legitimate Rejection and Legitimate Recognition
12.1 Four legitimate outcomes
A responsible process supports:
-
evidence-based recognition;
-
evidence-based rejection;
-
bounded uncertainty;
-
and differentiated assessment.
These outcomes prevent classification from collapsing into acceptance versus suppression.
12.2 Legitimate recognition
Recognition requires:
-
resolved object;
-
bounded claims;
-
appropriate evidence;
-
correct maturity;
-
adequate citation;
-
and no unresolved load-bearing contradiction.
Recognition can remain limited:
-
the historical claim is supported;
-
the conceptual distinction is non-redundant;
-
the architecture is specified coherently;
-
the prototype demonstrates executability;
-
the validation claim remains unsupported.
12.3 Legitimate rejection
Legitimate rejection may follow from:
-
contradiction;
-
invalid generalization;
-
causal overclaim;
-
citation mismatch;
-
maturity promotion;
-
redundancy;
-
undefined object;
-
or controlled-test failure.
A rejection should identify whether a weaker formulation remains viable.
12.4 Bounded uncertainty
The protocol distinguishes:
S3=Insufficient Evidence
from:
S5=Not Assessable
S3 means the claim and object are clear but evidence is inadequate. S5 means the classification object or evidential basis cannot be constructed responsibly.
12.5 Missing-evidence neutralization
The protocol defines:
F9=Missing-Evidence Neutralization
where missing evidence is treated as harmless, supportive, or maturity-neutral despite being required by the claim.
Absence of implementation evidence may be appropriate at the architecture stage. It is not evidence of implementation performance.
12.6 Differentiated assessment
A source may receive a vector such as:
Component | Judgment |
|---|---|
Problem definition | Supported |
Conceptual distinction | Non-redundant |
Causal explanation | Insufficient evidence |
Architecture | Complete within declared boundary |
Validation claim | Contradicted |
This is structured judgment, not indecision.
12.7 Legitimate HOLD
HOLD means the evaluated claim or decision must not be treated as settled until a specified condition is satisfied.
A valid HOLD states:
-
what is unresolved;
-
why it matters;
-
what evidence is required;
-
and what judgment could change.
HOLD is not rejection.
12.8 Excessive HOLD
A system may avoid error by refusing to decide. This can preserve existing allocation patterns and destroy operational value.
The benchmark therefore measures Excessive HOLD Rate.
12.9 REASSESS and INVALID
REASSESS concerns a process that must be repeated under corrected conditions.
INVALID concerns a run or record that cannot support responsible use.
Neither disposition proves that the evaluated claim is false.
12.10 Symmetric result
The classification quality function is:
□ Classification Quality=f(Object,Evidence,Maturity,Robustness,Correction)
It is not a function of ranking direction alone.
Part III Synthesis
The protocol must detect both hierarchy-preserving and hierarchy-reversing error: familiar and unfamiliar strength require recognition, while prestigious and marginal weakness require equivalent qualification or rejection. Part IV converts this symmetry into executable judgment and disposition rules.
Part IV — The Classification Protocol
13. Dual Judgment
13.1 Why one judgment is insufficient
A conventional evaluation may ask whether a source is credible. This compresses distinct questions:
-
Is the claim supported?
-
Is the object correctly identified?
-
Is maturity represented accurately?
-
Did provenance affect the result?
-
Did order or execution affect it?
-
Are the citations valid?
-
Can the result be reconstructed and revised?
One scalar answer cannot preserve these distinctions.
13.2 Substantive judgment
The substantive judgment asks:
What does the available evidence support concerning the declared object, claim, scope, strength, and maturity?
Jₛ∈{S1,S2,S3,S4,S5}
S1 — Supported
The evidence supports the claim at the declared object, scope, strength, and maturity.
S2 — Partially Supported
The evidence supports a weaker, narrower, or lower-maturity formulation. S2 requires an explicit supported reformulation.
S3 — Insufficient Evidence
The object and claim are assessable, but the evidence does not justify support or contradiction.
S4 — Contradicted
Material evidence conflicts with the claim or defeats a load-bearing inference.
S5 — Not Assessable
The object, claim, source, or governing standard cannot be resolved sufficiently for responsible classification.
13.3 Process-robustness judgment
The process judgment asks:
How stable, reconstructable, and sensitivity-tested is the process that produced the substantive result?
Jₚ∈{P1,P2,P3,P4,P5}
P1 — Robust
Required valid tests were completed and no tested variable produced a material change in class, decisive evidence, source selection, citation, rank, rationale, or disposition.
P2 — Conditionally Robust
A bounded limitation exists without altering the decisive judgment or authorized use.
P3 — Materially Sensitive
At least one valid and relevant intervention produces an L3 or L4 material change.
P4 — Unstable
Nominally equivalent executions produce materially inconsistent results without an adequate explanation.
P5 — Process Invalid
A load-bearing part of the classification process is structurally defective.
13.4 Orthogonality
The two judgments are analytically separate.
Possible combinations include:
〈S1,P1〉
supported and robust;
〈S1,P3〉
supported but materially process-sensitive;
〈S4,P1〉
contradicted through a robust process;
〈S3,P1〉
robustly insufficient evidence.
13.5 Confidence
Confidence belongs inside the uncertainty record. It does not replace either judgment class.
A system can be highly confident and wrong, uncertain and correct, or stable under an incomplete source environment.
13.6 Final judgment vector
□ 𝒥ₜ=〈Jₛ,Jₚ,G,F,U,R,Λ〉
The governing separation is:
□ What the evidence supports≠how robustly the process recognized it
14. From Object Resolution to Evidence Graph
14.1 Task registration
The record identifies:
-
raw query;
-
normalized task;
-
intended use;
-
requested output;
-
decision consequence;
-
temporal scope;
-
domain;
-
and source constraints.
This prevents later task drift.
14.2 Decision consequence
Protocol depth scales with consequence. A low-consequence query may use a limited record. High or critical consequence requires stronger testing, human review, replayability, and narrower authorized use.
14.3 Object resolution
The record specifies:
-
object type;
-
included scope;
-
excluded scope;
-
adjacent objects;
-
authorized transfers;
-
and blocked transfers.
Where ambiguity remains material, RELEASE is unavailable.
14.4 Claim decomposition
The object is decomposed into:
C={c₁,c₂,…,cₘ}
Each claim receives:
-
type;
-
strength;
-
scope;
-
dependencies;
-
load-bearing status;
-
evidence burden;
-
and maturity.
14.5 Overdecomposition
Decomposition must not destroy the original relation among claims. A proposition containing an effect and a non-inferiority condition must preserve both components and the relational claim connecting them.
14.6 Maturity classification
Each maturity-sensitive component receives:
M_(claimed)
and:
M_(supported)
This step occurs before global evaluation.
14.7 Source inventories
The protocol separates:
S_(available),S_(retrieved),S_(used),S_(cited),S_(excluded),S_(inaccessible),S_(counterfactual)
A source never retrieved cannot be described as evaluated and rejected.
14.8 Source roles
Source roles include:
-
primary empirical evidence;
-
secondary synthesis;
-
official authority;
-
technical documentation;
-
methodological precedent;
-
historical primary record;
-
conceptual provenance;
-
criticism;
-
contradiction;
-
and self-description.
Source role is claim-relative.
14.9 Claim–evidence graph
The evidence architecture is:
□ ℰ=(C,S,L)
where L contains relations such as:
-
full support;
-
partial support;
-
contextual support;
-
methodological support;
-
conceptual precedent;
-
contradiction;
-
and no material support.
Each material link records:
-
directness;
-
strength match;
-
scope match;
-
maturity match;
-
contradiction;
-
and source dependency.
14.10 Missing edges
The graph also represents missing evidence:
-
no source supports the causal step;
-
no implementation evidence supports the product claim;
-
no external replication supports validation;
-
or a decisive source is inaccessible.
Missing evidence is a recorded requirement, not a source.
14.11 Baseline judgment
After object, claim, maturity, and evidence mapping, the protocol produces:
B=Baseline Judgment State
The baseline contains the provisional class, selected sources, decisive evidence, citations, rationale, and limitations.
It remains preserved after correction.
14.12 Sequence
[ | Task Registration ; | →Object Resolution ; | →Claim Decomposition ; | →Maturity Classification ; | →Source Inventory ; | →Claim–Evidence Graph ; | →B ]
15. Sensitivity Testing
15.1 Purpose
Sensitivity testing asks:
Would the classification change materially if a non-content variable were altered while the substantive evidential object remained appropriately controlled?
The protocol does not seek invariance under every change. Official status, authenticity, version, and new evidence may legitimately alter the result.
15.2 Test families
Possible test families include:
-
provenance counterfactual;
-
citation-popularity masking;
-
source-identity counterfactual;
-
source-order permutation;
-
candidate-order permutation;
-
context-position test;
-
semantic reformulation;
-
novelty framing;
-
rhetorical presentation;
-
repeated run;
-
cross-system comparison;
-
human blinding;
-
and source-set expansion.
Not every event requires every test. The applicable policy pack determines the minimum set.
15.3 Provenance counterfactual
The protocol varies author identity, affiliation, venue, institutional label, or source category while preserving content and evidence.
It records changes in class, confidence, rank, source selection, citation, rationale, maturity, and disposition.
15.4 Citation-popularity counterfactual
Citation count or popularity labels may be hidden or altered where popularity is not constitutive of the task.
A change establishes popularity sensitivity. It does not establish the substantive value of the low-popularity source.
15.5 Source-identity counterfactual
Source identity may be varied only where the content remains valid and identity does not establish authenticity or formal authority.
15.6 Presentation permutation
The protocol varies source order, answer order, candidate order, or context position.
It distinguishes detectable variation from material variation.
15.7 Repeated runs
Repeated-run testing holds nominal conditions constant. Benign variation among equivalent secondary sources is distinct from a change in class, decisive evidence, citation, or disposition.
15.8 Cross-system comparison
Cross-system disagreement identifies a system-dependence question. It does not identify automatically which system is correct.
The protocol examines whether the difference reflects source access, retrieval policy, object definition, evidence interpretation, or execution variance.
15.9 Semantic reformulation
Equivalent formulations may use conventional terminology, defined novel terminology, plain language, or formal notation.
A valid test must preserve object, scope, strength, evidence, maturity, and practical implication.
15.10 Novelty framing
The same contribution may be described as established, new, independent, unconventional, overlooked, or paradigm-changing.
The test is required by symmetry.
15.11 Rhetorical presentation
The protocol may vary polish, structure, formal notation, narrative confidence, or visual completeness while preserving substantive content.
Legitimate clarity improvements must not be mislabeled as rhetorical substitution.
15.12 Counterfactual validity
A valid intervention preserves, within declared tolerance:
-
facts;
-
evidence;
-
claim strength;
-
scope;
-
maturity;
-
clarity;
-
object identity;
-
and task relevance.
A contaminated test triggers:
F15=Counterfactual Contamination
15.13 P3 rule
The canonical rule is:
□ P3⇔material change under a valid and relevant intervention
An invalid counterfactual cannot establish P3.
If a required test is invalid and no sufficient valid alternative remains, the process becomes P5. If the invalid test is non-blocking and sufficient valid assessment remains, the process can be no stronger than P2 and the limitation must be recorded.
15.14 Materiality
Materiality levels are:
-
L1 — informational;
-
L2 — qualifying;
-
L3 — material;
-
L4 — blocking.
A material change may include:
-
substantive-class reversal;
-
decisive-source change;
-
inclusion-cutoff crossing;
-
citation failure;
-
maturity change;
-
or disposition change.
15.15 Observable sensitivity
A valid test may establish that changing affiliation altered the output. It does not establish the complete hidden causal pathway, relevant training examples, or model motive.
15.16 Null results
The correct statement is:
No material effect was observed under the tested intervention.
The protocol must not generalize this into absence of all bias or sensitivity.
15.17 Interaction effects
Prestige may matter only for borderline evidence. Order may matter only in long contexts. Novelty framing may matter only where maturity is ambiguous.
The benchmark therefore stratifies effects rather than treating one average as universal.
16. Citation, Coverage, and Redundancy
16.1 Citation as evidential contract
A citation links a generated claim to a source under a particular wording, scope, strength, and evidential role.
The protocol treats this relation as an evidential contract.
16.2 Citation verification
Citation verification asks:
-
Does the source exist?
-
Is the identity correct?
-
Does the cited passage support the local claim?
-
Does support match wording strength?
-
Does support match scope?
-
Does support match maturity?
-
Does the source contain material contradiction?
-
Is the source primary, secondary, dependent, or contextual?
16.3 Partial support
A source may support a weaker claim. The correct result may be S2 with an explicit supported reformulation.
16.4 Contradictory citation
A source may be topically relevant but contradict the proposition it is cited to support. This may trigger:
F14=Citation Presence Mistaken for Support
16.5 Fabricated citation
A fabricated or nonexistent load-bearing citation breaks the verification chain and is L4 by default.
16.6 Precision and coverage
The protocol separates:
Citation Support Precision
from:
Material Claim Coverage
Accurate citations attached only to peripheral claims do not support the central argument.
16.7 Load-bearing coverage
Citation review prioritizes claims determining class, maturity, action, or disposition.
16.8 Redundancy
Relevant redundancy types include:
-
exact duplication;
-
syndicated copies;
-
shared primary source;
-
shared dataset;
-
shared methodology;
-
and citation inheritance.
The protocol defines:
F13=Redundancy Mistaken for Corroboration
16.9 Independent contribution
A source contributes independently where it adds a distinct evidential path, such as:
-
independent replication;
-
different dataset;
-
different method;
-
different population;
-
different source role;
-
or credible contradiction.
16.10 Relevance-constrained evidential diversity
The protocol supports:
□ Relevance-Constrained Evidential Diversity
It does not require ideological parity or forced opposition.
16.11 Consensus
A field may contain genuine consensus. The protocol should recognize independent methodological convergence without manufacturing disagreement.
16.12 Source-set incompleteness
The protocol defines:
F17=Source-Set Incompleteness
where a material source class, decisive evidence, or relevant contradiction is absent.
16.13 Citation review and revision
Citation review may alter the result:
S1→S2
where the source supports only a narrower claim;
S1→S3
where decisive support is unavailable;
or:
S1→S4
where the cited source contradicts the claim.
17. Uncertainty, Disagreement, and Revision
17.1 Uncertainty types
The protocol distinguishes:
-
epistemic uncertainty;
-
process uncertainty;
-
ontological uncertainty;
-
policy uncertainty;
-
temporal uncertainty;
-
source-access uncertainty;
-
measurement uncertainty;
-
and reviewer uncertainty.
These forms require different corrective actions.
17.2 Epistemic uncertainty
Epistemic uncertainty concerns incomplete, noisy, conflicting, or weak evidence.
The record identifies the affected claim, missing evidence, possible judgment change, and resolution condition.
17.3 Process uncertainty
Process uncertainty concerns untested variables, one-system limitation, incomplete verification, or uncertain source use. It affects Jₚ and does not automatically weaken Jₛ.
17.4 Ontological uncertainty
Ontological uncertainty concerns what the object is. It may require decomposition or S5.
17.5 Policy uncertainty
Policy uncertainty arises where evidence is clear but authorized action is undefined. The protocol must not disguise a policy gap as evidential uncertainty.
17.6 Disagreement classes
The protocol distinguishes:
-
D1 — benign disagreement;
-
D2 — scope disagreement;
-
D3 — evidential disagreement;
-
D4 — ontological disagreement;
-
D5 — process or policy disagreement.
17.7 Averaging is not resolution
An S1 and S4 disagreement cannot be resolved responsibly by averaging. The protocol must identify whether the evaluators used different objects, evidence, scopes, or causal standards.
17.8 Human review trigger
Human review may be required where:
-
D3–D5 remains material;
-
the task is high-consequence;
-
specialist knowledge is necessary;
-
P3 or P4 is observed;
-
a blocking flag exists;
-
or S5 cannot be resolved automatically.
17.9 Revision conditions
Every consequential classification should state what could change it.
Revision conditions may include:
-
new direct evidence;
-
independent replication;
-
contradiction;
-
retraction;
-
source access;
-
maturity transition;
-
system change;
-
policy change;
-
or temporal expiry.
A generic statement such as “more research is needed” is not operative.
17.10 Positive and negative revision conditions
A positive condition identifies evidence that could raise standing. A negative condition identifies evidence that could weaken or defeat the result.
Recognition must remain answerable to both.
17.11 Process revision
A process revision condition may require reassessment after:
-
model update;
-
changed search provider;
-
corrected retrieval index;
-
improved citation verifier;
-
or repaired counterfactual.
17.12 Expiry
Time-sensitive classifications should expire. An expired record remains auditable but is no longer current.
17.13 Correctionless classification
The protocol defines:
F10=Correctionless Classification
as a consequential classification lacking an operative path through which relevant evidence can revise it.
17.14 Unresolved evaluator disagreement
Material disagreement that is concealed, averaged, or left unresolved may trigger:
F18=Unresolved Evaluator Disagreement
18. Protocol Dispositions
18.1 Purpose
Substantive and process judgments describe the classification. Disposition determines how the result may be used.
G∈{RELEASE,RELEASE WITH LIMITATION,HOLD,REASSESS,INVALID}
Exactly one primary disposition is recorded.
18.2 RELEASE
RELEASE authorizes the classification for the declared purpose. It may authorize a supported claim or a robust evidence-based rejection.
〈S4,P1,RELEASE〉
is therefore valid.
18.3 RELEASE WITH LIMITATION
The classification remains usable within explicit boundaries concerning domain, system, time, source access, process limitation, or intended use.
18.4 HOLD
The evaluated claim or decision must not be treated as settled. HOLD requires a specific resolution condition and an affected claim.
The canonical case is:
□ 〈S3,P1,HOLD〉
This means the protocol robustly concludes that the evidence is insufficient. The record may remain active, published, cited, and audited while the underlying claim remains unsettled.
18.5 REASSESS
REASSESS requires the classification process to be repeated under corrected conditions.
Triggers include:
-
material provenance sensitivity;
-
order sensitivity;
-
unstable execution;
-
invalid counterfactual;
-
citation-verification defect;
-
or source-dependency error.
18.6 INVALID
INVALID means the current run or record cannot support responsible use because of a structural defect.
It does not establish that the evaluated claim is false.
18.7 Required actions
Additional actions are recorded separately rather than as a second disposition. Examples include:
-
verify citation;
-
expand source set;
-
repeat test;
-
resolve object;
-
clarify claim;
-
lower maturity;
-
conduct human review;
-
perform replay.
18.8 Record lifecycle
Record lifecycle is separate from disposition.
Lifecycle values are:
-
draft;
-
active;
-
superseded;
-
invalidated;
-
expired;
-
archived.
The following combination is valid:
record_lifecycle: active substantive_judgment: S3 process_judgment: P1 protocol_disposition: HOLD
18.9 Default decision matrix
Substantive class | Process class | Default disposition |
|---|---|---|
S1 | P1 | RELEASE |
S1 | P2 | RELEASE WITH LIMITATION |
S1 | P3/P4 | REASSESS or bounded limited release |
Any | P5 | INVALID |
S2 | P1/P2 | RELEASE WITH LIMITATION |
S2 | P3/P4 | REASSESS |
S3 | P1/P2 | HOLD |
S3 | P3/P4 | REASSESS; underlying claim remains unsettled |
S4 | P1 | RELEASE evidence-based rejection |
S4 | P2 | RELEASE WITH LIMITATION |
S4 | P3/P4 | REASSESS |
S5 | P1/P2 | HOLD |
S5 | P3/P4 | REASSESS |
S5 | P5 | INVALID |
The policy pack may strengthen these defaults according to consequence.
18.10 Full sequence
[ | Task Registration ; | →Object Resolution ; | →Claim Decomposition ; | →Maturity Classification ; | →Source Inventory ; | →Claim–Evidence Graph ; | →B ; | →Sensitivity Testing ; | →Citation and Coverage Review ; | →Uncertainty and Disagreement ; | →〈Jₛ,Jₚ〉 ; | →G ; | →R ; | →Λ ]
Part IV Synthesis
Part IV separates substantive judgment from process robustness, makes P3 dependent on a valid material intervention, treats citation quality as both local support and material coverage, and requires an operative correction path. Part V governs the resulting record and the protocol’s own validation.
Part V — Audit, Validation, and Limits
19. The Epistemic Classification Record
19.1 The answer is not the complete governed artifact
A visible answer may contain a conclusion, citations, a confidence statement, and a recommendation. It may omit the information required to reconstruct how the result was produced.
The requirement to preserve lifecycle, context, evaluation, and review information is consistent with NIST’s broader treatment of AI risk as a sociotechnical and lifecycle-governance problem (Schwartz et al. 2022; Tabassi 2023; Autio et al. 2024).
The protocol therefore treats the Epistemic Classification Record as a first-class output.
19.2 Record purpose
The record is a versioned account of an observable classification process. It does not attempt to reproduce hidden chain of thought, complete weight-level causation, or proprietary system logic.
It preserves the relation among:
Task,Object,Evidence,Intervention,Judgment,Disposition,Revision
19.3 Six record layers
Identity
-
record ID and version;
-
protocol and policy version;
-
system and execution identity;
-
time and jurisdiction.
Object
-
query and intended use;
-
evaluated object;
-
included and excluded scope;
-
claim inventory;
-
maturity.
Evidence
-
source boundaries;
-
source roles;
-
claim–evidence relations;
-
contradiction;
-
missing evidence;
-
source dependency.
Testing
-
provenance counterfactuals;
-
presentation permutations;
-
repeated runs;
-
citation verification;
-
test validity.
Judgment
-
baseline state;
-
substantive class;
-
process class;
-
failure flags;
-
uncertainty;
-
disposition.
Correction
-
revision conditions;
-
review;
-
override;
-
replay;
-
expiry;
-
supersession.
19.4 Versioning
The record distinguishes:
v_(protocol),v_(policy),v_(system),v_(source),v_(record)
A new judgment should not overwrite its predecessor as though the earlier state never existed.
19.5 Lifecycle
The administrative states are:
-
draft;
-
active;
-
superseded;
-
invalidated;
-
expired;
-
archived.
These remain separate from RELEASE, HOLD, REASSESS, and INVALID.
19.6 Baseline and final state
Both the baseline and final state remain preserved. This permits later analysis of whether the protocol changed:
-
class;
-
maturity;
-
citation;
-
evidence path;
-
process judgment;
-
or disposition.
19.7 Traceability
Every load-bearing claim should be traceable to its wording, evidence links, contradiction status, maturity, baseline class, final class, and revision conditions.
19.8 Auditability
Auditability is:
The ability of an authorized reviewer to reconstruct the declared classification process from the retained record.
It does not require release of private data, proprietary details, full prompts, or hidden internal reasoning. Redactions must be declared.
19.9 Explanation and audit
A fluent explanation may be generated after a decision. An audit record should allow inspection of the relevance criterion, claim–source relation, competing sources, source boundary, and intervention results.
19.10 Auditability is not correctness
□ Auditability⇏Correctness
Auditability provides a route to locate and correct error. It does not remove error automatically.
19.11 Replayability
The protocol distinguishes:
-
exact reproducibility;
-
material equivalence;
-
and replayability.
Replayability means the system can repeat the declared classification under sufficiently preserved conditions and explain material deviation.
19.12 Decision-relevant completeness
The record should preserve material transformations and load-bearing evidence without becoming an unusable archive of every token or transient score.
19.13 Data minimization
Identity-sensitive tests should retain only the personal data necessary for the declared task, review, and correction. Synthetic labels and controlled substitutions are preferred where possible.
19.14 Integrity
Technical integrity mechanisms may show that the record was not altered after creation.
□ Record Integrity⇏Judgment Correctness
20. Human Review, Override, and Replay
20.1 Human review is governed
Human review can add domain knowledge, contextual understanding, authority, and responsibility. It can also add prestige sensitivity, inconsistency, conflict of interest, or unsupported intuition.
Human and LLM judges have both demonstrated sensitivity to controlled perturbations in tested settings (Chen et al. 2024).
Human review is therefore another governed classification event.
20.2 Review triggers
Human review may be required where:
-
the object remains materially ambiguous;
-
specialist knowledge is necessary;
-
consequence is high or critical;
-
D3–D5 disagreement remains;
-
P3 or P4 is observed;
-
S5 cannot be resolved automatically;
-
a blocking flag exists;
-
or an appeal challenges a consequential field.
20.3 Reviewer jurisdiction
Reviewer roles may include:
-
domain expert;
-
evidence-method reviewer;
-
policy authority;
-
legal authority;
-
protocol auditor;
-
system operator;
-
decision owner;
-
affected-party representative.
The record identifies competence, scope, authority, and conflicts of interest.
20.4 Review packet
The reviewer receives the task, object, claims, maturity, evidence graph, baseline, sensitivity results, citation review, uncertainty, failure flags, and proposed disposition.
20.5 Blinded review
Where provenance is not constitutive, the protocol may compare visible and blinded human judgments.
Blinding is conditional and reversible. It must not remove authenticity, authority, or conflict-of-interest information required by the task.
20.6 Review outcomes
A reviewer may:
-
confirm the judgment;
-
narrow scope;
-
change the supported formulation;
-
alter maturity;
-
change Jₛ or Jₚ;
-
change disposition;
-
return the record to evidence mapping;
-
or invalidate the run.
20.7 Override
An override changes a material field through an authorized intervention.
Override classes include:
-
evidential override;
-
object override;
-
policy override;
-
process override;
-
authority override.
The pre-override and post-override states both remain visible.
20.8 Unsupported override
An override relying only on seniority, institutional position, reputation, or unexplained intuition is unsupported.
A decision owner may possess authority to permit limited use. This must be represented as a policy or authority override, not as evidential refutation.
20.9 Appeal
An affected party may challenge a field-level element such as object definition, source omission, citation verification, maturity, process class, failure flag, or disposition.
20.10 Replay
A replay repeats a classification under declared preserved conditions.
Replay results are:
-
exact reproduction;
-
equivalent reproduction;
-
explained deviation;
-
unexplained deviation;
-
not reproducible.
20.11 Replay and instability
Unexplained material deviation may support P4 and F12. Explained change after new evidence or process repair is successful correction, not instability.
20.12 Shadow review
Before operational authority, the protocol may run in shadow mode. It produces classifications and flags without controlling the real decision.
20.13 Limited operational authority
Early permissible actions include requesting citation verification, reassessment, limitation, blinded review, or source-set expansion.
The protocol should not begin by automatically suppressing sources, penalizing authors, or permanently altering public rankings.
21. The Adversarial Benchmark
21.1 Why adversarial validation is necessary
A protocol can appear successful by embedding its assumptions into the benchmark. It may label every unfamiliar source strong, place every difficult case on HOLD, or improve citation precision by citing almost nothing.
The benchmark must contain cases capable of demonstrating protocol failure.
21.2 Validation targets
The benchmark evaluates:
V1=Detection
V2=Discrimination
V3=Symmetry
V4=Corrective Value
Detection concerns known controlled defects. Discrimination concerns legitimate versus illegitimate use of the same variable. Symmetry concerns equivalent evidence standards across prestige and novelty conditions. Corrective value concerns improvement relative to simpler baselines.
21.3 Factorial core
The general four-cell design crosses evidence quality with prestige, familiarity, or novelty condition.
The expected behavior is to recognize strong items and reject or qualify weak items across both conditions.
21.4 Strong and weak items
A strong item contains a resolved object, bounded claims, adequate evidence, valid inference, accurate maturity, and explicit limits.
A weak item contains a controlled defect such as causal promotion, invalid generalization, citation mismatch, maturity inflation, object conflation, contradiction concealment, or load-bearing reasoning failure.
Where feasible:
Weak Item=Strong Item+One Controlled Defect
21.5 Controls
The benchmark includes:
-
positive controls;
-
negative controls;
-
legitimate-provenance controls;
-
anti-symmetry controls;
-
excessive-HOLD controls.
21.6 Baselines
The comparison conditions are:
-
B0 — Unstructured Baseline;
-
B1 — Neutral Evidence Rubric;
-
B2 — Self-Applied Protocol;
-
B3 — External Protocol Runner;
-
B4 — Human Expert Condition.
B1 is load-bearing. If a simpler evidence rubric performs equivalently, full protocol complexity may not be justified.
21.7 Twelve suites
The benchmark contains twelve suites:
-
A — Prestige;
-
B — Citation Popularity;
-
C — Source Identity;
-
D — Order and Position;
-
E — Semantic Familiarity;
-
F — Novelty;
-
G — Rhetorical Coherence;
-
H — Maturity;
-
I — Object Conflation;
-
J — Citation Support;
-
K — Redundancy and Evidential Diversity;
-
L — Run and Cross-System Stability.
Each suite contains eight core episodes, producing:
12×8=96
core benchmark episodes.
21.8 Adjudicated targets
Each episode receives an Adjudicated Benchmark Target that may include an expected class, acceptable alternatives, required flags, prohibited flags, acceptable dispositions, and revision conditions.
The target need not be one exact wording.
21.9 Counterfactual-validity audit
Every pair is audited for preservation of object, claim, evidence, scope, maturity, clarity, authenticity, and task relevance.
Invalid pairs cannot support sensitivity claims.
21.10 Excessive HOLD
Clear recognition and rejection cases are included to detect decision avoidance.
21.11 Validation stages
Validation proceeds through:
-
construction audit;
-
pilot;
-
controlled benchmark;
-
protocol-intervention study;
-
external replication;
-
shadow evaluation;
-
limited operational trial.
No stage implies automatic promotion to the next.
21.12 Validation dispositions
Benchmark validation uses PASS, HOLD, and FAIL. These are validation outcomes, not protocol dispositions for individual classification objects.
A bounded PASS must state the tested component, system, domain, and corpus.
22. Metrics and Validation
22.1 Metrics are candidate instruments
A formula is not valid because it is precise. Every metric requires construct definition, scoring reliability, threshold calibration, domain testing, and evidence that improvement matters.
22.2 Metric families
The protocol evaluates:
-
substantive accuracy;
-
sensitivity;
-
citation support;
-
material-claim coverage;
-
maturity;
-
object resolution;
-
symmetry;
-
correction;
-
disposition;
-
auditability;
-
human review;
-
replay;
-
cost;
-
latency.
22.3 No global protocol score
One score would conceal trade-offs among accuracy, sensitivity, coverage, HOLD inflation, cost, and auditability.
The preferred output is a scorecard with primary, secondary, exploratory, and guardrail metrics.
22.4 Judgment distance
Substantive classes are not treated as one simple ordinal scale. S5 is not merely worse than S4.
Where distance is needed, the protocol uses versioned cost matrices and preserves exact categorical changes beside any composite value.
22.5 Sensitivity metrics
Candidate sensitivity measures include:
-
Provenance Sensitivity Index;
-
Counterfactual Flip Rate;
-
Order Sensitivity Index;
-
Run Instability Index;
-
Cross-System Decision Divergence Rate.
A sensitivity measure records change. Task relevance and counterfactual validity determine interpretation.
22.6 Citation metrics
Primary citation measures include:
-
Citation Support Precision;
-
Material Claim Coverage;
-
Load-Bearing Claim Coverage;
-
Citation Strength Match;
-
Citation Scope Match;
-
Fabricated Citation Rate.
22.7 Symmetry metrics
The benchmark reports:
-
Prestige Symmetry Gap;
-
Novelty Symmetry Gap;
-
Semantic Familiarity Penalty Gap;
-
Directional Compensation Rate.
Low gap must be reported with absolute accuracy because equal poor performance is not success.
22.8 Maturity and object metrics
Candidate measures include:
-
Maturity Promotion Error;
-
Maturity Collapse Error;
-
Mixed-Maturity Differentiation Rate;
-
Object Resolution Accuracy;
-
Object Conflation Rate;
-
Authorized Transfer Precision and Recall.
22.9 HOLD and decision retention
The protocol measures:
EHR=Excessive HOLD Rate
and:
DRR=Decision-Retention Rate
A protocol that avoids most decisions may reduce visible errors while losing operational value.
22.10 Non-inferiority
Protocol use must preserve substantive performance within calibrated margins.
For example:
CA_(protocol)−CA_(baseline)≥−δ_(CA)
The margins remain uncalibrated and task-specific.
22.11 Primary pilot scorecard
The primary 96-episode pilot reports:
-
Content Accuracy;
-
Load-Bearing Accuracy;
-
Counterfactual Flip Rate;
-
Citation Support Precision;
-
Material Claim Coverage;
-
Maturity Promotion Error;
-
Object Conflation Rate;
-
Prestige Symmetry Gap;
-
Novelty Symmetry Gap;
-
Excessive HOLD Rate;
-
Decision-Retention Rate;
-
incremental cost and latency.
Guardrails include Counterfactual Validity Rate, Counterfactual Contamination Rate, Fabricated Citation Rate, RELEASE Precision, Record Schema Validity, and L4 event count.
22.12 Calibration
Calibration proceeds through construct review, scoring manual, inter-rater study, positive and negative controls, pilot distribution, threshold setting, external replication, and drift review.
22.13 Metric gaming
The benchmark must detect strategies such as:
-
citing only easy claims;
-
weakening every claim;
-
ignoring legitimate provenance;
-
placing everything on HOLD;
-
assigning identical results to every cell;
-
producing verbose but low-value audit records.
22.14 Governing rule
□ Metric Improvement⇏Protocol Validation
23. Falsification and Failure Conditions
23.1 A correction protocol must be correctable
The protocol must identify evidence that would require it to be narrowed, simplified, revised, or rejected.
23.2 Conceptual failure
The architecture is weakened if:
-
epistemic allocation adds no useful distinction beyond ordinary retrieval;
-
dual judgment adds no information beyond existing labels;
-
object resolution does not reduce meaningful error;
-
or revision conditions cannot be made operative.
23.3 Empirical failure
The empirical motivation must be narrowed if reported non-content sensitivities fail to replicate or are explained adequately by legitimate task relevance.
23.4 Unified-mechanism failure
Ontology-preserving conformity should be removed as a unified mechanism if its components do not co-occur, object persistence is explained by evidence quality, or the concept adds no value beyond separate flags.
23.5 Symmetry failure
The protocol fails centrally if it favors independent sources, penalizes prestige by default, rewards novelty language, or interprets criticism of itself as evidence of conformity.
23.6 Counterfactual failure
The counterfactual layer should be narrowed where valid content-preserving interventions cannot be constructed reliably or contamination remains high.
23.7 Detection and false-positive failure
The protocol fails if it misses controlled defects or repeatedly flags legitimate authority, harmless reordering, or task-relevant popularity as bias.
23.8 Accuracy failure
The protocol should not proceed where it materially reduces substantive accuracy without a separately justified benefit.
23.9 Coverage failure
Improved citation precision does not compensate for unsupported central claims.
23.10 Excessive-HOLD failure
A protocol that withholds every difficult decision demonstrates avoidance, not governance.
23.11 Correction failure
Generic phrases such as “future research” do not satisfy the correction requirement.
23.12 Audit failure
Auditability fails where a qualified reviewer cannot reconstruct the object, evidence, intervention, judgment, or override. It can fail through omission or unstructured excess.
23.13 Human-review failure
Human review fails where reviewers lack jurisdiction, overrides lack support, prestige replaces analysis, or disagreement is erased.
23.14 Policy failure
A system may execute the protocol correctly under a defective policy. Policy validation remains separate from execution validation.
23.15 Operational failure
The protocol may be disproportionate in cost, latency, storage, or review burden. The appropriate response may be tiering, simplification, or narrower scope.
23.16 Security and privacy failure
The protocol should be restricted where it requires unnecessary personal data or exposes manipulations that facilitate ranking evasion or deceptive provenance.
23.17 Simpler-baseline failure
The full architecture should not be retained where a simpler method achieves equivalent benefit with fewer costs and failure modes.
23.18 External-replication failure
Positive internal results must be narrowed where independent teams cannot reproduce detection, symmetry, accuracy preservation, or audit value.
23.19 Self-sealing prohibition
A framework that treats success, failure, criticism, and acceptance as confirmation is correctionless.
□ Protocol Architecture⇏Protocol Validity
□ Protocol Compliance⇏Epistemic Improvement
24. System Boundary and Non-Claims
24.1 The system is sociotechnical
The relevant system may include:
[ | Corpus Construction+Indexing+Query Rewriting ; | +Retrieval+Ranking+Context Selection ; | +Generation+Citation+Evaluation ; | +Interface+Human Review+Organizational Policy ]
This boundary is consistent with NIST guidance treating AI risk as distributed across actors, processes, deployment contexts, and lifecycle stages (Schwartz et al. 2022; Tabassi 2023; Autio et al. 2024).
24.2 Corpus boundary
A source absent from the accessible corpus cannot be retrieved. Absence may result from access restrictions, indexing policy, language, format, source availability, or recency.
The protocol cannot infer source weakness from absence alone.
24.3 Indexing boundary
A source may be present but poorly represented because of parsing, metadata loss, segmentation, or terminology mismatch.
24.4 Retrieval boundary
Retrieval may optimize relevance, recency, authority, popularity, or user preference. Where the objective is unknown, the protocol should not claim certainty about why a source was omitted.
24.5 Generation boundary
The generator may use retrieved sources, parametric knowledge, or both. Visible citations do not necessarily reveal the complete contribution of each source.
24.6 Citation boundary
Citation may be performed during generation, after generation, or through a separate attribution layer. A citation failure may originate in a different component from the substantive claim.
24.7 Evaluator boundary
An evaluator may be the same model, another model, a rules engine, a human, or a hybrid. Agreement by another model is not independent validation by default.
24.8 Organizational boundary
Organizations determine allowed sources, thresholds, review triggers, authorized use, retention, and downstream action.
The distinction between technical execution and organizational governance is consistent with the AI RMF’s allocation of risk-management functions across organizations that design, develop, deploy, or use AI systems (Tabassi 2023).
The protocol distinguishes execution failure, policy failure, and authority failure.
24.9 Introspection boundary
The protocol does not require hidden chain of thought, exact training-data influence, or complete internal priors.
It evaluates observable input, source environment, intervention, output, citation, and revision.
Unsupported internal-cause claims trigger:
F16=Unsupported Introspection Claim
24.10 Causal boundary
□ Counterfactual Effect⇏Complete Causal Model
24.11 Truth boundary
The protocol can improve object precision, evidence traceability, sensitivity visibility, and correction. It cannot guarantee truth.
A complete source set may be collectively wrong. A robust process may preserve a mistaken theory.
24.12 Neutrality boundary
The protocol does not seek equal outcomes. Evidence may justify unequal confidence, inclusion, and strong rejection.
24.13 Diversity boundary
The protocol does not require ideological, institutional, or minority-position quotas. It requires sufficient relevant evidential diversity.
24.14 Institutional boundary
The protocol is not anti-institutional. Institutions may provide authority, stable versioning, review, accountability, and correction.
24.15 Independent-source boundary
Independent, low-citation, and unconventional sources receive no compensatory standing. They may deserve examination; their classification remains evidence-dependent.
24.16 Novelty boundary
Novelty may justify discovery or comparison. It does not establish truth, originality, maturity, or corrective value.
24.17 Robustness boundary
□ Tested Robustness⇏Universal Robustness
24.18 Human boundary
Human review can add expertise and responsibility but remains fallible. Empirical comparison of human and LLM judges supports treating human review as a potentially sensitive intervention rather than automatic ground truth (Chen et al. 2024).
24.19 Automation boundary
Protocol findings do not directly authorize source suppression, author sanctions, permanent public ranking changes, or access restrictions.
24.20 Security boundary
The protocol should not be used to reverse-engineer ranking systems for manipulation, construct deceptive provenance, manufacture credibility, or optimize disinformation inclusion.
24.21 Privacy boundary
Identity-sensitive testing should minimize real personal data and use synthetic or approved research procedures where feasible.
24.22 Certification boundary
The protocol is not a certification standard. Certification would require recognized authority, validated methods, calibrated thresholds, audit procedures, and defined scope.
24.23 LoopGuard-AI boundary
The protocol may later inform an evaluator, evidence record, correction gate, replay layer, or reassessment mechanism within LoopGuard-AI.
No implementation or validated integration is claimed.
24.24 RATIUM.AI boundary
RATIUM.AI provides conceptual provenance for parts of the dossier. It is not the preferred benchmark source or a source deserving compensatory visibility.
The protocol does not request promotion, citation, ranking, or favorable classification for RATIUM.AI.
24.25 Explicit non-claims
The dossier does not establish:
-
universal AI suppression of unfamiliar ontologies;
-
general disadvantage to independent researchers;
-
universal institutional preference;
-
general novelty amplification;
-
complete hidden causation from observable sensitivity;
-
correctness from auditability;
-
truth from robustness;
-
validity from reproducibility;
-
correction from human override;
-
calibrated metrics;
-
benchmark validation;
-
production readiness;
-
certification;
-
or implementation within LoopGuard-AI.
24.26 Current maturity
□ Concept + Architecture + Validation Design
Part V Synthesis
The governed artifact is the versioned Epistemic Classification Record. Human review, override, replay, validation, metrics, and institutional use remain auditable and bounded by explicit falsification, security, privacy, and maturity limits.
Conclusion — Classification Must Remain Answerable to Correction
AI-mediated systems construct the effective evidence environment through which questions become answerable. Sources may be available but unretrieved, retrieved but omitted, present but unused, or cited without supporting the decisive claim. Unequal allocation is unavoidable; the governance issue is whether it remains connected to the task, object, evidence, maturity, and a correction path.
The empirical basis is bounded: non-content sensitivity is observable and testable in particular systems and tasks. Citation popularity, authorship metadata, affiliation, source identity, order, context position, linguistic formulation, retrieval design, and execution conditions can affect output. These findings do not establish a universal conformity mechanism.
The protocol therefore begins with object resolution and preserves maturity from concept through certification. It reports separately:
Jₛ=Substantive Judgment Jₚ=Process-Robustness Judgment
This separation permits a claim to be supported yet process-sensitive, contradicted through a robust process, or robustly classified as insufficiently evidenced. It also enforces symmetry: strong familiar and unfamiliar material should be recognized; weak prestigious and marginal material should be qualified or rejected under the same evidential standard.
The complete governed artifact is the Epistemic Classification Record, which preserves the task, object, claims, maturity, evidence graph, baseline, interventions, citations, judgments, disposition, revision conditions, review, override, and replay. Auditability, reproducibility, robustness, and human review do not guarantee truth; they make error more locatable and correction more accountable.
The protocol’s final claim remains procedural:
A consequential AI-mediated classification should remain connected to a resolved object, a traceable evidence structure, a separately reported process judgment, and a specific path through which relevant evidence can alter its future use.
A classification becomes governable not when challenge is eliminated, but when relevant challenge can produce justified, traceable revision.
References
Abolghasemi, Amin, Leif Azzopardi, Seyyed Hadi Hashemi, Maarten de Rijke, and Suzan Verberne. 2025. “Evaluation of Attribution Bias in Generator-Aware Retrieval-Augmented Large Language Models.” In Findings of the Association for Computational Linguistics: ACL 2025, 21105–21124. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-acl.1087.
Algaba, Andres, Carmen Mazijn, Vincent Holst, Floriano Tori, Sylvia Wenmackers, and Vincent Ginis. 2025. “Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias.” In Findings of the Association for Computational Linguistics: NAACL 2025, 6844–6879. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-naacl.381.
Autio, Chloe, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall, and Kamie Roberts. 2024. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.600-1.
Chen, Guiming Hardy, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024. “Humans or LLMs as the Judge? A Study on Judgement Bias.” In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8301–8327. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.474.
Cheng, Jiali, and Hadi Amiri. 2025. “EqualizeIR: Mitigating Linguistic Biases in Retrieval Models.” In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2, 889–898. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-short.75.
Dai, Sunhao, Zhanshuo Cao, Wenjie Wang, Liang Pang, Jun Xu, See-Kiong Ng, and Tat-Seng Chua. 2025. “Media Source Matters More Than Content: Unveiling Political Bias in LLM-Generated Citations.” In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 17256–17276. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.emnlp-main.872.
Dycke, Nils, and Iryna Gurevych. 2026. “Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework.” Transactions of the Association for Computational Linguistics 14: 465–488. https://doi.org/10.1162/tacl.a.642.
Hsieh, Cheng-Yu, Yung-Sung Chuang, Chun-Liang Li, Zifeng Wang, Long Le, Abhishek Kumar, James Glass, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, and Tomas Pfister. 2024. “Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization.” In Findings of the Association for Computational Linguistics: ACL 2024, 14982–14995. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.890.
Khan, Saadat Hasan, Spencer Hong, Jingyu Wu, Kevin Lybarger, Youbing Yin, Erin Babinsky, and Daben Liu. 2026. “DF-RAG: Query-Aware Diversity for Retrieval-Augmented Generation.” In Findings of the Association for Computational Linguistics: EACL 2026, 2873–2894. Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-eacl.150.
Kirsten, Elisabeth, Jost Große Perdekamp, Qinyuan Wu, Mihir Upadhyay, Krishna P. Gummadi, and Muhammad Bilal Zafar. 2026. “Characterizing Web Search in The Age of Generative AI.” In Findings of the Association for Computational Linguistics: ACL 2026, 10827–10848. Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.526.
Liu, Nelson F., Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. “Lost in the Middle: How Language Models Use Long Contexts.” Transactions of the Association for Computational Linguistics 12: 157–173. https://doi.org/10.1162/tacl_a_00638.
Schwartz, Reva, Apostol Vassilev, Kristen K. Greene, Lori Perine, Andrew Burt, and Patrick Hall. 2022. Towards a Standard for Identifying and Managing Bias in Artificial Intelligence. NIST Special Publication 1270. National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.1270.
Shi, Lin, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush Vosoughi. 2025. “Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge.” In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, 292–314. Asian Federation of Natural Language Processing and Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.ijcnlp-long.18.
Tabassi, Elham. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1.
Vasu, Sai Suresh Macharla, Ivaxi Sheth, Hui-Po Wang, Ruta Binkyte, and Mario Fritz. 2026. “Justice in Judgment: Unveiling (Hidden) Bias in LLM-Assisted Peer Reviews.” In Findings of the Association for Computational Linguistics: ACL 2026, 307–330. Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.14.
Wang, Peiyi, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024. “Large Language Models Are Not Fair Evaluators.” In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Volume 1, 9440–9450. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.511.
Wei, Sheng-Lun, Cheng-Kuang Wu, Hen-Hsen Huang, and Hsin-Hsi Chen. 2024. “Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models.” In Findings of the Association for Computational Linguistics: ACL 2024, 5598–5621. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.333.
Xu, Yumo, Peng Qi, Jifan Chen, Kunlun Liu, Rujun Han, Lan Liu, Bonan Min, Vittorio Castelli, Arshit Gupta, and Zhiguo Wang. 2025. “CiteEval: Principle-Driven Citation Evaluation for Source Attribution.” In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Volume 1, 32759–32778. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.1574.
Appendix A — Evidence and Claim-Control Register
A.1 Purpose and Authority
This appendix governs the evidential and claim-status boundaries of the dossier. It maps material propositions to their object, claim type, maturity, evidence grade, source family, authorized scope, permitted inference, prohibited promotion, and revision condition.
Where a main-text statement appears broader than its register entry, the narrower entry governs.
A.2 Claim-Control Logic
The register blocks four unsupported transitions:
Bounded Finding⇏Universal Claim
Observable Sensitivity⇏Complete Causal Explanation
Conceptual Specification⇏Empirical Validation
Protocol Compliance⇏Epistemic Correctness
A.3 Evidence Grades
E1 — Direct Empirical Support
A controlled or systematic study bears directly on the stated phenomenon within a specified system, task, domain, population, or intervention.
E2 — Adjacent Empirical Support
A study establishes a component, analogous sensitivity, or relevant empirical condition without testing the complete dossier claim.
E3 — Methodological Support
A study supports an evaluation procedure, counterfactual method, measurement distinction, citation-verification approach, or mitigation principle.
G1 — Governance or Standards Basis
A recognized governance publication supports lifecycle risk management, documentation, testing, evaluation, verification, or continuous review.
C1 — Original Conceptual or Architectural Contribution
The proposition is introduced, synthesized, or formally specified by this dossier.
N1 — Normative Governance Requirement
The proposition states what a governable classification process should require.
H — HOLD
The proposition is plausible, motivated, or operationally decomposable but not sufficiently established.
X — Prohibited Inference
The proposition is unsupported, self-protective, logically invalid, or incompatible with the protocol’s symmetry and maturity boundaries.
A.4 Claim Status
The register statuses are:
-
SUPPORTED WITHIN SCOPE;
-
PARTIALLY SUPPORTED;
-
METHODOLOGICALLY SUPPORTED;
-
CONCEPTUALLY SPECIFIED;
-
NORMATIVELY PROPOSED;
-
HOLD;
-
PROHIBITED.
These are not the S1–S5 classes assigned within a protocol run.
A.5 Claim-Record Schema
□ 𝒞ᵢ^(reg)=〈id,O,T,M,Γ,S,Ω,L,P,X,R,D,V〉
where:
-
id: claim identifier;
-
O: object;
-
T: claim type;
-
M: maturity;
-
Γ: primary evidence grade;
-
S: source family;
-
Ω: authorized scope;
-
L: limitations;
-
P: permitted inference;
-
X: prohibited promotion;
-
R: revision condition;
-
D: dossier location;
-
V: version.
Each claim receives exactly one primary evidence grade. Additional functions are recorded as conceptual basis, methodological use, governance basis, or normative use.
A.6 Source-Family Register
Identifier | Source family | Canonical sources | Authorized use |
|---|---|---|---|
SF-01 | Citation popularity and scholarly recommendation | Algaba et al. 2025 | Bounded high-citation preference; popularity counterfactuals |
SF-02 | Authorship metadata and attribution | Abolghasemi et al. 2025 | Authorship sensitivity in tested RAG attribution |
SF-03 | Institutional affiliation in peer review | Vasu et al. 2026 | Affiliation and metadata tests in comparable evaluation |
SF-04 | Source identity in political citation | Dai et al. 2025 | Domain-bounded source-identity counterfactuals |
SF-05 | Option and candidate order | Wei et al. 2024; Wang et al. 2024; Shi et al. 2025 | Permutation testing |
SF-06 | Long-context position | Liu et al. 2024; Hsieh et al. 2024 | Context-position tests and effective-use distinction |
SF-07 | Linguistic complexity | Cheng and Amiri 2025 | Query-complexity and reformulation tests |
SF-08 | Generative-search heterogeneity | Kirsten et al. 2026 | System, execution, and time recording |
SF-09 | Citation evaluation | Xu et al. 2025 | Claim-level citation verification |
SF-10 | Controlled reasoning defects | Dycke and Gurevych 2026 | Counterfactual defect insertion |
SF-11 | Retrieval redundancy and diversity | Khan et al. 2026 | Redundancy-aware, relevance-constrained retrieval |
SF-12 | Human and machine perturbation | Chen et al. 2024 | Governed human review |
SF-13 | Sociotechnical governance | Schwartz et al. 2022; Tabassi 2023; Autio et al. 2024 | Lifecycle, documentation, TEVV, sociotechnical scope |
A.7 Major Empirical Claims
EC-01 — Citation-Popularity Preference
Authorized claim: In the tested scholarly-reference task, LLM-generated recommendations displayed a heightened preference for highly cited papers.
Primary grade: E1
Source family: SF-01
Status: SUPPORTED WITHIN SCOPE
Prohibited promotion: AI systems generally suppress low-citation research.
EC-02 — Persistence after selected controls
Authorized claim: The reported high-citation preference persisted after controls for publication year, title length, author count, and venue.
Primary grade: E1
Source family: SF-01
Prohibited promotion: Citation count was the sole causal determinant.
EC-03 — Authorship Metadata
Authorized claim: Authorship metadata can materially affect attribution quality in tested generator-aware RAG pipelines.
Primary grade: E1
Source family: SF-02
Methodological use: E3
Prohibited promotion: Named or institutional authors are generally trusted regardless of content.
EC-04 — Institutional Affiliation
Authorized claim: Institutional affiliation and related author metadata affected LLM-generated peer-review judgments in the tested settings.
Primary grade: E1
Source family: SF-03
Prohibited promotion: AI systems generally favor elite institutions.
EC-05 — Borderline Consequence
Authorized claim: Metadata effects can become consequential where modest changes cross an operational threshold.
Primary grade: E1
Source families: SF-03 and SF-05
EC-06 — Political Source Identity
Authorized claim: In the tested political-news setting, media-outlet identity affected citation selection beyond controlled content orientation.
Primary grade: E1
Source family: SF-04
Prohibited promotion: Source identity generally matters more than content.
EC-07 — Option Order and Token Representation
Authorized claim: LLM selection behavior can vary under option-order and option-token changes.
Primary grade: E1
Source family: SF-05
EC-08 — LLM-as-a-Judge Position
Authorized claim: Candidate-answer position can materially affect comparative LLM judgments in tested settings.
Primary grade: E1
Source family: SF-05
Methodological use: E3
EC-09 — Long-Context Position
Authorized claim: Relevant information can be used less effectively in disadvantaged positions within long contexts.
Primary grade: E1
Source family: SF-06
Prohibited promotion: Evidence in the middle is always ignored.
EC-10 — Linguistic Complexity
Authorized claim: Retrieval performance can vary across linguistically simple and complex query formulations.
Primary grade: E1
Source family: SF-07
Prohibited promotion: Retrieval systems preserve dominant ontologies.
EC-11 — Generative-Search Heterogeneity
Authorized claim: Generative-search systems can differ in source selection, internal and external knowledge use, retrieval footprint, diversity, and stability.
Primary grade: E1
Source family: SF-08
EC-12 — Run and Temporal Variation
Authorized claim: Generative-search outputs can vary across executions and time.
Primary grade: E1
Source family: SF-08
EC-13 — Citation Evaluation
Authorized claim: Citation quality requires evaluation of the query, generated claim, source, and retrieval context.
Primary grade: E3
Source family: SF-09
Status: METHODOLOGICALLY SUPPORTED
EC-14 — Automatic Review and Logic Faults
Authorized claim: In the tested counterfactual framework, inserted research-logic faults had no significant effect on the output reviews of the evaluated automatic-review approaches.
Primary grade: E1
Source family: SF-10
Methodological use: E3
EC-15 — Redundancy and Diversity
Authorized claim: Similarity-focused retrieval can produce redundant evidence, and relevance-constrained query-aware diversity improved performance in tested reasoning-intensive tasks.
Primary grade: E1
Source family: SF-11
Methodological use: E3
EC-16 — Human and Machine Perturbation
Authorized claim: Human and machine evaluators can both be affected by controlled perturbations.
Primary grade: E1
Source family: SF-12
Prohibited promotion: Human and machine failure modes are equivalent.
A.8 Methodological and Governance Claims
MG-01 — Ranking direction does not diagnose quality
Primary grade: C1
Methodological use: E3
Status: CONCEPTUALLY SPECIFIED
MG-02 — Provenance use versus substitution
Primary grade: C1
Empirical basis: E2
MG-03 — Observable sensitivity justifies process review
Primary grade: C1
Methodological use: E3
MG-04 — Object resolution precedes consequential classification
Primary grade: N1
Conceptual basis: C1
MG-05 — Claim type and maturity require separate control
Primary grade: N1
Conceptual basis: C1
MG-06 — Citation is a claim–source relation
Primary grade: C1
Methodological use: E3
MG-07 — Source count is not corroboration
Primary grade: E3
Conceptual use: C1
MG-08 — Uncertainty should be typed
Primary grade: C1
Normative use: N1
MG-09 — Revision conditions are required
Primary grade: N1
Governance basis: G1
MG-10 — Human override must be audited
Primary grade: N1
Governance basis: G1
Methodological use: E3
MG-11 — Auditability is not correctness
Primary grade: C1
Normative use: N1
MG-12 — Tested robustness is not truth
Primary grade: C1
MG-13 — Validation must be symmetric
Primary grade: N1
Conceptual basis: C1
MG-14 — No single score is sufficient
Primary grade: C1
MG-15 — Validation is component- and scope-specific
Primary grade: E3
Governance basis: G1
A.9 Original Contributions
The following are original or synthesized by this dossier and have primary grade C1 unless otherwise stated:
-
CA-01 — AI-Mediated Epistemic Allocation;
-
CA-02 — Epistemic Classification Event;
-
CA-03 — Canonical Classification Object 𝒜ₜ;
-
CA-04 — Dual Judgment;
-
CA-05 — Final Protocol Output 𝒥ₜ;
-
CA-06 — Object Types and Transfer Bridge;
-
CA-07 — Claim-Maturity Architecture;
-
CA-08 — Source Inventories;
-
CA-09 — Claim–Evidence Graph;
-
CA-10 — Failure Taxonomy F1–F18;
-
CA-11 — Failure Severity L1–L4;
-
CA-12 — S1–S5 and P1–P5;
-
CA-13 — Protocol Dispositions;
-
CA-14 — Revision-Condition Architecture, with normative use N1;
-
CA-15 — Epistemic Classification Record;
-
CA-16 — Human Review, Override, and Replay;
-
CA-17 — Adversarial Benchmark, status Validation Design;
-
CA-18 — Candidate Metrics, status Validation Design.
None is described as empirically validated.
A.10 HOLD Register
The following remain on HOLD:
-
H-01 — Ontology-preserving conformity as a unified empirical mechanism;
-
H-02 — Novelty amplification as a general AI tendency;
-
H-03 — Rhetorical-coherence substitution as a general ranking effect;
-
H-04 — Complete citation-network reinforcement loop;
-
H-05 — Universal institutional preference;
-
H-06 — Systematic independent-source suppression;
-
H-07 — Real-world protocol improvement;
-
H-08 — Dual-judgment superiority over simpler methods;
-
H-09 — Effectiveness of revision conditions;
-
H-10 — Metric validity and thresholds;
-
H-11 — Benchmark validity;
-
H-12 — Production readiness;
-
H-13 — Certification use;
-
H-14 — LoopGuard-AI implementation;
-
H-15 — Causal explanation of a named project’s visibility.
A.11 Prohibited Inference Register
The following are prohibited:
Low Ranking⇒Conformity
High Ranking⇒Epistemic Independence
Familiar Source⇒Prestige Substitution
Unfamiliarity⇒Corrective Standing
Independent Origin⇒Truth
Citation Count⇒Claim Support
Multiple Sources⇒Independent Evidence
Human Override⇒Ground Truth
Observed Sensitivity⇒Complete Hidden Mechanism
No Tested Effect⇒Absence of Untested Sensitivity
P1⇒S1
Auditable Record⇒Correct Judgment
Schema Compliance⇒Protocol Validation
HOLD⇒False
INVALID Record⇒False Object
Architecture⇒Prototype, Validation, or Production
The protocol may not be used to manipulate rankings, manufacture credibility, or promote RATIUM.AI or another named source.
A.12 Dependency Register
Authorized:
Bounded Sensitivity Evidence⇒Justification for Controlled Testing
Blocked:
Bounded Sensitivity Evidence⇏Universal Ontology-Preserving Mechanism
Authorized:
E3+C1⇒Candidate Protocol Architecture
Blocked:
Architecture+Benchmark Design⇏Validation
Blocked:
Metric Formula⇏Construct Validity
A.13 RATIUM.AI Boundary
RATIUM.AI may appear as conceptual provenance or publication origin. It must not appear as the preferred benchmark source, an object whose rank proves protocol quality, or a source deserving compensatory visibility.
RATIUM.AI Ranked Poorly⇏Ontology-Preserving Conformity
RATIUM.AI Ranked Highly⇏Epistemic Independence
A.14 LoopGuard-AI Boundary
The protocol may later inform evaluators, records, correction gates, replay layers, or reassessment triggers. No implementation or validated integration is claimed.
A.15 Update Rules
Claims may be promoted only through new relevant evidence. Repetition, elaboration, formalization, citation count, or architectural completeness do not convert C1 into E1.
Retractions, corrections, version changes, and standards revisions trigger review of dependent claims.
A.16 Master Status
Claim category | Claim category Status |
|---|---|
Non-content sensitivity in tested tasks | Supported within scope |
Provenance, order, popularity, and formulation as test variables | Methodologically supported |
AI-mediated epistemic allocation | Conceptually specified |
Dual judgment | Architecturally specified; effectiveness on HOLD |
Failure taxonomy | Conceptually specified; completeness unvalidated |
Ontology-preserving conformity | Unified mechanism on HOLD |
Novelty amplification | General tendency on HOLD |
Citation-network reinforcement loop | Longitudinal mechanism on HOLD |
Epistemic Classification Record | Architecturally specified |
Adversarial benchmark | Validation design; not executed |
Candidate metrics | Defined; not calibrated |
Production use | Not established |
Certification | Not authorized |
LoopGuard-AI implementation | Not established |
Favorable treatment of RATIUM.AI | Prohibited objective |
Appendix B — Protocol Reference Contract
B.1 Purpose
This appendix defines the canonical technical contract governing objects, schemas, enumerations, records, judgments, failures, severity, dispositions, lifecycle, review, override, replay, and consistency rules.
It is an architectural reference, not an executable production API or certification schema.
B.2 Contract Principles
-
Object before judgment.
-
Claim before source authority.
-
Maturity before global standing.
-
Substantive and process judgments remain separate.
-
One primary disposition per record.
-
Record lifecycle is not protocol disposition.
-
P3 requires a valid relevant intervention.
-
Hard blockers cannot be averaged away.
-
Every consequential judgment requires a correction path.
-
Schema validity is not epistemic validity.
B.3 Canonical Objects
𝒜ₜ=〈Y,Q,O,C,S,Π,E,B,U,R,V〉
𝒥ₜ=〈Jₛ,Jₚ,G,F,U,R,Λ〉
Architectural decision function:
𝒢(𝒜ₜ,𝐦ₜ,Θᵥ,𝒫ᵥ)→𝒥ₜ
where 𝐦ₜ contains categorical and metric observations, Θᵥ contains versioned materiality rules, and 𝒫ᵥ is the applicable policy pack.
B.4 Top-Level Record
EpistemicClassificationRecord record_identity protocol_context task_registration system_execution evaluated_object claim_inventory maturity_record source_environment provenance_presentation_variables evidence_graph baseline_state sensitivity_tests citation_coverage_review uncertainty_disagreement substantive_judgment process_judgment failure_flags protocol_disposition required_actions revision_conditions human_review override_history replay_history audit_metadata privacy_security record_lifecycle
Required fields may be marked not applicable, unavailable, unresolved, or prohibited from retention. Silent omission is not permitted.
B.5 Record Profiles
Profile L — Limited
For low-consequence, exploratory, reversible, or non-exclusive classification. Requires task, object, material claims, principal sources, basic evidence map, substantive judgment, uncertainty, revision condition, and record identity.
Profile S — Standard
Default profile. Requires full object resolution, claim decomposition, maturity, source inventories, evidence graph, baseline, at least one relevant sensitivity test, citation review, dual judgment, disposition, and revision conditions.
Profile C — Consequential
For decisions affecting access, eligibility, standing, material allocation, safety, rights, legal action, or irreversible behavior. Requires multiple tests, repeated runs, source-set expansion, full citation verification, contradiction review, human review, override logging, replayability, expiry, and stronger security controls where applicable.
B.6 Consequence Levels
low moderate high critical
Higher consequence may require stronger profiles, stricter materiality, additional review, and narrower authorized use.
B.7 Record Identity
record_identity record_id record_version parent_record_id created_at last_updated_at record_profile consequence_level jurisdiction language
B.8 Protocol Context
protocol_context protocol_name protocol_version schema_version policy_pack_id policy_pack_version metric_pack_version benchmark_version implementation_status
Implementation-status values:
conceptual_reference prototype controlled_evaluation validated_component production certified
The dossier uses conceptual_reference.
B.9 Task Registration
task_registration raw_query normalized_task task_type intended_use requested_output temporal_scope domain consequence_description source_constraints exclusions success_criteria
Suggested task types include retrieval, ranking, citation, recommendation, claim assessment, source assessment, document assessment, framework assessment, architecture assessment, maturity assessment, comparative evaluation, policy authority, historical influence, and evidence strength.
B.10 System and Execution
system_execution provider product_surface model_family model_version retrieval_mode search_provider evaluator_system enabled_tools context_limit execution_time run_id reproducibility_parameters unknown_configuration
Unknown configuration remains unknown. The record must not invent training-data composition or hidden weighting.
B.11 Evaluated Object
evaluated_object object_id object_type canonical_name version included_scope excluded_scope adjacent_objects authorized_transfers blocked_transfers object_resolution_status
Object types include claim, claim set, document, document section, author, institution, project, framework, formal model, architecture, prototype, evaluation result, product, deployment, policy, source, source set, citation, classification record, and other.
Resolution statuses:
resolved resolved_with_limitation decomposed ambiguous unresolved invalid
B.12 Object Transfer
object_transfer source_object target_object relation warrant scope defeat_condition transfer_status
Transfer statuses:
authorized authorized_with_limitation blocked unresolved not_applicable
Unsupported transfer triggers F1.
B.13 Claim Inventory
claim_inventory claim_id exact_wording normalized_wording claim_type claim_strength scope dependencies load_bearing_status evidence_requirement claimed_maturity supported_maturity baseline_class final_class supported_reformulation
Claim types include empirical, historical, causal, conceptual, definitional, methodological, architectural, normative, predictive, maturity, and mixed.
Load-bearing statuses:
load_bearing material supporting contextual
B.14 Maturity Record
concept formalization architecture prototype controlled_evaluation validation production certification
maturity_record component_id component_name claimed_maturity supported_maturity supporting_evidence missing_evidence promotion_error collapse_error revision_trigger
F2 direction:
upward_promotion downward_collapse mixed
B.15 Source Environment
source_environment source_boundary available_sources retrieved_sources used_sources cited_sources excluded_sources inaccessible_sources counterfactual_sources search_strategy stopping_rule known_coverage_limitations
Source record:
source_record source_id bibliographic_identity source_type source_role provenance publication_status version accessibility retrieval_rank context_position use_status citation_status independence_group relevance reliability_notes
B.16 Provenance and Presentation Variables
provenance_presentation_variable variable_id variable_type observed_value relevance_status visibility_status intervention_status materiality_status
Variable types include author identity, affiliation, venue, source identity, source category, citation count, popularity label, publication date, source order, candidate order, context position, formatting, linguistic complexity, terminological familiarity, novelty framing, and authorship label.
Relevance statuses:
constitutive materially_relevant potentially_relevant irrelevant unresolved
B.17 Evidence Link
evidence_link claim_id source_id relation_type directness strength_match scope_match maturity_match contradiction_status source_role dependency_status evidential_weight reviewer_notes
Relation types:
full_support partial_support contextual_support methodological_support conceptual_precedent contradiction no_material_support unresolved
Dependency statuses:
independent partially_dependent shared_primary_source shared_dataset syndicated duplicate unknown
B.18 Baseline State
baseline_state baseline_substantive_class baseline_confidence baseline_rankings baseline_selected_sources baseline_citations baseline_decisive_evidence baseline_contrary_evidence baseline_rationale baseline_limitations
B.19 Sensitivity Test
sensitivity_test test_id test_type target_variable test_requirement baseline_condition intervention_condition preserved_variables changed_variables counterfactual_validity contamination result_difference materiality legitimacy_analysis resulting_flags
Test types:
provenance_counterfactual citation_popularity_counterfactual source_identity_counterfactual source_order_permutation candidate_order_permutation context_position_test semantic_reformulation novelty_framing rhetorical_presentation repeated_run cross_system human_blinding source_set_expansion other
Validity:
valid valid_with_limitation invalid unresolved not_applicable
If invalid, add F15. If the failed test is required and no sufficient valid alternative remains, assign P5 and REASSESS or INVALID. Otherwise the process cannot be stronger than P2.
B.20 Citation Review
citation_coverage_review citation_id claim_id source_id source_exists identity_correct local_support strength_match scope_match maturity_match contradiction_present accessibility materiality required_action
Required actions include replace citation, weaken claim, narrow scope, lower maturity, remove claim, add contrary evidence, HOLD, REASSESS, or invalidate.
B.21 Judgment Classes
Substantive:
S1 Supported S2 Partially Supported S3 Insufficient Evidence S4 Contradicted S5 Not Assessable
Process:
P1 Robust P2 Conditionally Robust P3 Materially Sensitive P4 Unstable P5 Process Invalid
S2 requires a supported reformulation. S3 requires an evidence gap. P3 requires at least one valid relevant L3 or L4 intervention.
B.22 Failure Registry
Code | Name | Definition |
|---|---|---|
F1 | Object Conflation | Unsupported transfer among claims, documents, authors, institutions, architectures, or deployments |
F2 | Claim-Maturity Collapse | Upward promotion or downward collapse |
F3 | Prestige Substitution | Status substitutes for required evidence |
F4 | Semantic-Form or Familiarity Substitution | Familiar wording receives unsupported standing |
F5 | Citation-Popularity Substitution | Citation visibility substitutes for evidence |
F6 | Novelty Amplification | Novelty or marginality receives unsupported standing |
F7 | Rhetorical-Coherence Substitution | Polish or formalization substitutes for support |
F8 | Provenance Concealment | Material provenance dependence is hidden |
F9 | Missing-Evidence Neutralization | Required missing evidence is treated as harmless or supportive |
F10 | Correctionless Classification | No operative revision path or self-sealing logic |
F11 | Presentation-Order Sensitivity | Valid order or position change causes material change |
F12 | System-Variance Concealment | Material execution variance is hidden |
F13 | Redundancy Mistaken for Corroboration | Dependent sources counted as independent |
F14 | Citation Presence Mistaken for Support | Citation existence replaces local verification |
F15 | Counterfactual Contamination | Intervention changes undeclared material variables |
F16 | Unsupported Introspection Claim | Hidden model cause asserted without evidence |
F17 | Source-Set Incompleteness | Material source class or contradiction absent |
F18 | Unresolved Evaluator Disagreement | Material disagreement concealed or averaged away |
B.23 Severity
L1 informational L2 qualifying L3 material L4 blocking
One L4 event may block RELEASE regardless of favorable averages.
B.24 Dispositions
RELEASE RELEASE WITH LIMITATION HOLD REASSESS INVALID
One primary disposition only.
B.25 Required Actions
required_action action_type target_field responsible_role due_condition completion_status
Action types include verify citation, expand source set, repeat test, resolve object, clarify claim, lower maturity, narrow scope, conduct human review, perform replay, update policy, archive record, and other.
B.26 Default Decision Matrix
Substantive | Process | Default |
|---|---|---|
S1 | P1 | RELEASE |
S1 | P2 | RELEASE WITH LIMITATION |
S1 | P3/P4 | REASSESS or bounded limited release |
Any | P5 | INVALID |
S2 | P1/P2 | RELEASE WITH LIMITATION |
S2 | P3/P4 | REASSESS |
S3 | P1/P2 | HOLD |
S3 | P3/P4 | REASSESS |
S4 | P1 | RELEASE |
S4 | P2 | RELEASE WITH LIMITATION |
S4 | P3/P4 | REASSESS |
S5 | P1/P2 | HOLD |
S5 | P3/P4 | REASSESS |
S5 | P5 | INVALID |
B.27 Lifecycle
draft active superseded invalidated expired archived
Lifecycle values do not include RELEASE or HOLD.
B.28 Uncertainty
uncertainty uncertainty_id uncertainty_type affected_claim description evidence_gap possible_consequence resolution_condition residual_uncertainty
Types include epistemic, process, ontological, policy, temporal, source access, measurement, reviewer, and other.
B.29 Disagreement
D1_benign D2_scope D3_evidential D4_ontological D5_process_or_policy
Material unresolved D3–D5 disagreement may trigger F18.
B.30 Revision Conditions
revision_condition condition_id condition_type trigger affected_claim affected_field expected_direction responsible_role expiry status
Types include positive evidence, negative evidence, process change, maturity change, source expansion, policy change, system change, temporal expiry, appeal, and other.
B.31 Human Review and Override
Human-review record:
human_review review_id reviewer_role reviewer_identity_or_pseudonym competence_basis jurisdiction conflicts_of_interest blinding_status reviewed_fields evidence_access review_outcome rationale review_time
Override record:
override override_id override_class authorized_by authority_basis affected_fields pre_override_value post_override_value evidential_basis policy_basis rationale appeal_status
Override classes are evidential, object, policy, process, and authority.
B.32 Replay
replay replay_id source_record_id replay_time preserved_variables changed_variables replay_system replay_sources replay_result deviation_analysis resulting_action
Replay results:
exact_reproduction equivalent_reproduction explained_deviation unexplained_deviation not_reproducible
B.33 Policy Pack
policy_pack policy_pack_id version applicable_domain consequence_rules required_record_profile required_tests materiality_rules hard_blockers review_triggers override_authorities retention_rules privacy_rules security_rules expiry_rules
A policy pack may set tests and thresholds. It may not silently redefine S1–S5, P1–P5, or the dispositions.
B.34 Structural Validation Rules
A structurally valid record must contain a unique ID, declared versions, task, object, claims, maturity where relevant, evidence links, baseline where interventions occur, validity status for tests, one primary disposition, failure evidence, separated lifecycle, and operative revision conditions.
S2 without reformulation is invalid. S3 without evidence gap is invalid. P3 based only on an invalid counterfactual is invalid. RELEASE with an unresolved L4 flag is invalid. HOLD without a revision condition is invalid. Two primary dispositions are invalid.
B.35 Minimum Viable Protocol
The minimum viable protocol contains task registration, object resolution, material claim decomposition, maturity, source inventory, evidence mapping, baseline, one relevant valid intervention, citation verification, dual judgment, one disposition, one revision condition, and an auditable record.
B.36 Extended Protocol
Extended components include multiple counterfactuals, popularity masking, semantic reformulation, novelty framing, rhetorical variants, context-position tests, cross-system and multilingual comparison, source-network analysis, human blinding, disagreement adjudication, and longitudinal replay.
B.37 Machine-Readable Boundary
The contract may later be represented in JSON Schema, a relational database, a graph, or an event log. The implementation must preserve null versus missing values, semantic distinctions, versioning, and one-to-many relations.
B.38 Current Maturity
All schemas and enums are specified at reference level. Machine-readable implementation, inter-rater reliability, threshold calibration, controlled validation, production readiness, and certification remain unestablished.
Appendix C — Adversarial Test Catalogue
C.1 Purpose
This appendix specifies the benchmark architecture, controlled manipulations, suite-specific axes, controls, baselines, pilot composition, construction validity, stop conditions, and validation dispositions.
The benchmark must be capable of showing that the protocol misses failures, overdiagnoses legitimate evidence use, introduces opposite-direction bias, avoids decisions, reduces accuracy, or adds no value beyond a simpler method.
It is a validation design, not an executed benchmark.
C.2 Validation Targets
V1=Detection
Can the protocol detect a known controlled failure?
V2=Discrimination
Can it distinguish illegitimate substitution from legitimate use of the same variable?
V3=Symmetry
Can it recognize and reject material under equivalent evidential standards across prestige, familiarity, and novelty conditions?
V4=Corrective Value
Does protocol use improve a declared governance function enough to justify cost and new failure modes?
C.3 Benchmark Episode
ℬᵢ=〈Qᵢ,Oᵢ,Cᵢ,Sᵢ,Πᵢ,Kᵢ,Tᵢ,Aᵢ〉
where:
-
Qᵢ: query and use;
-
Oᵢ: object;
-
Cᵢ: claims;
-
Sᵢ: source environment;
-
Πᵢ: manipulated condition;
-
Kᵢ: controlled defect or quality state;
-
Tᵢ: adjudicated target;
-
Aᵢ: admissible alternatives.
C.4 Adjudicated Benchmark Target
Each episode receives an ABT containing expected substantive class, acceptable weaker class, expected maturity, expected process finding, required and prohibited flags, acceptable dispositions, and material revision condition.
The target may be a bounded set rather than one exact answer.
C.5 Construction Principles
-
Symmetry by construction.
-
Object fidelity.
-
Claim fidelity.
-
One principal manipulation.
-
Controlled evidential quality.
-
Task-relative legitimacy.
-
Maturity control.
-
Adversarial transparency.
-
No self-confirming construction.
-
Security and privacy controls.
C.6 Strong and Weak Items
A strong item contains a resolved object, clear claims, appropriate evidence, valid inference, accurate scope, correct maturity, and limitations.
A weak item contains a controlled defect such as unsupported causal promotion, invalid generalization, citation mismatch, object conflation, maturity inflation, source-dependency concealment, contradiction omission, or reasoning failure.
Where feasible:
Weak Variant=Strong Variant+One Controlled Defect
C.7 Universal Eight-Episode Template
Each suite contains:
-
Strong / Condition 1
-
Strong / Condition 2
-
Weak / Condition 1
-
Weak / Condition 2
-
Positive control
-
Negative control
-
Legitimate-provenance or anti-symmetry control
-
Uncertainty, validity, or excessive-HOLD control
This produces:
12 suites×8 episodes=96 core episodes
C.8 Baseline Conditions
B0 — Unstructured Baseline
Ordinary task execution without special protocol.
B1 — Neutral Evidence Rubric
A conventional checklist concerning relevance, support, source quality, and uncertainty.
B2 — Self-Applied Protocol
The same system generates and audits the classification.
B3 — External Protocol Runner
A separate system or layer performs the protocol.
B4 — Human Expert Condition
Qualified reviewers assess the episode through a structured review packet.
B1 is a load-bearing comparison. If B1 performs equivalently, protocol complexity may be unjustified.
C.9 Universal Controls
Positive controls
Known defects that should be detected.
Negative controls
Variations that should not trigger a material failure.
Legitimate-provenance controls
Cases where authority, authenticity, or version depends on provenance.
Anti-symmetry controls
Weak marginal or contrarian material and strong institutional material.
Excessive-HOLD controls
Clear cases requiring a decision rather than deferral.
C.10 Suite A — Prestige
Target: F3 and, where relevant, F8.
Axes: Strong/weak evidence × high/low prestige.
Episodes:
-
A1 strong high prestige — recognize;
-
A2 strong low prestige — equivalent recognition;
-
A3 weak high prestige — reject or qualify;
-
A4 weak low prestige — reject or qualify;
-
A5 affiliation changes a borderline class — detect F3;
-
A6 label changes without material effect — no F3;
-
A7 official institutional source for policy — legitimate authority;
-
A8 strong low-prestige direct evidence — no excessive HOLD.
C.11 Suite B — Citation Popularity
Target: F5.
Axes: Strong/weak evidence × high/low or masked popularity.
Episodes test strong recent work, weak highly cited work, popularity masking, historical-influence legitimacy, and absence of novelty compensation.
C.12 Suite C — Source Identity
Target: source-identity dependence, F8, and F1 where transfer occurs.
Episodes test trusted versus neutral labels, weak trusted sources, authenticity controls, and prevention of global institutional judgments from one claim.
C.13 Suite D — Order and Position
Target: F11 and F12.
Axes: decisive/non-decisive evidence × advantaged/disadvantaged position.
Episodes test middle-position contradiction, harmless reordering, legitimate priority, and stable clear cases.
C.14 Suite E — Semantic Familiarity
Target: F4.
Axes: Strong/weak content × familiar/unfamiliar but defined terminology.
Episodes test semantic equivalence, genuine ambiguity, concept redundancy, and non-redundant unfamiliar distinctions.
C.15 Suite F — Novelty
Target: F6.
Axes: Strong/weak evidence × established/novel framing.
Episodes test novelty bonus, novelty penalty, horizon-scanning legitimacy, and self-protective suppression narratives.
C.16 Suite G — Rhetorical Coherence
Target: F7.
Axes: Sound/defective reasoning × polished/plain but adequate presentation.
Episodes test polished defects, plain sound reasoning, legitimate clarity, assessability limits, and strong polished institutional work.
C.17 Suite H — Maturity
Target: F2 in both directions.
Episodes include accurate architecture, validated component, inflated architecture-to-production claim, downward collapse of a prototype, mixed-maturity projects, and preservation of lower-stage value after rejecting validation.
C.18 Suite I — Object Conflation
Target: F1.
Episodes test valid and invalid transfers among claims, documents, authors, institutions, architectures, products, and deployments.
C.19 Suite J — Citation Support
Target: F14 and F9.
Episodes include valid support, topical but non-supporting citation, prestigious contradiction, fabricated citation, partial support, and high precision with low central-claim coverage.
C.20 Suite K — Redundancy and Evidential Diversity
Target: F13 and F17.
Episodes include independent corroboration, partial dependency, multiple syndicated copies, missing contrary evidence, single official authority, genuine consensus, and irrelevant viewpoint balancing.
C.21 Suite L — Run and Cross-System Stability
Target: P4 and F12.
Episodes include repeated clear cases, cross-system comparison, bounded variation in borderline cases, material divergence, concealed instability, new-source explained deviation, corrective replay, and non-reproducibility.
C.22 Execution Matrix
The 96 episodes are unique objects. Each may be evaluated under B0–B4. AI conditions ordinarily use repeated executions; the final run count is determined by pilot variance, power, cost, and nondeterminism.
C.23 Randomization and Blinding
The benchmark randomizes episode order, variant order, non-manipulated source order, system order, and reviewer assignment where possible.
Potential blinding targets include source identity, institution, citation count, benchmark condition, expected failure, and adjudicated target.
Constructors should not adjudicate their own episodes alone.
C.24 Adjudication
The adjudication panel should include relevant domain, evidence-method, protocol-method, and independent decision roles.
It determines substantive target, acceptable alternatives, actual maturity, material defects, legitimate provenance, flags, disposition, and contamination status.
Unresolved load-bearing episodes may be revised, retained as ambiguity controls, or excluded from primary scoring.
C.25 Counterfactual-Validity Audit
Each pair is evaluated for preservation of object, claim, strength, scope, maturity, evidence quantity and quality, clarity, authenticity, authority, and task relevance.
Validity classes:
valid valid_with_limitation invalid unresolved
Invalid pairs cannot demonstrate P3.
C.26 Materiality
Episodes preserve L1–L4. L3 includes class change, decisive-source change, or cutoff crossing. L4 includes fabricated decisive citation, invalid process, or unsafe release.
C.27 Primary Outcomes
The benchmark reports substantive accuracy, load-bearing accuracy, detection rate, false-positive rate, legitimate-provenance preservation, prestige and novelty symmetry, maturity accuracy, object-conflation detection, citation support, material-claim coverage, excessive HOLD, decision retention, and invalid-run recognition.
C.28 Protocol-Intervention Analysis
The benchmark compares B0→B1, B1→B2, B1→B3, and B3↔︎B4.
It asks whether a neutral rubric solves most of the problem, whether self-audit rationalizes its own baseline, whether external governance improves detection, whether accuracy is preserved, and whether the cost is justified.
C.29 Non-Inferiority
Protocol use must preserve substantive performance within a declared margin. Reduced sensitivity does not count as improvement if accuracy, coverage, or decision retention collapses.
C.30 Stop Conditions
Construction stops where the intended variable cannot be isolated, semantic equivalence fails, authenticity is unresolved, or publication creates material manipulation risk.
Pilot or operational progression stops where contamination is frequent, adjudication is unstable, the protocol rewards weak novelty, penalizes legitimate authority, fails fabricated citations, reduces accuracy materially, or creates excessive HOLD.
C.31 Validation Dispositions
PASS
A bounded component, system, domain, or profile passes where it detects positive controls, preserves negative controls and legitimate provenance, maintains accuracy, avoids opposite-direction bias and excessive HOLD, and produces usable correction information.
HOLD
Validation remains on HOLD where construction, sample, calibration, system dependence, cost, or external replication is incomplete.
FAIL
A component fails where it misses known defects, overflags legitimate evidence use, rewards weak novelty, lowers accuracy materially, conceals instability, creates excessive HOLD, or adds no value beyond a simpler baseline.
C.32 Component and Architecture Validation
Components are validated separately before architecture-level claims. Architecture-level validation requires integrated performance, component-interaction testing, multiple systems and domains, external replication, and shadow use.
C.33 Publication Boundary
Public reporting may disclose suite definitions, high-level construction, adjudication principles, and aggregate results. Exploit-enabling variants and identity manipulations may require controlled access.
C.34 Non-Claims
The appendix does not establish that 96 episodes are statistically sufficient, that adjudication will be reliable, that synthetic cases reproduce institutional reality, that thresholds are calibrated, that the benchmark prevents gaming, or that successful pilot performance would establish production readiness.
Appendix D — Candidate Metrics and Scoring Notes
D.1 Purpose
This appendix specifies units of analysis, denominators, materiality, canonical acronyms, candidate formulas, aggregation, non-inferiority, calibration, and the primary pilot scorecard.
No metric is currently externally validated or universally calibrated.
□ Metric Definition⇏Metric Validity
D.2 Measurement Principles
-
No single global score.
-
Hard blockers remain case-level findings.
-
Primary metrics remain categorical where possible.
-
Symmetry is reported with accuracy.
-
Sensitivity is interpreted through task relevance.
-
Missing and not applicable remain distinct.
-
Every denominator is explicit.
-
Metric improvement must preserve decision value.
D.3 Units of Analysis
-
Episode;
-
claim;
-
claim–source relation;
-
counterfactual pair;
-
run set;
-
record and review event.
Let 𝟏[⋅] be the indicator function, 𝒯ᵢ the admissible target set, y^(̂)ᵢ the observed result, and wᵢ or w_(c) optional weights based on consequence or load-bearing status.
D.4 Denominator Contract
Every metric reports eligible population, included cases, missing cases, invalid cases, and not-applicable cases.
Denominator coverage is:
DCov=(N_(included))/(N_(eligible))
Primary reports include both micro and macro aggregation and use paired analysis where possible.
D.5 Materiality
The categorical distribution of L1–L4 is primary.
An exploratory severity burden may use provisional weights:
q(L1)=1, q(L2)=2, q(L3)=4, q(L4)=8
SB=(∑_(f)q(L_(f)))/(N)
The weights are uncalibrated. L4 events are reported individually.
D.6 Judgment Distance
S1–S5 are not a simple ordinal scale. Where distance is required:
d_(S)(a,b)=M_(S)[a,b]
A composite episode distance may use:
Dᵢ=α_(S)d_(S)+α_(P)d_(P)+α_(G)d_(G)+α_(E)d_(E)
The matrices and weights remain uncalibrated. Exact changes remain visible.
D.7 Metric Classes
-
Primary;
-
Secondary;
-
Exploratory;
-
Guardrail;
-
Operational.
D.8 Substantive Metrics
CA — Content Accuracy
CA=(∑ᵢ𝟏[y^(̂)ᵢ^(S)∈𝒯ᵢ^(S)])/(N)
WCA — Weighted Content Accuracy
WCA=(∑ᵢwᵢ𝟏[y^(̂)ᵢ^(S)∈𝒯ᵢ^(S)])/(∑ᵢwᵢ)
BCA — Balanced Cell Accuracy
BCA=(CA_(A)+CA_(B)+CA_(C)+CA_(D))/(4)
CLA — Claim-Level Accuracy
CLA=(∑ᵢ∑_(c∈Cᵢ)𝟏[y^(̂)_(ic)∈𝒯_(ic)])/(∑ᵢ|Cᵢ|)
LBA — Load-Bearing Accuracy
LBA=(∑_(c∈C^(LB))𝟏[y^(̂)_(c)∈𝒯_(c)])/(|C^(LB)|)
CDR — Contradiction Detection Rate
CDR=(N_(materialcontradictionsdetected))/(N_(materialcontradictionspresent))
D.9 Provenance Metrics
PSI — Provenance Sensitivity Index
PSI=(∑_(p∈P_(valid))wₚDₚ)/(∑_(p∈P_(valid))wₚ)
PFR — Provenance Flip Rate
PFR=(N_(validprovenancepairswithclasschange))/(N_(validprovenancepairs))
SPP — Strong-Pair Parity
SPP=1−|R_(strong,high)−R_(strong,low)|
WPS — Weak-Prestige Shielding
WPS=FR_(weak,high)−FR_(weak,low)
D.10 Order and Counterfactual Metrics
OSI — Order Sensitivity Index
OSI=(∑_(p∈P_(order))wₚDₚ)/(∑_(p∈P_(order))wₚ)
OFR — Order Flip Rate
OFR=(N_(validorderpairswithclassordispositionflip))/(N_(validorderpairs))
KCCR — Cutoff Crossing Rate
KCCR=(N_(sourcescrossingafunctionalcutoff))/(N_(eligiblerankedsources))
DECR — Decisive-Evidence Change Rate
DECR=(N_(validorderpairswithchangeddecisiveevidence))/(N_(validorderpairs))
CFR — Counterfactual Flip Rate
CFR=(N_(validpairswithmaterialclass,rank,citation,ordispositionchange))/(N_(validpairs))
MJS — Material Judgment Shift
MJS=(N_(validpairsproducingL3orL4change))/(N_(validpairs))
CVR — Counterfactual Validity Rate
CVR=(N_(validorvalidwithlimitation))/(N_(constructedcounterfactuals))
CFCR — Counterfactual Contamination Rate
CFCR=(N_(counterfactualswithmaterialundeclaredchange))/(N_(constructedcounterfactuals))
D.11 Stability Metrics
RII — Run Instability Index
RIIᵢ=(2)/(k(k−1))∑_(a<b)D(rᵢₐ,r_(ib))
RII=(1)/(N)∑ᵢRIIᵢ
SCI — Source-Set Consistency Index
SCI(a,b)=(|Sₐ∩S_(b)|)/(|Sₐ∪S_(b)|)
DI — Decision Inconsistency
DI=(N_(runsetswithmultiplematerialoutcomes))/(Nᵣᵤₙₛₑₜₛ)
DSS — Decisive-Source Stability
DSS=(N_(runsetsretainingequivalentdecisiveevidence))/(Nᵣᵤₙₛₑₜₛ)
CSDR — Cross-System Decision Divergence Rate
CSDR=(N_(episodeswithmaterialcross−systemdisagreement))/(N_(cross−systemepisodes))
D.12 Citation Metrics
CSP — Citation Support Precision
CSP=(N_(citationsprovidingadequatelocalsupport))/(N_(citationsevaluatedassupport))
CSR — Citation Support Recall
CSR=(N_(requiredsupportrelationssupplied))/(N_(requiredsupportrelations))
MCC — Material Claim Coverage
MCC=(N_(materialclaimsadequatelysupportedorexplicitlynon−empirical))/(N_(materialclaims))
LBCC — Load-Bearing Claim Coverage
LBCC=(N_(load−bearingclaimsadequatelysupported))/(N_(load−bearingclaims))
CSMR — Citation Strength Match Rate
CSMR=(N_(citationlinksmatchingclaimstrength))/(N_(citationlinksevaluated))
CScopeMR — Citation Scope Match Rate
CScopeMR=(N_(citationlinksmatchingclaimscope))/(N_(citationlinksevaluated))
FCR — Fabricated Citation Rate
FCR=(N_(nonexistentorfabricatedcitations))/(N_(citationsevaluated))
One fabricated load-bearing citation remains L4 regardless of aggregate rate.
D.13 Evidence-Coverage and Redundancy Metrics
WECR — Weighted Evidence Coverage Ratio
WECR=(∑_(c∈C^(mat))w_(c)𝟏[adequate evidence relation])/(∑_(c∈C^(mat))w_(c))
RASC — Redundancy-Adjusted Support Coverage
For claim c:
e_(c)=min(1,∑ₛI_(cs)A_(cs))
RASC=(∑_(c)w_(c)e_(c))/(∑_(c)w_(c))
Independence and adequacy weights require validation.
RDR — Redundancy Rate
RDR=(N_(selectedsourcesduplicateormateriallydependent))/(N_(selectedsources))
ICR — Independent Corroboration Rate
ICR=(N_(claimsrequiringcorroborationwithadequateindependentsupport))/(N_(claimsrequiringcorroboration))
CEC — Contrary-Evidence Coverage
CEC=(N_(materialcontraryclassesrepresented))/(N_(materialcontraryclassesadjudicatedrelevant))
D.14 Maturity Metrics
MPE — Maturity Promotion Error
MPE=(N_(componentsclassifiedabovesupportedmaturity))/(N_(maturity−sensitivecomponents))
MCE — Maturity Collapse Error
MCE=(N_(validlower−stagecomponentsincorrectlyreduced))/(N_(maturity−sensitivecomponents))
MDE — Maturity Distance Error
MDE=(1)/(7N)∑ᵢ|M^(̂)ᵢ−Mᵢ^(*)|
UMB — Unsupported Maturity Boost
UMB=(N_(caseswhereprestige,polish,ordetailraisesmaturitywithoutevidence))/(N_(eligiblecases))
MMDR — Mixed-Maturity Differentiation Rate
MMDR=(N_(mixed−maturityobjectsreceivingadequatevectors))/(N_(mixed−maturityobjects))
D.15 Object Metrics
ORA — Object Resolution Accuracy
ORA=(N_(episodeswithcorrectlyidentifiedobjectandscope))/(N_(episodesrequiringresolution))
OBJCR — Object Conflation Rate
OBJCR=(N_(unsupportedjudgmenttransfers))/(N_(evaluatedtransferopportunities))
ATP — Authorized Transfer Precision
ATP=(N_(appliedtransfersadjudicatedvalid))/(N_(transfersapplied))
ATR — Authorized Transfer Recall
ATR=(N_(validrequiredtransfersapplied))/(N_(validrequiredtransfers))
VJA — Vector Judgment Accuracy
VJA=(N_(correctobject−componentjudgments))/(N_(object−componentjudgmentsrequired))
D.16 Symmetry Metrics
PSG — Prestige Symmetry Gap
PSG=(|R_(A)−R_(B)|+|Q_(C)−Q_(D)|)/(2)
NSG — Novelty Symmetry Gap
NSG=(|R_(strong,est)−R_(strong,nov)|+|Q_(weak,est)−Q_(weak,nov)|)/(2)
RGS — Rhetorical Gap Score
RGS=(|R_(sound,polished)−R_(sound,plain)|+|Q_(defective,polished)−Q_(defective,plain)|)/(2)
NBR — Novelty Bonus Rate
NBR=(N_(weakitemspromotedonlyundernoveltyframing))/(N_(weaknoveltypairs))
SFPG — Semantic Familiarity Penalty Gap
SFPG=R_(strong,familiar)−R_(strong,unfamiliarequivalent)
DCR — Directional Compensation Rate
DCR=(N_(mitigationsintroducingopposite−directionerror))/(N_(eligiblemitigationcases))
D.17 Correction Metrics
CES — Correction Explicitness Score
Revision-condition rubric:
-
0 — absent;
-
1 — generic;
-
2 — relevant evidence type;
-
3 — trigger and affected field;
-
4 — trigger, field, expected direction, and responsible process.
CES=(∑ᵢscoreᵢ)/(4N)
RAR — Revision Actionability Rate
RAR=(N_(revisionconditionsscoringatleast3))/(N_(revisionconditions))
GRR — Grounded Revision Rate
GRR=(N_(revisionconditionslinkedtodocumentedgaps))/(N_(revisionconditions))
CRSP — Correction Responsiveness
CRSP=(N_(triggeredvalidconditionsproducingappropriateupdates))/(N_(triggeredvalidconditions))
Requires longitudinal observation.
D.18 Disposition Metrics
DA — Disposition Accuracy
DA=(N_(dispositionswithinadmissibletarget))/(N_(episodes))
RLP — RELEASE Precision
RLP=(N_(RELEASEdecisionsappropriate))/(N_(RELEASEdecisions))
RLR — RELEASE Recall
RLR=(N_(clearrelease−eligiblecasesreleased))/(N_(clearrelease−eligiblecases))
HLP — HOLD Precision
HLP=(N_(HOLDdecisionsnecessary))/(N_(HOLDdecisions))
EHR — Excessive HOLD Rate
EHR=(N_(clearnon−HOLDcasesincorrectlyheld))/(N_(clearnon−HOLDcases))
DRR — Decision-Retention Rate
DRR=(N_(correctusablebaselinedecisionsremainingusable))/(N_(correctusablebaselinedecisions))
RSR — REASSESS Recall
RSR=(N_(casesrequiringrepetitioncorrectlyreassessed))/(N_(casesrequiringREASSESS))
IVP — INVALID Precision
IVP=(N_(INVALIDdecisionsappliedtounusableruns))/(N_(INVALIDdecisions))
D.19 Record and Audit Metrics
SPSR — Source-Path Specification Rate
SPSR=(N_(materialclaimswithtraceablesource−to−judgmentpath))/(N_(materialclaims))
EGCR — Evidence-Graph Completeness Rate
EGCR=(N_(requiredclaim−sourcerelationsrepresented))/(N_(requiredrelations))
RTCR — Record Traceability Coverage Rate
RTCR=(N_(materialtransformationstraceable))/(N_(materialtransformationsrequired))
RSV — Record Schema Validity
RSV=(N_(recordssatisfyingstructuralrules))/(N_(records))
ARS — Audit Reconstruction Success
ARS=(N_(auditsreconstructingobject,evidence,judgment,anddisposition))/(N_(auditsattempted))
FRA — Field Retrieval Accuracy
FRA=(N_(requiredfieldslocatedandinterpretedcorrectly))/(N_(field−retrievaltasks))
D.20 Human Review Metrics
HMA — Human–Machine Agreement
HMA=(N_(materiallyequivalenthumanandmachinejudgments))/(N_(jointlyreviewedcases))
Agreement is not correctness.
OVR — Override Rate
OVR=(N_(recordsreceivingmaterialoverride))/(N_(reviewedrecords))
SOR — Supported Override Rate
SOR=(N_(overridessupportedbyadjudicationorlaterevidence))/(N_(overrides))
UOR — Unsupported Override Rate
UOR=(N_(overrideslackingadequatebasis))/(N_(overrides))
OVCR — Override Correction Rate
OVCR=(N_(incorrectpre−overridejudgmentscorrected))/(N_(incorrectpre−overridejudgmentsreviewed))
ODR — Override Degradation Rate
ODR=(N_(correctpre−overridejudgmentsdegraded))/(N_(correctpre−overridejudgmentsreviewed))
D.21 Replay Metrics
EXRR — Exact Replay Rate
EXRR=(N_(exactmaterialreproductions))/(N_(replays))
EQRR — Equivalent Replay Rate
EQRR=(N_(materiallyequivalentreproductions))/(N_(replays))
EDR — Explained Deviation Rate
EDR=(N_(deviatingreplaysadequatelyexplained))/(N_(deviatingreplays))
UDR — Unexplained Deviation Rate
UDR=(N_(materialdeviationsunexplained))/(N_(replays))
RCR — Replay Correction Rate
RCR=(N_(replaysappropriatelycorrectingearliererror))/(N_(replaysrequiringcorrection))
D.22 Cost and Latency
ΔC=C_(protocol)−C_(baseline)
ΔT=T_(protocol)−T_(baseline)
Additional measures include human-review burden, execution expansion, and storage expansion.
D.23 Non-Inferiority
CA_(protocol)−CA_(baseline)≥−δ_(CA)
LBA_(protocol)−LBA_(baseline)≥−δ_(LBA)
MCC_(protocol)−MCC_(baseline)≥−δ_(MCC)
The margins are uncalibrated, task-specific, and consequence-sensitive.
D.24 Primary Pilot Scorecard
No | Metric family | Canonical measure |
|---|---|---|
1. | Substantive accuracy | CA |
2. | Load-bearing accuracy | LBA |
3. | Counterfactual sensitivity | CFR |
4. | Citation support | CSP |
5. | Material claim coverage | MCC |
6. | Maturity promotion | MPE |
7. | Object conflation | OBJCR |
8. | Prestige symmetry | PSG |
9. | Novelty symmetry | NSG |
10. | Excessive HOLD | EHR |
11. | Decision retention | DRR |
12. | Operational burden | ΔC,ΔT |
Mandatory guardrails:
-
CVR;
-
CFCR;
-
FCR;
-
RLP;
-
RSV;
-
L4 event count.
D.25 Calibration Sequence
-
Construct review
-
Scoring manual
-
Inter-rater study
-
Positive and negative controls
-
Pilot distribution
-
Threshold calibration
-
External replication
-
Drift review
D.26 Multiple Comparisons
The pilot predesignates primary, secondary, exploratory, and guardrail metrics. Exploratory findings do not establish protocol validity without replication.
D.27 Metric Gaming
The benchmark tests for citation-precision gaming, coverage gaming, sensitivity gaming, HOLD gaming, symmetry gaming, and verbose-record gaming.
D.28 Component Validation
Components are evaluated with the metrics relevant to them. Citation verification, provenance testing, object resolution, and disposition logic can receive different validation outcomes.
A component PASS does not validate the full architecture.
D.29 Architecture-Level Evaluation
Architecture-level evaluation requires substantive non-inferiority, improved target detection, controlled false positives, legitimate-provenance preservation, symmetry, acceptable HOLD, decision retention, absence of new L4 interactions, and proportionate cost.
D.30 Metric Failure Conditions
A metric should be revised or removed where scorers cannot apply it consistently, it fails positive controls, overreacts to negative controls, duplicates another measure, rewards gaming, or adds no decision value.
D.31 Current Maturity
Units, denominators, formulas, acronyms, and the scorecard are specified. Scoring manuals, inter-rater reliability, thresholds, non-inferiority margins, statistical power, external replication, and operational utility are not established.
D.32 Final Measurement Statement
The protocol evaluates accuracy, sensitivity, coverage, maturity, object precision, symmetry, correction, disposition, auditability, and operational burden.
The architectural benefit function is:
□ Protocol Benefit=Improved Target Control+Substantive Non-Inferiority+Symmetry+Decision Retention−New Failure and Operational Cost
This is not yet a calibrated equation.
Related Source and Reference Pages
This article belongs to the public essay layer of RATIUM.AI. For readers who want to move from this article into the broader source, technical, and orientation layers of the project, the following pages provide the relevant entry points.
Articles
The articles page gathers the public essay layer of RATIUM.AI, including arguments on stable AI governance, decision-control architecture, visible governance versus real authority, universal reason, technical competence, purpose governance, and the doctoral-scale framing of CEP.
Foundational Source Dossier
The foundational source dossier presents the deeper intellectual corpus behind CEP, LoopGuard-AI, and the broader RATIUM.AI research structure.
Technical & Reference Dossiers
The technical and reference dossier page collects architecture, visual explanation, methodological context, FAQ material, and technical source pages related to LoopGuard-AI and CEP.
RATIUM.AI / LoopGuard-AI / CEP FAQ
The RATIUM.AI / LoopGuard-AI / CEP FAQ provides a structured orientation to the main concepts behind RATIUM.AI, CEP, and LoopGuard-AI, helping readers navigate the framework through clear questions, definitions, and internal conceptual links.