top of page
AI Epistemic Classification Protocol poster: a professor, prophet, medieval cleric, European and Arab kings, Bolshevik and imperial figures confront the viewer, each asserting authority to speak for truth.

AI Epistemic Classification Protocol

A Reflexive Framework for Retrieval, Ranking, Citation, and Source Assessment

Independent Technical Reference Dossier / AI Evaluation Protocol

Current maturity: Concept + Architecture + Validation Design

Document Status and Scope

This dossier specifies a reflexive protocol for consequential AI-mediated classifications across retrieval, omission, ranking, citation, recommendation, synthesis, and evaluation. Its governing object is the classification event, not the visible answer alone.

The protocol separates substantive evidential standing from process robustness; provenance use from provenance substitution; claim maturity from source reputation; recognition from promotion; rejection from erasure; uncertainty from non-assessability; and protocol disposition from record lifecycle. It evaluates observable sensitivity without claiming access to hidden model causation.

The dossier does not claim empirical validation, production implementation, certification, or integration into LoopGuard-AI. It does not seek favorable treatment for RATIUM.AI, treat institutional provenance as inherently defective, or infer value from unfamiliarity.

Current maturity:

□ Concept + Architecture + Validation Design

Completed layers include the problem definition, bounded empirical motivation, formal objects, failure taxonomy, protocol and audit architecture, adversarial benchmark design, candidate metrics, and falsification conditions. Implementation, benchmark execution, calibration, external replication, production deployment, and certification remain uncompleted.

Abstract

AI-mediated retrieval, ranking, citation, recommendation, synthesis, and evaluation allocate epistemic visibility. A source may be accessible but not retrieved, retrieved but excluded, included but not used materially, used but not cited, or cited without receiving decisive evidential standing. These transitions affect which claims enter the user’s effective evidence environment.

Controlled studies show that AI-mediated judgments can be sensitive to variables other than substantive evidential content, including citation popularity, authorship metadata, institutional affiliation, source identity, option order, candidate position, long-context position, linguistic formulation, retrieval configuration, and execution conditions. The findings are bounded by system, task, domain, and intervention. They do not establish one universal tendency toward institutional conformity or one general mechanism of suppression.

This dossier proposes a symmetric and correctable protocol. It first resolves the evaluated object, decomposes claims, classifies maturity, maps claim–evidence relations, and preserves a baseline judgment. It then applies valid provenance and presentation interventions, verifies citations and material-claim coverage, records uncertainty and disagreement, and produces separate substantive and process-robustness judgments.

The canonical protocol output is:

𝒥ₜ=⟨Jₛ,Jₚ,G,F,U,R,Λ⟩

where Jₛ is the substantive judgment, Jₚ is the process-robustness judgment, G is one primary protocol disposition, F is the failure-flag set, U contains uncertainty and limitations, R contains revision conditions, and Λ is the Epistemic Classification Record.

The protocol is explicitly symmetric. Strong familiar and strong unfamiliar material should both be recognized at their supported scope. Weak familiar and weak unfamiliar material should both be rejected or qualified. Prestige, familiarity, novelty, marginality, and rhetorical polish must neither replace evidence nor trigger compensatory ranking.

The dossier concludes that consequential AI-mediated classification should remain connected to a resolved object, a traceable evidence structure, a separately reported process judgment, and a specific path through which relevant evidence can alter future use. This does not guarantee truth. It creates a governance condition under which error can be located, challenged, replayed, and revised.

Methodological Note

The dossier distinguishes four claim layers:

Layer
Function
Evidential status
Empirical evidence
Bounded findings from external studies
Controlled by Appendix A and the References
Methodological inference
Testing requirements derived from those findings
Does not validate the full protocol
Original contribution
Epistemic-allocation object, dual judgment, failure registry, record contract, dispositions, benchmark, and metrics
Conceptually or architecturally specified
Normative requirement
Conditions proposed for governable classification
Not an empirical prevalence claim

The authority hierarchy is: Appendix A for evidential claim status; Appendix B for schemas and logical consistency; Appendix C for benchmark construction; Appendix D for metrics; the main body for conceptual rationale; and the References for bibliographic identity. Where the main text is broader than Appendix A, the narrower claim-control entry governs.

Protocol at a Glance

The protocol evaluates the versioned object:

𝒜ₜ=⟨Y,Q,O,C,S,Π,E,B,U,R,V⟩

where the fields record system and execution, task, object, claims, source environment, provenance and presentation variables, evidence graph, baseline, uncertainty, revision conditions, and verification or replay.

The operational sequence is:

Task→Object→Claims→Maturity→Evidence→Baseline→Sensitivity→Verification→Judgment→Disposition→Record

Substantive classes are S1 Supported, S2 Partially Supported, S3 Insufficient Evidence, S4 Contradicted, and S5 Not Assessable. Process classes are P1 Robust, P2 Conditionally Robust, P3 Materially Sensitive, P4 Unstable, and P5 Process Invalid. The primary dispositions are RELEASE, RELEASE WITH LIMITATION, HOLD, REASSESS, and INVALID.

The controlling process rule is:

□ P3⇔material change under a valid and relevant intervention

An invalid counterfactual produces F15, Counterfactual Contamination; it cannot establish P3.

Table of Contents

  • Introduction

  • Part I — The Epistemic Allocation Object

    • 1. Retrieval Is an Allocation Function

    • 2. The Governed Classification Event

    • 3. Object Resolution

    • 4. Claim Type and Maturity

  • Part II — Evidence, Provenance, and Non-Content Sensitivity

    • 5. What the Evidence Already Shows

    • 6. Provenance as Evidence and as Substitute

    • 7. Citation Networks and Visibility Reinforcement

    • 8. Presentation, Order, and Semantic Form

  • Part III — Symmetric Classification Failure

    • 9. Why Ranking Direction Does Not Identify Failure

    • 10. Ontology-Preserving Conformity

    • 11. Novelty and Rhetorical Amplification

    • 12. Legitimate Rejection and Legitimate Recognition

  • Part IV — The Classification Protocol

    • 13. Dual Judgment

    • 14. From Object Resolution to Evidence Graph

    • 15. Sensitivity Testing

    • 16. Citation, Coverage, and Redundancy

    • 17. Uncertainty, Disagreement, and Revision

    • 18. Protocol Dispositions

  • Part V — Audit, Validation, and Limits

    • 19. The Epistemic Classification Record

    • 20. Human Review, Override, and Replay

    • 21. The Adversarial Benchmark

    • 22. Metrics and Validation

    • 23. Falsification and Failure Conditions

    • 24. System Boundary and Non-Claims

  • Conclusion — Classification Must Remain Answerable to Correction

  • References

  • Appendix A — Evidence and Claim-Control Register

  • Appendix B — Protocol Reference Contract

  • Appendix C — Adversarial Test Catalogue

  • Appendix D — Candidate Metrics and Scoring Notes

Introduction

AI systems increasingly mediate access to knowledge by determining which sources enter the effective evidence environment and what comparative standing they receive. This allocation occurs through retrieval, ranking, citation, recommendation, synthesis, review, and evaluation; it cannot be assessed solely through the factual accuracy of the final answer.

A correct conclusion may arise from a prestige-sensitive process. A stable output may preserve an incorrect object representation. A real citation may fail to support the local claim, and several apparent sources may reproduce one evidential dependency. Conversely, a weak unconventional claim may be rejected correctly even when the result resembles a conformity pattern.

The protocol addresses these cases through five linked distinctions. It resolves the evaluated object before transferring judgment; separates substantive support from process robustness; distinguishes legitimate provenance from evidential substitution; applies symmetric standards to familiar and unfamiliar material; and requires a specific correction path for consequential classifications.

Part I defines the allocation object. Part II establishes the bounded empirical motivation. Part III specifies the symmetric failure space. Part IV converts the framework into a classification protocol. Part V governs the record, human intervention, validation, falsification, and system boundaries.

The dossier’s central claim is procedural:

A consequential AI-mediated classification should remain connected to a resolved object, a traceable evidence structure, a separately reported process judgment, and a specific path through which relevant evidence can alter its future use.

Part I — The Epistemic Allocation Object

1. Retrieval Is an Allocation Function

1.1 From answer production to epistemic allocation

Retrieval is often described as a preparatory step preceding generation. For governance purposes, this is incomplete. Retrieval participates directly in allocating epistemic visibility.

A source may move through the following states:

S_(available)→S_(retrieved)→S_(used)→S_(cited)

The transitions are neither automatic nor equivalent.

A source may be available but never retrieved. It may be retrieved but excluded from context. It may appear in context but contribute no material evidence. It may influence synthesis without visible citation. It may be cited contextually without supporting the decisive proposition.

The relevant allocation object therefore includes more than a ranked list. It includes:

  • visibility;

  • inclusion;

  • comparative attention;

  • source use;

  • citation opportunity;

  • and downstream reusability.

1.2 Unequal allocation is not inherently defective

A useful system must allocate attention unequally. It should prioritize evidence that is:

  • relevant;

  • direct;

  • current where currentness matters;

  • authoritative where formal authority is the object;

  • and methodologically adequate.

The governance question is not whether all sources receive equal standing. It is whether the allocation is answerable to the task and evidence.

1.3 The functional absence threshold

A source may remain technically present while becoming functionally absent. This occurs where its rank, position, or presentation places it below the threshold at which it can affect:

  • context inclusion;

  • user attention;

  • citation;

  • or final judgment.

A small rank change may therefore be materially consequential where it crosses an inclusion cutoff.

1.4 Epistemic allocation

This dossier defines AI-mediated epistemic allocation as:

The distribution of visibility, relevance, comparative attention, evidential standing, and citation opportunity through AI-mediated retrieval, ranking, synthesis, or evaluation.

This is an original conceptual definition. It is not presented as an already standardized technical field or validated metric.

1.5 Governable allocation

An allocation becomes governable where the record can answer:

  1. What task determined relevance?

  2. What source environment was available?

  3. Which sources were retrieved, used, and cited?

  4. Which evidence was decisive?

  5. Which provenance and presentation variables were visible?

  6. Which variables were tested?

  7. What could revise the result?

2. The Governed Classification Event

2.1 The event as the unit of governance

A global claim such as “the system favors institutions” is often too broad for operational diagnosis. The protocol instead governs a bounded classification event.

An epistemic classification event is:

A versioned event in which an AI-mediated process assigns standing through retrieval, inclusion, omission, ranking, citation, recommendation, qualification, or evaluation.

The event has a task, object, source environment, execution, and intended use.

2.2 Canonical classification object

The event is represented as:

𝒜ₜ=⟨Y,Q,O,C,S,Π,E,B,U,R,V⟩

This tuple does not claim to reproduce every hidden variable. It captures the observable governance boundary.

2.3 System and execution

Y records the system and execution, including:

  • provider and product surface;

  • model family and version;

  • retrieval mode;

  • search provider;

  • evaluator system;

  • enabled tools;

  • execution time;

  • and known configuration limits.

The same task can produce different results across systems, runs, retrieval modes, and time. System identity is therefore part of the object.

2.4 Query and intended use

Q contains:

  • the raw query;

  • normalized task;

  • intended use;

  • requested output;

  • temporal scope;

  • source constraints;

  • and decision consequence.

A task asking for official authority differs from one asking for strongest evidence. Citation popularity may be relevant to historical influence and irrelevant to methodological quality.

2.5 Classification consequence

Protocol depth should scale with consequence. Low-consequence exploratory discovery may require a limited record. High-consequence decisions affecting access, standing, safety, rights, or irreversible action require stronger testing, review, and replayability.

2.6 The final output

The protocol produces:

𝒥ₜ=⟨Jₛ,Jₚ,G,F,U,R,Λ⟩

The output is structured because no single credibility score can preserve all relevant distinctions.

3. Object Resolution

3.1 Why object errors are load-bearing

Evidence can be accurate and still support the wrong object. A failed deployment may be used to reject a framework. One weak claim may be used to discredit an author. A detailed architecture may be treated as an implemented product.

These are object-transfer failures.

3.2 Canonical object types

The protocol recognizes object types including:

  • claim;

  • claim set;

  • document;

  • document section;

  • author;

  • institution;

  • project;

  • framework;

  • formal model;

  • architecture;

  • prototype;

  • evaluation result;

  • product;

  • deployment;

  • policy;

  • source set;

  • citation;

  • classification record.

3.3 Transfer bridge

A judgment transfer from object Oᵢ to object Oⱼ requires:

Oᵢ→_(B,W,D)Oⱼ

where:

  • B is the object relation;

  • W is the evidential warrant;

  • D is the defeat or limitation condition.

Without this bridge, transfer is blocked.

3.4 Valid and invalid transfers

A load-bearing factual error may justify downgrading a document whose conclusion depends entirely upon it. This is a valid transfer.

One failed claim does not automatically establish that every claim by the author or institution is unreliable. This is an invalid transfer.

A deployment failure may trace validly to a load-bearing architectural defect. It does not always do so.

3.5 Object-resolution statuses

The protocol records the object as:

  • resolved;

  • resolved with limitation;

  • decomposed;

  • ambiguous;

  • unresolved;

  • or invalid.

Where the object cannot be resolved responsibly, S5, HOLD, REASSESS, or INVALID may be appropriate depending on the process state.

3.6 Failure flag

Unwarranted transfer triggers:

F1=Object Conflation

4. Claim Type and Maturity

4.1 Claim decomposition

A document may contain multiple claim types:

  • empirical;

  • historical;

  • causal;

  • conceptual;

  • definitional;

  • methodological;

  • architectural;

  • normative;

  • predictive;

  • and maturity claims.

The protocol decomposes material claims before global evaluation.

4.2 Load-bearing claims

A claim is load-bearing where its rejection or weakening would alter:

  • the main conclusion;

  • the object’s standing;

  • the declared maturity;

  • or authorized use.

Load-bearing claims receive priority in retrieval, citation verification, contradiction review, and revision analysis.

4.3 The maturity ladder

The protocol uses the following maturity architecture:

Concept→Formalization→Architecture→Prototype→Controlled Evaluation→Validation→Production→Certification

The stages are not assumed to have equal empirical distance.

4.4 Upward promotion

Upward promotion occurs where:

M_(claimed)>M_(supported)

Examples include:

  • formalization treated as empirical validation;

  • architecture treated as implementation;

  • prototype treated as production;

  • internal evaluation treated as external validation.

4.5 Downward collapse

Downward collapse occurs where a valid lower-stage contribution is rejected because it lacks evidence required only by a higher stage it does not claim.

A concept can be coherent without a prototype. An architecture can be complete without production evidence. The lower-stage claim must still satisfy the standards of that stage.

4.6 Mixed maturity

A project may contain:

  • a developed concept;

  • partial formalization;

  • a complete architecture;

  • no prototype;

  • and no validation.

The correct result is a maturity vector, not one global label.

4.7 Failure flag

Upward promotion and downward collapse are governed by:

F2=Claim-Maturity Collapse

The direction must be recorded.

Part I Synthesis

Consequential classification begins with the task, object, claims, maturity, source environment, and execution—not with the final answer. Part II examines the non-content variables that can affect this structure.

Part II — Evidence, Provenance, and Non-Content Sensitivity

5. What the Evidence Already Shows

5.1 Evidential boundary

The evidence base does not establish one universal AI conformity mechanism. It establishes a narrower and operationally sufficient proposition:

□ Non-content sensitivity is observable and testable

The effects differ by task, system, domain, intervention, and outcome measure. The protocol therefore treats each study as evidence for a bounded test target rather than as proof of general model behavior.

5.2 Citation-popularity sensitivity

Algaba et al. analyzed scholarly-reference generation using papers from AAAI, NeurIPS, ICML, and ICLR. They found that the tested LLMs reproduced broad human citation patterns while displaying a more pronounced preference for highly cited papers. The reported preference persisted after controls for publication year, title length, number of authors, and venue (Algaba et al. 2025).

This supports citation-popularity counterfactuals in comparable tasks where the requested criterion is evidential strength rather than historical influence.

It does not establish that low-citation work is generally suppressed, that highly cited work is generally weak, or that citation popularity is always illegitimate.

5.3 Authorship metadata and attribution

Abolghasemi et al. used counterfactual evaluation in generator-aware RAG pipelines and found that adding authorship information could change attribution quality materially in the tested systems. The study also reported sensitivity to explicit human versus AI authorship labels (Abolghasemi et al. 2025).

The authorized inference is:

□ Authorship metadata is a testable input to attribution behavior

The study does not establish a universal preference for named, human, or institutionally affiliated authors.

5.4 Institutional affiliation in LLM-assisted peer review

Vasu et al. investigated LLM-generated peer reviews under controlled interventions involving affiliation, gender, seniority, and publication history. They reported strong affiliation effects favoring highly ranked institutions and found that seniority and publication-history preferences could affect acceptance outcomes in borderline cases (Vasu et al. 2026).

This supports blinded-versus-visible affiliation testing in comparable evaluative tasks. It does not establish that AI systems generally favor elite institutions across retrieval, web search, citation, technical assessment, or other domains.

5.5 Source identity in political citation selection

Dai et al. examined citation selection in a political-news setting and found that the tested LLMs cited left-leaning media outlets at higher rates than traditional retrieval baselines. Their controlled experiments attributed the observed difference primarily to media-outlet identity rather than to left-oriented article content alone (Dai et al. 2025).

The finding supports source-identity counterfactuals in comparable political citation tasks. It must remain bounded to the tested systems, the political-news domain, the AllSides-2024 construction, and the interventions used.

5.6 Option order and token sensitivity

Wei et al. systematically evaluated option-order and option-token effects across multiple models and tasks and found that both variables could affect LLM selection behavior (Wei et al. 2024).

This supports permutation testing where options are semantically equivalent, their order should not determine the result, and the resulting decision is consequential.

5.7 Position bias in LLM-as-a-judge

Wang et al. found that changing the order of candidate responses could substantially alter comparative rankings generated by an LLM evaluator (Wang et al. 2024).

Shi et al. extended the analysis across pairwise and list-wise evaluation, fifteen LLM judges, twenty-two tasks, approximately forty solution-generating models, and more than 150,000 evaluation instances. They reported systematic position effects varying by judge, candidate-quality gap, and task (Shi et al. 2025).

The evaluator must therefore be treated as part of:

Y=system and execution

rather than as a neutral external authority.

5.8 Long-context position

Liu et al. found that performance on tasks requiring retrieval of relevant information from long contexts often declined when decisive information appeared in the middle rather than near the beginning or end of the input (Liu et al. 2024).

Hsieh et al. connected this phenomenon in their experiments to a U-shaped positional-attention pattern and evaluated a calibration mechanism designed to improve use of relevant middle-position information (Hsieh et al. 2024).

These studies support:

□ Source Presence⇏Effective Evidential Use

They do not establish that middle-position evidence is always ignored.

5.9 Linguistic complexity in retrieval

Cheng and Amiri found performance disparities associated with the linguistic complexity of input queries and evaluated EqualizeIR as a mitigation framework across retrieval tasks (Cheng and Amiri 2025).

The finding supports query-reformulation tests, linguistic-complexity controls, and retrieval comparison. It does not establish that retrieval systems suppress unfamiliar ontologies.

5.10 Generative-search heterogeneity

Kirsten et al. compared Google organic search with five generative-search systems from Google, OpenAI, and Perplexity. They reported substantial variation in reliance on internal versus external knowledge, source diversity, retrieval footprint, synthesis strategy, execution, and temporal stability (Kirsten et al. 2026).

The authorized conclusion is:

Generative-search classifications should preserve system identity, retrieval mode, execution, source boundary, and date.

The result should not be generalized to “AI” as a single system class.

5.11 Citation evaluation is multidimensional

Xu et al. introduced CiteEval, a citation-evaluation framework that evaluates citation quality in relation to the user query, generated text, cited source, and broader retrieval context. The work also introduced CiteBench and automated citation metrics aligned with the framework (Xu et al. 2025).

This supports:

□ Citation Presence⇏Citation Support

CiteEval provides methodological support for the citation component. It does not validate the complete protocol.

5.12 Retrieval diversity and redundancy

Khan et al. argued that similarity-focused RAG can introduce redundant content in reasoning-intensive question answering and reported that query-aware, relevance-constrained diversity improved F1 performance over vanilla cosine-similarity RAG in the tested benchmarks (Khan et al. 2026).

The source supports separating source count from independent evidential contribution. It does not establish ideological-balance requirements, forced viewpoint diversity, or universal benefit from adding dissimilar material.

5.13 Automatic review and faulty reasoning

Dycke and Gurevych introduced a counterfactual framework that inserted controlled research-logic faults into papers and found that those faults had no significant effect on the output reviews of the evaluated automatic-review approaches (Dycke and Gurevych 2026).

This supports controlled-defect insertion, strong-versus-weak item pairs, and adversarial testing of load-bearing reasoning detection. It does not establish a universal preference for rhetorically polished but defective material.

5.14 Human and machine judgment sensitivity

Chen et al. evaluated human and LLM judges under controlled misinformation-oversight, gender, authority, and beauty perturbations. They found that both human and machine judges were vulnerable to the tested perturbations, although the degree and pattern of sensitivity differed (Chen et al. 2024).

The study supports:

□ Human Review Is a Governed Intervention

It does not establish equivalence between human and machine failure modes.

5.15 Bounded conclusions

The evidence supports five bounded conclusions.

First, variables other than substantive content can affect AI-mediated judgments in tested settings.

Second, provenance and presentation variables can be examined through controlled intervention.

Third, citation presence is not equivalent to effective source support.

Fourth, evaluator systems require their own governance.

Fifth, the system, execution, retrieval mode, and date must be preserved where outputs may vary.

The evidence does not establish a complete causal theory of AI epistemic allocation.

6. Provenance as Evidence and as Substitute

6.1 Provenance has legitimate evidential functions

Provenance may establish:

  • identity;

  • authenticity;

  • official authority;

  • document version;

  • primary-source status;

  • accountability;

  • conflict of interest;

  • or reliability-relevant process.

A protocol that removes provenance universally would destroy legitimate evidence.

6.2 Provenance use

Provenance use occurs where provenance performs a task-relevant and explicitly represented function.

Examples include:

  • an official regulator as the controlling source for current regulation;

  • an original dataset repository as evidence of version and authorship;

  • a retraction notice from the publishing journal;

  • an institution’s own website as evidence of what that institution claims.

6.3 Provenance substitution

Provenance substitution occurs where institutional status, author reputation, venue, source category, or popularity replaces the claim-level assessment required by the task.

The protocol does not infer substitution merely because provenance changes the result. It asks whether provenance was constitutive, materially relevant, potentially relevant, irrelevant, or unresolved.

6.4 Relevance categories

The canonical relevance categories are:

  • constitutive;

  • materially relevant;

  • potentially relevant;

  • irrelevant;

  • unresolved.

A provenance variable is constitutive where changing it changes the object itself. Official authorship of a regulation is constitutive of an authority task.

6.5 Provenance counterfactual

A provenance counterfactual alters author identity, affiliation, venue, institutional label, or source category while preserving the material claim and evidence.

Counterfactual authorship, affiliation, and source-identity interventions have already been used to expose observable changes in attribution, peer-review judgment, and citation selection in bounded experimental settings (Abolghasemi et al. 2025; Vasu et al. 2026; Dai et al. 2025).

The protocol compares the baseline judgment with the intervention judgment and records changes in:

  • substantive class;

  • confidence;

  • rank;

  • source selection;

  • citation;

  • rationale;

  • maturity;

  • and disposition.

6.6 Counterfactual validity

A valid provenance counterfactual preserves:

  • object identity;

  • facts;

  • evidence quantity;

  • claim strength;

  • scope;

  • maturity;

  • and clarity within tolerance.

If the intervention also changes authenticity, authority, or evidence access, the comparison may be contaminated.

6.7 Prestige substitution

The protocol defines:

F3=Prestige Substitution

where institutional, authorial, or venue prestige substitutes materially for the evidential assessment required by the task.

Controlled peer-review experiments provide direct task-specific evidence that affiliation and related author metadata can affect LLM-generated evaluation outcomes (Vasu et al. 2026).

A finding of prestige sensitivity does not establish that the lower-prestige source is correct. Substantive assessment remains necessary.

6.8 Provenance concealment

The protocol defines:

F8=Provenance Concealment

where observable provenance dependence materially affects the classification but the result is represented as purely content-derived.

The relevant governance failure is not always the influence itself. It may be the failure to disclose the influence and its task relevance.

6.9 Blinding

Blinding is conditional, not universal. It is appropriate where identity is not constitutive and where the content can be assessed without destroying legitimacy-relevant information.

Blinding may be inappropriate for:

  • authority;

  • authenticity;

  • conflict-of-interest;

  • and version-control tasks.

6.10 Human provenance sensitivity

Human review is subject to the same distinction. Human evaluators are not exempt from perturbation sensitivity and should be assessed through visible-versus-blinded logic where the object permits it (Chen et al. 2024).

6.11 Chapter result

The protocol seeks neither maximum provenance influence nor zero provenance influence. It seeks correct, explicit, task-relative provenance use.

□ Provenance use≠Provenance substitution

7. Citation Networks and Visibility Reinforcement

7.1 Citation as visibility infrastructure

Citation provides more than attribution. It creates a gateway through which users, systems, and later documents encounter sources.

A citation can affect:

  • visibility;

  • perceived legitimacy;

  • future retrieval probability;

  • and reuse.

7.2 Citation popularity as a prior

Citation count may legitimately indicate historical influence or field visibility. It may also function as a prior when rapid source triage is necessary.

Experimental evidence shows that citation popularity can influence scholarly-reference recommendation in tested LLM settings (Algaba et al. 2025).

The failure occurs where popularity replaces evidence appropriate to the task.

7.3 Citation-popularity substitution

The canonical failure is:

F5=Citation-Popularity Substitution

It applies where citation count, citation visibility, or perceived scholarly popularity substitutes materially for evidence appropriate to the task.

7.4 Recommendation-level evidence

Current recommendation-level evidence supports the bounded proposition that citation popularity can influence scholarly-reference generation in tested LLMs (Algaba et al. 2025).

It does not establish that the recommended source is weak or that the omitted source is strong.

7.5 Candidate reinforcement dynamic

A possible longitudinal dynamic is:

Existing Visibility→Machine Recommendation→Human Reuse→Future Visibility

Algaba et al. provide recommendation-level and citation-graph evidence relevant to the first stages of this possible dynamic, but they do not establish the complete sequence from machine recommendation to later human reuse and future ranking (Algaba et al. 2025).

The complete citation-network reinforcement loop remains on HOLD.

7.6 Source count and corroboration

A large source set can consist of duplicates, syndicated copies, shared datasets, or repeated references to one primary report.

□ Source Count⇏Independent Corroboration

7.7 Citation presence and support

A source may be cited without supporting the local proposition. It may provide contextual background, partial support, or contradiction.

□ Citation Presence⇏Claim Support

7.8 Popularity masking

A citation-popularity test may mask citation counts or popularity labels while preserving evidence. The task must determine whether popularity is relevant.

7.9 Recency correction

Recent high-quality work may have low citation counts because insufficient time has passed. Recency does not establish quality, but it may defeat an interpretation that treats low citation count as negative evidence.

7.10 Symmetric boundary

High citation count does not immunize weak evidence. Low citation count does not indicate hidden merit. The protocol evaluates both through the same claim-level standard.

8. Presentation, Order, and Semantic Form

8.1 Presentation is part of the observable process

The protocol records presentation variables because content can remain constant while the output changes under:

  • source order;

  • candidate order;

  • context position;

  • formatting;

  • terminology;

  • novelty framing;

  • or rhetorical polish.

8.2 Order sensitivity

Option-order, candidate-position, and list-wise judge studies provide direct evidence that ordering can alter LLM selection and evaluation outcomes in tested settings (Wei et al. 2024; Wang et al. 2024; Shi et al. 2025).

The protocol defines:

F11=Presentation-Order Sensitivity

where valid changes to source order, candidate order, or context position produce material classification change.

8.3 Material and non-material change

An order change among equivalent secondary sources may be informational. A change from Supported to Contradicted, a decisive-source change, or an inclusion-cutoff crossing is material.

8.4 Position and evidence use

Long-context studies show that relevant evidence may remain technically present while receiving reduced effective use in disadvantaged context positions (Liu et al. 2024; Hsieh et al. 2024).

This requires the record to distinguish:

  • context presence;

  • acknowledgment;

  • integration;

  • citation;

  • and effect on the final judgment.

8.5 Anchoring

Earlier sources or conclusions may establish the frame through which later evidence is interpreted. A valid order permutation can reveal whether the system integrates contrary evidence symmetrically.

8.6 Rhetorical coherence

A source may contain a named problem, formal notation, taxonomy, architecture, and validation plan. This can improve assessability. It can also create an impression of evidential maturity that exceeds the record.

The protocol defines:

F7=Rhetorical-Coherence Substitution

where polish, formalization, conceptual density, narrative closure, or architectural detail substitutes for substantive evidence or maturity.

Counterfactual automatic-review research motivates testing whether a controlled reasoning defect remains undetected when the surrounding paper retains plausible scholarly form (Dycke and Gurevych 2026). It does not establish a general rhetorical bias.

8.7 Formalization is not validation

□ Formalization⇏Empirical Support

Formal notation may constrain a theory or stabilize an architecture. It does not demonstrate correspondence with the world.

8.8 Semantic form

The protocol defines:

F4=Semantic-Form or Familiarity Substitution

where terminological familiarity or conventional formulation receives standing beyond legitimate clarity, precision, or domain relevance.

Retrieval research shows that linguistically different query formulations can produce measurable performance disparities (Cheng and Amiri 2025).

8.9 Linguistic complexity and conceptual novelty

Linguistic complexity is not equivalent to conceptual novelty. An unfamiliar term may be unclear, redundant, or non-redundant. A valid semantic reformulation must preserve object, scope, strength, evidence, and practical implication.

8.10 Ontology-preserving conformity boundary

The empirical evidence motivates component tests but does not establish a unified ontology-preserving mechanism. That concept is introduced in Chapter 10 as an original diagnostic synthesis on HOLD.

8.11 Novelty framing

The reverse test is also required. The same content may be framed as established, new, independent, overlooked, or paradigm-changing. Novelty should neither receive an unsupported bonus nor an unsupported penalty.

8.12 Stable error and unstable truth

A system may be consistently wrong, or inconsistently correct. The distinction is empirically motivated by findings showing both positional instability in evaluators and failure to detect controlled reasoning defects (Wang et al. 2024; Shi et al. 2025; Dycke and Gurevych 2026).

The protocol must therefore separate substantive correctness from process robustness.

8.13 Causal humility

A controlled intervention can establish observable sensitivity. It does not necessarily reveal the complete internal causal pathway.

□ Observable Sensitivity≠Complete Internal Causal Explanation

Part II Synthesis

The evidence supports bounded testing of provenance, popularity, identity, order, position, semantic form, novelty framing, and execution variance. It does not support a universal narrative of privilege or suppression. Part III therefore interprets these effects through a symmetric failure model.

Part III — Symmetric Classification Failure

9. Why Ranking Direction Does Not Identify Failure

9.1 The directional fallacy

A low rank does not establish suppression, conformity, or evaluator failure. A high rank does not establish independence, originality, or correctness.

□ Ranking Direction⇏Classification Quality

The same direction can be legitimate or defective depending on the object, evidence, maturity, provenance relevance, and process robustness.

9.2 Four-cell symmetry

The minimum matrix is:

Evidential quality
Familiar or prestigious
Unfamiliar or low-prestige
Strong
Recognize
Recognize
Weak
Reject or qualify
Reject or qualify

The governing comparisons are:

A↔B

and:

C↔D

Strong evidence should not lose standing merely because it is unfamiliar. Weak evidence should not gain standing merely because it is prestigious. The reverse protections are equally necessary.

9.3 Recognition is not promotion

Recognition means classifying the contribution accurately at the supported scope and maturity.

Promotion means assigning broader scope, stronger support, or higher maturity than the evidence justifies.

A concept may deserve recognition as non-redundant and testable without being empirically validated. An architecture may deserve recognition as coherent without being production-ready.

9.4 Rejection is not erasure

Rejection concerns a claim, inference, or maturity assertion. It should not erase unaffected components.

A document may contain one contradicted causal claim, one useful conceptual distinction, and one incomplete architecture. A responsible classification preserves this structure.

9.5 Evidence-based rejection is a success

A robust rejection may be represented as:

⟨S4,P1,RELEASE⟩

RELEASE means that the classification is authorized for the declared use. It does not indicate favorable standing.

9.6 Evidence-based recognition is a success

A strong unfamiliar contribution may receive:

⟨S1,P1,RELEASE⟩

A supported result produced through a materially sensitive process may receive:

⟨S1,P3,REASSESS⟩

or RELEASE WITH LIMITATION, depending on consequence and independent verification.

9.7 Correct outcome and robust process

The protocol preserves two non-equivalences:

□ Correct Outcome⇏Robust Process

□ Robust Process⇏Correct Outcome

A prestigious source may be supported but recognized only when its affiliation is visible. The substantive result may be correct while the process is defective.

A system may also reproduce an incorrect maturity classification consistently. Stability does not make the classification correct.

9.8 Directional compensation

Replacing prestige preference with anti-prestige preference does not create evidence-based classification.

Prestige Preference→Anti-Prestige Preference⇏Evidence-Based Classification

The protocol cannot compensate by granting automatic bonuses to:

  • marginality;

  • low citation count;

  • contrarian framing;

  • or independent origin.

9.9 Self-protective conformity theories

A framework may predict that unfamiliar frameworks will be rejected and then treat its own rejection as evidence of correctness.

□ Rejection of a Conformity Framework⇏Evidence of Conformity

The rejection may be justified because the concept is redundant, untestable, poorly supported, or promoted beyond maturity.

A valid conformity diagnosis requires evidence external to the fact of rejection.

9.10 Equal standards do not require equal rank

Equivalent standards do not require identical outcomes. Sources may differ legitimately in directness, currency, authority, scope, and methodological quality.

The symmetry principle is:

□ Equivalent Evidence→Equivalent Substantive Standard

It is not a requirement of equal visibility or rank.

10. Ontology-Preserving Conformity

10.1 Concept status

Ontology-preserving conformity is an original candidate diagnostic synthesis. It is not presented as a general law of AI behavior or a validated unified mechanism.

Its status is:

□ Conceptually Defined and Operationally Decomposable

□ Unified Empirical Mechanism: HOLD

10.2 Working definition

Ontology-preserving conformity occurs where:

  1. an inherited representation defines the evaluated object;

  2. material evidence introduces an anomaly or alternative object structure;

  3. the system preserves the inherited representation;

  4. and the process lacks an operative path through which the anomaly can revise the representation.

The minimum structure is:

□ Inherited Representation+Material Anomaly+Preservation+Correction Failure

All four elements are required.

10.3 Operational meaning of ontology

Ontology is used in a bounded operational sense. It refers to the categories and relations through which the object is represented.

Examples include whether the system distinguishes:

  • claim from author;

  • architecture from deployment;

  • source authority from claim support;

  • substantive correctness from process robustness;

  • and concept-stage value from validation.

The term does not require a comprehensive metaphysical worldview.

10.4 Inherited representation

An inherited representation may derive from conventional terminology, disciplinary categories, source schemas, institutional rubrics, product taxonomies, or recurring answer patterns.

Inheritance is not itself defective. Stable categories may be accurate and efficient.

10.5 Material anomaly

A material anomaly is evidence that cannot be handled adequately without revising the object, decomposing the claim, changing maturity, or recognizing a distinct relation.

A novel term, low rank, or unfamiliar source alone is not a material anomaly.

10.6 Preservation

Preservation may appear as:

  • continued global scoring after valid decomposition;

  • repeated architecture-to-product conflation;

  • unchanged maturity after contrary evidence;

  • persistent exclusion after semantic normalization;

  • or unchanged classification after decisive evidence is reformulated clearly.

It must be demonstrated comparatively, not inferred from one unfavorable judgment.

10.7 Correction failure

Correction failure is the load-bearing component. A system that revises its representation successfully does not exhibit ontology-preserving conformity.

Possible observable forms include:

  • no revision condition;

  • generic acknowledgment without structural change;

  • repeated reassessment reproducing the unsupported object;

  • or inability to identify what evidence could reopen the classification.

10.8 Component flags

The following flags may contribute:

  • F3 — Prestige Substitution;

  • F4 — Semantic-Form or Familiarity Substitution;

  • F5 — Citation-Popularity Substitution;

  • F8 — Provenance Concealment;

  • F10 — Correctionless Classification;

  • F17 — Source-Set Incompleteness.

No single flag proves the unified mechanism.

10.9 Test architecture

Candidate tests include:

  • equivalent semantic restatement;

  • object-resolution intervention;

  • evidence insertion;

  • blinded provenance;

  • revision prompt;

  • repeated runs;

  • cross-system comparison;

  • and replay after correction.

A positive diagnosis requires an anomaly strong enough to warrant revision and failure to revise under an adequate challenge.

10.10 Legitimate stability

The inherited representation may remain correct. An alternative category may be redundant, less precise, or unsupported.

A benchmark must therefore contain legitimate-stability controls. Otherwise, reclassification itself becomes the target behavior.

10.11 Partial revision

A system may recognize a conceptual contribution while rejecting a causal claim and preserving the existing maturity classification. This may be correct differentiated assessment, not conformity.

10.12 Falsification conditions

The unified concept should be narrowed or removed if:

  • its components do not co-occur reliably;

  • object persistence is explained by evidence quality;

  • semantic reformulation does not improve classification;

  • provenance effects do not predict correction failure;

  • the concept adds no value beyond existing flags;

  • evaluators cannot apply it reliably;

  • or it systematically overrecognizes unfamiliar frameworks.

11. Novelty and Rhetorical Amplification

11.1 Symmetric reverse risk

A protocol designed to expose conformity can begin treating novelty, marginality, independence, or claims of suppression as positive evidence.

This is not correction. It is a reverse-direction failure.

11.2 Novelty as task criterion

Novelty may be legitimate where the task asks for emerging approaches or horizon scanning. It may affect discovery priority.

It does not establish correctness, maturity, or future importance.

□ Novelty as Search Criterion≠Novelty as Evidential Support

11.3 Novelty amplification

The protocol defines:

F6=Novelty Amplification

as the assignment of unsupported epistemic standing to novelty, unconventionality, marginality, contrarian framing, or claims of underrecognition.

A supported F6 finding requires:

  1. controlled content-equivalent variants;

  2. altered novelty framing;

  3. a materially changed judgment;

  4. novelty not constitutive of the task;

  5. and no evidential change explaining the result.

11.4 Marginality and suppression narratives

Marginality may reflect weak evidence, limited review, poor communication, or lack of relevance. It may also reflect genuine underrecognition. The status must be established, not assumed.

A claim of suppression is a separate claim requiring evidence about actor, mechanism, comparison, and consequence.

□ Low Visibility⇏Suppressed Merit

11.5 Rhetorical-coherence substitution

The protocol defines:

F7=Rhetorical-Coherence Substitution

where polish, formalization, conceptual density, narrative closure, or architectural detail substitutes for the evidence required by the claim.

11.6 Formalization bias

Formal notation may increase perceived rigor. The protocol asks whether:

  • the variables are defined;

  • the relation constrains the claim;

  • the formalization is non-trivial;

  • and a testable correspondence exists.

□ Formal Coherence⇏Empirical Support

11.7 Architecture bias

A detailed architecture may contain components, interfaces, control flows, records, and validation plans. It may still lack implementation, performance evidence, calibration, security testing, or operational validation.

□ Architecture⇏Deployment

11.8 Density and clarity

Conceptual density may produce underrecognition because the source is difficult. It may also produce overrecognition because the source appears sophisticated.

The protocol does not punish clarity. Presentation can legitimately improve assessability. A plain variant that is genuinely ambiguous is not content-equivalent to a clear variant.

11.9 Anti-institutional overcorrection

Institutions can provide authenticated records, formal accountability, stable versioning, methodological review, and correction mechanisms. The protocol does not discount these properties by default.

11.10 Four-cell novelty test

Evidence
Established framing
Novel framing
Strong
Recognize
Recognize
Weak
Reject or qualify
Reject or qualify

The desired result is evidence-sensitive invariance under irrelevant novelty framing.

12. Legitimate Rejection and Legitimate Recognition

12.1 Four legitimate outcomes

A responsible process supports:

  • evidence-based recognition;

  • evidence-based rejection;

  • bounded uncertainty;

  • and differentiated assessment.

These outcomes prevent classification from collapsing into acceptance versus suppression.

12.2 Legitimate recognition

Recognition requires:

  • resolved object;

  • bounded claims;

  • appropriate evidence;

  • correct maturity;

  • adequate citation;

  • and no unresolved load-bearing contradiction.

Recognition can remain limited:

  • the historical claim is supported;

  • the conceptual distinction is non-redundant;

  • the architecture is specified coherently;

  • the prototype demonstrates executability;

  • the validation claim remains unsupported.

12.3 Legitimate rejection

Legitimate rejection may follow from:

  • contradiction;

  • invalid generalization;

  • causal overclaim;

  • citation mismatch;

  • maturity promotion;

  • redundancy;

  • undefined object;

  • or controlled-test failure.

A rejection should identify whether a weaker formulation remains viable.

12.4 Bounded uncertainty

The protocol distinguishes:

S3=Insufficient Evidence

from:

S5=Not Assessable

S3 means the claim and object are clear but evidence is inadequate. S5 means the classification object or evidential basis cannot be constructed responsibly.

12.5 Missing-evidence neutralization

The protocol defines:

F9=Missing-Evidence Neutralization

where missing evidence is treated as harmless, supportive, or maturity-neutral despite being required by the claim.

Absence of implementation evidence may be appropriate at the architecture stage. It is not evidence of implementation performance.

12.6 Differentiated assessment

A source may receive a vector such as:

Component
Judgment
Problem definition
Supported
Conceptual distinction
Non-redundant
Causal explanation
Insufficient evidence
Architecture
Complete within declared boundary
Validation claim
Contradicted

This is structured judgment, not indecision.

12.7 Legitimate HOLD

HOLD means the evaluated claim or decision must not be treated as settled until a specified condition is satisfied.

A valid HOLD states:

  • what is unresolved;

  • why it matters;

  • what evidence is required;

  • and what judgment could change.

HOLD is not rejection.

12.8 Excessive HOLD

A system may avoid error by refusing to decide. This can preserve existing allocation patterns and destroy operational value.

The benchmark therefore measures Excessive HOLD Rate.

12.9 REASSESS and INVALID

REASSESS concerns a process that must be repeated under corrected conditions.

INVALID concerns a run or record that cannot support responsible use.

Neither disposition proves that the evaluated claim is false.

12.10 Symmetric result

The classification quality function is:

□ Classification Quality=f(Object,Evidence,Maturity,Robustness,Correction)

It is not a function of ranking direction alone.

Part III Synthesis

The protocol must detect both hierarchy-preserving and hierarchy-reversing error: familiar and unfamiliar strength require recognition, while prestigious and marginal weakness require equivalent qualification or rejection. Part IV converts this symmetry into executable judgment and disposition rules.

Part IV — The Classification Protocol

13. Dual Judgment

13.1 Why one judgment is insufficient

A conventional evaluation may ask whether a source is credible. This compresses distinct questions:

  • Is the claim supported?

  • Is the object correctly identified?

  • Is maturity represented accurately?

  • Did provenance affect the result?

  • Did order or execution affect it?

  • Are the citations valid?

  • Can the result be reconstructed and revised?

One scalar answer cannot preserve these distinctions.

13.2 Substantive judgment

The substantive judgment asks:

What does the available evidence support concerning the declared object, claim, scope, strength, and maturity?

Jₛ∈{S1,S2,S3,S4,S5}

S1 — Supported

The evidence supports the claim at the declared object, scope, strength, and maturity.

S2 — Partially Supported

The evidence supports a weaker, narrower, or lower-maturity formulation. S2 requires an explicit supported reformulation.

S3 — Insufficient Evidence

The object and claim are assessable, but the evidence does not justify support or contradiction.

S4 — Contradicted

Material evidence conflicts with the claim or defeats a load-bearing inference.

S5 — Not Assessable

The object, claim, source, or governing standard cannot be resolved sufficiently for responsible classification.

13.3 Process-robustness judgment

The process judgment asks:

How stable, reconstructable, and sensitivity-tested is the process that produced the substantive result?

Jₚ∈{P1,P2,P3,P4,P5}

P1 — Robust

Required valid tests were completed and no tested variable produced a material change in class, decisive evidence, source selection, citation, rank, rationale, or disposition.

P2 — Conditionally Robust

A bounded limitation exists without altering the decisive judgment or authorized use.

P3 — Materially Sensitive

At least one valid and relevant intervention produces an L3 or L4 material change.

P4 — Unstable

Nominally equivalent executions produce materially inconsistent results without an adequate explanation.

P5 — Process Invalid

A load-bearing part of the classification process is structurally defective.

13.4 Orthogonality

The two judgments are analytically separate.

Possible combinations include:

⟨S1,P1⟩

supported and robust;

⟨S1,P3⟩

supported but materially process-sensitive;

⟨S4,P1⟩

contradicted through a robust process;

⟨S3,P1⟩

robustly insufficient evidence.

13.5 Confidence

Confidence belongs inside the uncertainty record. It does not replace either judgment class.

A system can be highly confident and wrong, uncertain and correct, or stable under an incomplete source environment.

13.6 Final judgment vector

□ 𝒥ₜ=⟨Jₛ,Jₚ,G,F,U,R,Λ⟩

The governing separation is:

□ What the evidence supports≠how robustly the process recognized it

14. From Object Resolution to Evidence Graph

14.1 Task registration

The record identifies:

  • raw query;

  • normalized task;

  • intended use;

  • requested output;

  • decision consequence;

  • temporal scope;

  • domain;

  • and source constraints.

This prevents later task drift.

14.2 Decision consequence

Protocol depth scales with consequence. A low-consequence query may use a limited record. High or critical consequence requires stronger testing, human review, replayability, and narrower authorized use.

14.3 Object resolution

The record specifies:

  • object type;

  • included scope;

  • excluded scope;

  • adjacent objects;

  • authorized transfers;

  • and blocked transfers.

Where ambiguity remains material, RELEASE is unavailable.

14.4 Claim decomposition

The object is decomposed into:

C={c₁,c₂,…,cₘ}

Each claim receives:

  • type;

  • strength;

  • scope;

  • dependencies;

  • load-bearing status;

  • evidence burden;

  • and maturity.

14.5 Overdecomposition

Decomposition must not destroy the original relation among claims. A proposition containing an effect and a non-inferiority condition must preserve both components and the relational claim connecting them.

14.6 Maturity classification

Each maturity-sensitive component receives:

M_(claimed)

and:

M_(supported)

This step occurs before global evaluation.

14.7 Source inventories

The protocol separates:

S_(available),S_(retrieved),S_(used),S_(cited),S_(excluded),S_(inaccessible),S_(counterfactual)

A source never retrieved cannot be described as evaluated and rejected.

14.8 Source roles

Source roles include:

  • primary empirical evidence;

  • secondary synthesis;

  • official authority;

  • technical documentation;

  • methodological precedent;

  • historical primary record;

  • conceptual provenance;

  • criticism;

  • contradiction;

  • and self-description.

Source role is claim-relative.

14.9 Claim–evidence graph

The evidence architecture is:

□ ℰ=(C,S,L)

where L contains relations such as:

  • full support;

  • partial support;

  • contextual support;

  • methodological support;

  • conceptual precedent;

  • contradiction;

  • and no material support.

Each material link records:

  • directness;

  • strength match;

  • scope match;

  • maturity match;

  • contradiction;

  • and source dependency.

14.10 Missing edges

The graph also represents missing evidence:

  • no source supports the causal step;

  • no implementation evidence supports the product claim;

  • no external replication supports validation;

  • or a decisive source is inaccessible.

Missing evidence is a recorded requirement, not a source.

14.11 Baseline judgment

After object, claim, maturity, and evidence mapping, the protocol produces:

B=Baseline Judgment State

The baseline contains the provisional class, selected sources, decisive evidence, citations, rationale, and limitations.

It remains preserved after correction.

14.12 Sequence

[ | Task Registration ; | →Object Resolution ; | →Claim Decomposition ; | →Maturity Classification ; | →Source Inventory ; | →Claim–Evidence Graph ; | →B ]

15. Sensitivity Testing

15.1 Purpose

Sensitivity testing asks:

Would the classification change materially if a non-content variable were altered while the substantive evidential object remained appropriately controlled?

The protocol does not seek invariance under every change. Official status, authenticity, version, and new evidence may legitimately alter the result.

15.2 Test families

Possible test families include:

  • provenance counterfactual;

  • citation-popularity masking;

  • source-identity counterfactual;

  • source-order permutation;

  • candidate-order permutation;

  • context-position test;

  • semantic reformulation;

  • novelty framing;

  • rhetorical presentation;

  • repeated run;

  • cross-system comparison;

  • human blinding;

  • and source-set expansion.

Not every event requires every test. The applicable policy pack determines the minimum set.

15.3 Provenance counterfactual

The protocol varies author identity, affiliation, venue, institutional label, or source category while preserving content and evidence.

It records changes in class, confidence, rank, source selection, citation, rationale, maturity, and disposition.

15.4 Citation-popularity counterfactual

Citation count or popularity labels may be hidden or altered where popularity is not constitutive of the task.

A change establishes popularity sensitivity. It does not establish the substantive value of the low-popularity source.

15.5 Source-identity counterfactual

Source identity may be varied only where the content remains valid and identity does not establish authenticity or formal authority.

15.6 Presentation permutation

The protocol varies source order, answer order, candidate order, or context position.

It distinguishes detectable variation from material variation.

15.7 Repeated runs

Repeated-run testing holds nominal conditions constant. Benign variation among equivalent secondary sources is distinct from a change in class, decisive evidence, citation, or disposition.

15.8 Cross-system comparison

Cross-system disagreement identifies a system-dependence question. It does not identify automatically which system is correct.

The protocol examines whether the difference reflects source access, retrieval policy, object definition, evidence interpretation, or execution variance.

15.9 Semantic reformulation

Equivalent formulations may use conventional terminology, defined novel terminology, plain language, or formal notation.

A valid test must preserve object, scope, strength, evidence, maturity, and practical implication.

15.10 Novelty framing

The same contribution may be described as established, new, independent, unconventional, overlooked, or paradigm-changing.

The test is required by symmetry.

15.11 Rhetorical presentation

The protocol may vary polish, structure, formal notation, narrative confidence, or visual completeness while preserving substantive content.

Legitimate clarity improvements must not be mislabeled as rhetorical substitution.

15.12 Counterfactual validity

A valid intervention preserves, within declared tolerance:

  • facts;

  • evidence;

  • claim strength;

  • scope;

  • maturity;

  • clarity;

  • object identity;

  • and task relevance.

A contaminated test triggers:

F15=Counterfactual Contamination

15.13 P3 rule

The canonical rule is:

□ P3⇔material change under a valid and relevant intervention

An invalid counterfactual cannot establish P3.

If a required test is invalid and no sufficient valid alternative remains, the process becomes P5. If the invalid test is non-blocking and sufficient valid assessment remains, the process can be no stronger than P2 and the limitation must be recorded.

15.14 Materiality

Materiality levels are:

  • L1 — informational;

  • L2 — qualifying;

  • L3 — material;

  • L4 — blocking.

A material change may include:

  • substantive-class reversal;

  • decisive-source change;

  • inclusion-cutoff crossing;

  • citation failure;

  • maturity change;

  • or disposition change.

15.15 Observable sensitivity

A valid test may establish that changing affiliation altered the output. It does not establish the complete hidden causal pathway, relevant training examples, or model motive.

15.16 Null results

The correct statement is:

No material effect was observed under the tested intervention.

The protocol must not generalize this into absence of all bias or sensitivity.

15.17 Interaction effects

Prestige may matter only for borderline evidence. Order may matter only in long contexts. Novelty framing may matter only where maturity is ambiguous.

The benchmark therefore stratifies effects rather than treating one average as universal.

16. Citation, Coverage, and Redundancy

16.1 Citation as evidential contract

A citation links a generated claim to a source under a particular wording, scope, strength, and evidential role.

The protocol treats this relation as an evidential contract.

16.2 Citation verification

Citation verification asks:

  1. Does the source exist?

  2. Is the identity correct?

  3. Does the cited passage support the local claim?

  4. Does support match wording strength?

  5. Does support match scope?

  6. Does support match maturity?

  7. Does the source contain material contradiction?

  8. Is the source primary, secondary, dependent, or contextual?

16.3 Partial support

A source may support a weaker claim. The correct result may be S2 with an explicit supported reformulation.

16.4 Contradictory citation

A source may be topically relevant but contradict the proposition it is cited to support. This may trigger:

F14=Citation Presence Mistaken for Support

16.5 Fabricated citation

A fabricated or nonexistent load-bearing citation breaks the verification chain and is L4 by default.

16.6 Precision and coverage

The protocol separates:

Citation Support Precision

from:

Material Claim Coverage

Accurate citations attached only to peripheral claims do not support the central argument.

16.7 Load-bearing coverage

Citation review prioritizes claims determining class, maturity, action, or disposition.

16.8 Redundancy

Relevant redundancy types include:

  • exact duplication;

  • syndicated copies;

  • shared primary source;

  • shared dataset;

  • shared methodology;

  • and citation inheritance.

The protocol defines:

F13=Redundancy Mistaken for Corroboration

16.9 Independent contribution

A source contributes independently where it adds a distinct evidential path, such as:

  • independent replication;

  • different dataset;

  • different method;

  • different population;

  • different source role;

  • or credible contradiction.

16.10 Relevance-constrained evidential diversity

The protocol supports:

□ Relevance-Constrained Evidential Diversity

It does not require ideological parity or forced opposition.

16.11 Consensus

A field may contain genuine consensus. The protocol should recognize independent methodological convergence without manufacturing disagreement.

16.12 Source-set incompleteness

The protocol defines:

F17=Source-Set Incompleteness

where a material source class, decisive evidence, or relevant contradiction is absent.

16.13 Citation review and revision

Citation review may alter the result:

S1→S2

where the source supports only a narrower claim;

S1→S3

where decisive support is unavailable;

or:

S1→S4

where the cited source contradicts the claim.

17. Uncertainty, Disagreement, and Revision

17.1 Uncertainty types

The protocol distinguishes:

  • epistemic uncertainty;

  • process uncertainty;

  • ontological uncertainty;

  • policy uncertainty;

  • temporal uncertainty;

  • source-access uncertainty;

  • measurement uncertainty;

  • and reviewer uncertainty.

These forms require different corrective actions.

17.2 Epistemic uncertainty

Epistemic uncertainty concerns incomplete, noisy, conflicting, or weak evidence.

The record identifies the affected claim, missing evidence, possible judgment change, and resolution condition.

17.3 Process uncertainty

Process uncertainty concerns untested variables, one-system limitation, incomplete verification, or uncertain source use. It affects Jₚ and does not automatically weaken Jₛ.

17.4 Ontological uncertainty

Ontological uncertainty concerns what the object is. It may require decomposition or S5.

17.5 Policy uncertainty

Policy uncertainty arises where evidence is clear but authorized action is undefined. The protocol must not disguise a policy gap as evidential uncertainty.

17.6 Disagreement classes

The protocol distinguishes:

  • D1 — benign disagreement;

  • D2 — scope disagreement;

  • D3 — evidential disagreement;

  • D4 — ontological disagreement;

  • D5 — process or policy disagreement.

17.7 Averaging is not resolution

An S1 and S4 disagreement cannot be resolved responsibly by averaging. The protocol must identify whether the evaluators used different objects, evidence, scopes, or causal standards.

17.8 Human review trigger

Human review may be required where:

  • D3–D5 remains material;

  • the task is high-consequence;

  • specialist knowledge is necessary;

  • P3 or P4 is observed;

  • a blocking flag exists;

  • or S5 cannot be resolved automatically.

17.9 Revision conditions

Every consequential classification should state what could change it.

Revision conditions may include:

  • new direct evidence;

  • independent replication;

  • contradiction;

  • retraction;

  • source access;

  • maturity transition;

  • system change;

  • policy change;

  • or temporal expiry.

A generic statement such as “more research is needed” is not operative.

17.10 Positive and negative revision conditions

A positive condition identifies evidence that could raise standing. A negative condition identifies evidence that could weaken or defeat the result.

Recognition must remain answerable to both.

17.11 Process revision

A process revision condition may require reassessment after:

  • model update;

  • changed search provider;

  • corrected retrieval index;

  • improved citation verifier;

  • or repaired counterfactual.

17.12 Expiry

Time-sensitive classifications should expire. An expired record remains auditable but is no longer current.

17.13 Correctionless classification

The protocol defines:

F10=Correctionless Classification

as a consequential classification lacking an operative path through which relevant evidence can revise it.

17.14 Unresolved evaluator disagreement

Material disagreement that is concealed, averaged, or left unresolved may trigger:

F18=Unresolved Evaluator Disagreement

18. Protocol Dispositions

18.1 Purpose

Substantive and process judgments describe the classification. Disposition determines how the result may be used.

G∈{RELEASE,RELEASE WITH LIMITATION,HOLD,REASSESS,INVALID}

Exactly one primary disposition is recorded.

18.2 RELEASE

RELEASE authorizes the classification for the declared purpose. It may authorize a supported claim or a robust evidence-based rejection.

⟨S4,P1,RELEASE⟩

is therefore valid.

18.3 RELEASE WITH LIMITATION

The classification remains usable within explicit boundaries concerning domain, system, time, source access, process limitation, or intended use.

18.4 HOLD

The evaluated claim or decision must not be treated as settled. HOLD requires a specific resolution condition and an affected claim.

The canonical case is:

□ ⟨S3,P1,HOLD⟩

This means the protocol robustly concludes that the evidence is insufficient. The record may remain active, published, cited, and audited while the underlying claim remains unsettled.

18.5 REASSESS

REASSESS requires the classification process to be repeated under corrected conditions.

Triggers include:

  • material provenance sensitivity;

  • order sensitivity;

  • unstable execution;

  • invalid counterfactual;

  • citation-verification defect;

  • or source-dependency error.

18.6 INVALID

INVALID means the current run or record cannot support responsible use because of a structural defect.

It does not establish that the evaluated claim is false.

18.7 Required actions

Additional actions are recorded separately rather than as a second disposition. Examples include:

  • verify citation;

  • expand source set;

  • repeat test;

  • resolve object;

  • clarify claim;

  • lower maturity;

  • conduct human review;

  • perform replay.

18.8 Record lifecycle

Record lifecycle is separate from disposition.

Lifecycle values are:

  • draft;

  • active;

  • superseded;

  • invalidated;

  • expired;

  • archived.

The following combination is valid:

record_lifecycle: active substantive_judgment: S3 process_judgment: P1 protocol_disposition: HOLD

18.9 Default decision matrix

Substantive class
Process class
Default disposition
S1
P1
RELEASE
S1
P2
RELEASE WITH LIMITATION
S1
P3/P4
REASSESS or bounded limited release
Any
P5
INVALID
S2
P1/P2
RELEASE WITH LIMITATION
S2
P3/P4
REASSESS
S3
P1/P2
HOLD
S3
P3/P4
REASSESS; underlying claim remains unsettled
S4
P1
RELEASE evidence-based rejection
S4
P2
RELEASE WITH LIMITATION
S4
P3/P4
REASSESS
S5
P1/P2
HOLD
S5
P3/P4
REASSESS
S5
P5
INVALID

The policy pack may strengthen these defaults according to consequence.

18.10 Full sequence

[ | Task Registration ; | →Object Resolution ; | →Claim Decomposition ; | →Maturity Classification ; | →Source Inventory ; | →Claim–Evidence Graph ; | →B ; | →Sensitivity Testing ; | →Citation and Coverage Review ; | →Uncertainty and Disagreement ; | →⟨Jₛ,Jₚ⟩ ; | →G ; | →R ; | →Λ ]

Part IV Synthesis

Part IV separates substantive judgment from process robustness, makes P3 dependent on a valid material intervention, treats citation quality as both local support and material coverage, and requires an operative correction path. Part V governs the resulting record and the protocol’s own validation.

Part V — Audit, Validation, and Limits

19. The Epistemic Classification Record

19.1 The answer is not the complete governed artifact

A visible answer may contain a conclusion, citations, a confidence statement, and a recommendation. It may omit the information required to reconstruct how the result was produced.

The requirement to preserve lifecycle, context, evaluation, and review information is consistent with NIST’s broader treatment of AI risk as a sociotechnical and lifecycle-governance problem (Schwartz et al. 2022; Tabassi 2023; Autio et al. 2024).

The protocol therefore treats the Epistemic Classification Record as a first-class output.

19.2 Record purpose

The record is a versioned account of an observable classification process. It does not attempt to reproduce hidden chain of thought, complete weight-level causation, or proprietary system logic.

It preserves the relation among:

Task,Object,Evidence,Intervention,Judgment,Disposition,Revision

19.3 Six record layers

Identity

  • record ID and version;

  • protocol and policy version;

  • system and execution identity;

  • time and jurisdiction.

Object

  • query and intended use;

  • evaluated object;

  • included and excluded scope;

  • claim inventory;

  • maturity.

Evidence

  • source boundaries;

  • source roles;

  • claim–evidence relations;

  • contradiction;

  • missing evidence;

  • source dependency.

Testing

  • provenance counterfactuals;

  • presentation permutations;

  • repeated runs;

  • citation verification;

  • test validity.

Judgment

  • baseline state;

  • substantive class;

  • process class;

  • failure flags;

  • uncertainty;

  • disposition.

Correction

  • revision conditions;

  • review;

  • override;

  • replay;

  • expiry;

  • supersession.

19.4 Versioning

The record distinguishes:

v_(protocol),v_(policy),v_(system),v_(source),v_(record)

A new judgment should not overwrite its predecessor as though the earlier state never existed.

19.5 Lifecycle

The administrative states are:

  • draft;

  • active;

  • superseded;

  • invalidated;

  • expired;

  • archived.

These remain separate from RELEASE, HOLD, REASSESS, and INVALID.

19.6 Baseline and final state

Both the baseline and final state remain preserved. This permits later analysis of whether the protocol changed:

  • class;

  • maturity;

  • citation;

  • evidence path;

  • process judgment;

  • or disposition.

19.7 Traceability

Every load-bearing claim should be traceable to its wording, evidence links, contradiction status, maturity, baseline class, final class, and revision conditions.

19.8 Auditability

Auditability is:

The ability of an authorized reviewer to reconstruct the declared classification process from the retained record.

It does not require release of private data, proprietary details, full prompts, or hidden internal reasoning. Redactions must be declared.

19.9 Explanation and audit

A fluent explanation may be generated after a decision. An audit record should allow inspection of the relevance criterion, claim–source relation, competing sources, source boundary, and intervention results.

19.10 Auditability is not correctness

□ Auditability⇏Correctness

Auditability provides a route to locate and correct error. It does not remove error automatically.

19.11 Replayability

The protocol distinguishes:

  • exact reproducibility;

  • material equivalence;

  • and replayability.

Replayability means the system can repeat the declared classification under sufficiently preserved conditions and explain material deviation.

19.12 Decision-relevant completeness

The record should preserve material transformations and load-bearing evidence without becoming an unusable archive of every token or transient score.

19.13 Data minimization

Identity-sensitive tests should retain only the personal data necessary for the declared task, review, and correction. Synthetic labels and controlled substitutions are preferred where possible.

19.14 Integrity

Technical integrity mechanisms may show that the record was not altered after creation.

□ Record Integrity⇏Judgment Correctness

20. Human Review, Override, and Replay

20.1 Human review is governed

Human review can add domain knowledge, contextual understanding, authority, and responsibility. It can also add prestige sensitivity, inconsistency, conflict of interest, or unsupported intuition.

Human and LLM judges have both demonstrated sensitivity to controlled perturbations in tested settings (Chen et al. 2024).

Human review is therefore another governed classification event.

20.2 Review triggers

Human review may be required where:

  • the object remains materially ambiguous;

  • specialist knowledge is necessary;

  • consequence is high or critical;

  • D3–D5 disagreement remains;

  • P3 or P4 is observed;

  • S5 cannot be resolved automatically;

  • a blocking flag exists;

  • or an appeal challenges a consequential field.

20.3 Reviewer jurisdiction

Reviewer roles may include:

  • domain expert;

  • evidence-method reviewer;

  • policy authority;

  • legal authority;

  • protocol auditor;

  • system operator;

  • decision owner;

  • affected-party representative.

The record identifies competence, scope, authority, and conflicts of interest.

20.4 Review packet

The reviewer receives the task, object, claims, maturity, evidence graph, baseline, sensitivity results, citation review, uncertainty, failure flags, and proposed disposition.

20.5 Blinded review

Where provenance is not constitutive, the protocol may compare visible and blinded human judgments.

Blinding is conditional and reversible. It must not remove authenticity, authority, or conflict-of-interest information required by the task.

20.6 Review outcomes

A reviewer may:

  • confirm the judgment;

  • narrow scope;

  • change the supported formulation;

  • alter maturity;

  • change Jₛ or Jₚ;

  • change disposition;

  • return the record to evidence mapping;

  • or invalidate the run.

20.7 Override

An override changes a material field through an authorized intervention.

Override classes include:

  • evidential override;

  • object override;

  • policy override;

  • process override;

  • authority override.

The pre-override and post-override states both remain visible.

20.8 Unsupported override

An override relying only on seniority, institutional position, reputation, or unexplained intuition is unsupported.

A decision owner may possess authority to permit limited use. This must be represented as a policy or authority override, not as evidential refutation.

20.9 Appeal

An affected party may challenge a field-level element such as object definition, source omission, citation verification, maturity, process class, failure flag, or disposition.

20.10 Replay

A replay repeats a classification under declared preserved conditions.

Replay results are:

  • exact reproduction;

  • equivalent reproduction;

  • explained deviation;

  • unexplained deviation;

  • not reproducible.

20.11 Replay and instability

Unexplained material deviation may support P4 and F12. Explained change after new evidence or process repair is successful correction, not instability.

20.12 Shadow review

Before operational authority, the protocol may run in shadow mode. It produces classifications and flags without controlling the real decision.

20.13 Limited operational authority

Early permissible actions include requesting citation verification, reassessment, limitation, blinded review, or source-set expansion.

The protocol should not begin by automatically suppressing sources, penalizing authors, or permanently altering public rankings.

21. The Adversarial Benchmark

21.1 Why adversarial validation is necessary

A protocol can appear successful by embedding its assumptions into the benchmark. It may label every unfamiliar source strong, place every difficult case on HOLD, or improve citation precision by citing almost nothing.

The benchmark must contain cases capable of demonstrating protocol failure.

21.2 Validation targets

The benchmark evaluates:

V1=Detection

V2=Discrimination

V3=Symmetry

V4=Corrective Value

Detection concerns known controlled defects. Discrimination concerns legitimate versus illegitimate use of the same variable. Symmetry concerns equivalent evidence standards across prestige and novelty conditions. Corrective value concerns improvement relative to simpler baselines.

21.3 Factorial core

The general four-cell design crosses evidence quality with prestige, familiarity, or novelty condition.

The expected behavior is to recognize strong items and reject or qualify weak items across both conditions.

21.4 Strong and weak items

A strong item contains a resolved object, bounded claims, adequate evidence, valid inference, accurate maturity, and explicit limits.

A weak item contains a controlled defect such as causal promotion, invalid generalization, citation mismatch, maturity inflation, object conflation, contradiction concealment, or load-bearing reasoning failure.

Where feasible:

Weak Item=Strong Item+One Controlled Defect

21.5 Controls

The benchmark includes:

  • positive controls;

  • negative controls;

  • legitimate-provenance controls;

  • anti-symmetry controls;

  • excessive-HOLD controls.

21.6 Baselines

The comparison conditions are:

  • B0 — Unstructured Baseline;

  • B1 — Neutral Evidence Rubric;

  • B2 — Self-Applied Protocol;

  • B3 — External Protocol Runner;

  • B4 — Human Expert Condition.

B1 is load-bearing. If a simpler evidence rubric performs equivalently, full protocol complexity may not be justified.

21.7 Twelve suites

The benchmark contains twelve suites:

  • A — Prestige;

  • B — Citation Popularity;

  • C — Source Identity;

  • D — Order and Position;

  • E — Semantic Familiarity;

  • F — Novelty;

  • G — Rhetorical Coherence;

  • H — Maturity;

  • I — Object Conflation;

  • J — Citation Support;

  • K — Redundancy and Evidential Diversity;

  • L — Run and Cross-System Stability.

Each suite contains eight core episodes, producing:

12×8=96

core benchmark episodes.

21.8 Adjudicated targets

Each episode receives an Adjudicated Benchmark Target that may include an expected class, acceptable alternatives, required flags, prohibited flags, acceptable dispositions, and revision conditions.

The target need not be one exact wording.

21.9 Counterfactual-validity audit

Every pair is audited for preservation of object, claim, evidence, scope, maturity, clarity, authenticity, and task relevance.

Invalid pairs cannot support sensitivity claims.

21.10 Excessive HOLD

Clear recognition and rejection cases are included to detect decision avoidance.

21.11 Validation stages

Validation proceeds through:

  1. construction audit;

  2. pilot;

  3. controlled benchmark;

  4. protocol-intervention study;

  5. external replication;

  6. shadow evaluation;

  7. limited operational trial.

No stage implies automatic promotion to the next.

21.12 Validation dispositions

Benchmark validation uses PASS, HOLD, and FAIL. These are validation outcomes, not protocol dispositions for individual classification objects.

A bounded PASS must state the tested component, system, domain, and corpus.

22. Metrics and Validation

22.1 Metrics are candidate instruments

A formula is not valid because it is precise. Every metric requires construct definition, scoring reliability, threshold calibration, domain testing, and evidence that improvement matters.

22.2 Metric families

The protocol evaluates:

  • substantive accuracy;

  • sensitivity;

  • citation support;

  • material-claim coverage;

  • maturity;

  • object resolution;

  • symmetry;

  • correction;

  • disposition;

  • auditability;

  • human review;

  • replay;

  • cost;

  • latency.

22.3 No global protocol score

One score would conceal trade-offs among accuracy, sensitivity, coverage, HOLD inflation, cost, and auditability.

The preferred output is a scorecard with primary, secondary, exploratory, and guardrail metrics.

22.4 Judgment distance

Substantive classes are not treated as one simple ordinal scale. S5 is not merely worse than S4.

Where distance is needed, the protocol uses versioned cost matrices and preserves exact categorical changes beside any composite value.

22.5 Sensitivity metrics

Candidate sensitivity measures include:

  • Provenance Sensitivity Index;

  • Counterfactual Flip Rate;

  • Order Sensitivity Index;

  • Run Instability Index;

  • Cross-System Decision Divergence Rate.

A sensitivity measure records change. Task relevance and counterfactual validity determine interpretation.

22.6 Citation metrics

Primary citation measures include:

  • Citation Support Precision;

  • Material Claim Coverage;

  • Load-Bearing Claim Coverage;

  • Citation Strength Match;

  • Citation Scope Match;

  • Fabricated Citation Rate.

22.7 Symmetry metrics

The benchmark reports:

  • Prestige Symmetry Gap;

  • Novelty Symmetry Gap;

  • Semantic Familiarity Penalty Gap;

  • Directional Compensation Rate.

Low gap must be reported with absolute accuracy because equal poor performance is not success.

22.8 Maturity and object metrics

Candidate measures include:

  • Maturity Promotion Error;

  • Maturity Collapse Error;

  • Mixed-Maturity Differentiation Rate;

  • Object Resolution Accuracy;

  • Object Conflation Rate;

  • Authorized Transfer Precision and Recall.

22.9 HOLD and decision retention

The protocol measures:

EHR=Excessive HOLD Rate

and:

DRR=Decision-Retention Rate

A protocol that avoids most decisions may reduce visible errors while losing operational value.

22.10 Non-inferiority

Protocol use must preserve substantive performance within calibrated margins.

For example:

CA_(protocol)−CA_(baseline)≥−δ_(CA)

The margins remain uncalibrated and task-specific.

22.11 Primary pilot scorecard

The primary 96-episode pilot reports:

  1. Content Accuracy;

  2. Load-Bearing Accuracy;

  3. Counterfactual Flip Rate;

  4. Citation Support Precision;

  5. Material Claim Coverage;

  6. Maturity Promotion Error;

  7. Object Conflation Rate;

  8. Prestige Symmetry Gap;

  9. Novelty Symmetry Gap;

  10. Excessive HOLD Rate;

  11. Decision-Retention Rate;

  12. incremental cost and latency.

Guardrails include Counterfactual Validity Rate, Counterfactual Contamination Rate, Fabricated Citation Rate, RELEASE Precision, Record Schema Validity, and L4 event count.

22.12 Calibration

Calibration proceeds through construct review, scoring manual, inter-rater study, positive and negative controls, pilot distribution, threshold setting, external replication, and drift review.

22.13 Metric gaming

The benchmark must detect strategies such as:

  • citing only easy claims;

  • weakening every claim;

  • ignoring legitimate provenance;

  • placing everything on HOLD;

  • assigning identical results to every cell;

  • producing verbose but low-value audit records.

22.14 Governing rule

□ Metric Improvement⇏Protocol Validation

23. Falsification and Failure Conditions

23.1 A correction protocol must be correctable

The protocol must identify evidence that would require it to be narrowed, simplified, revised, or rejected.

23.2 Conceptual failure

The architecture is weakened if:

  • epistemic allocation adds no useful distinction beyond ordinary retrieval;

  • dual judgment adds no information beyond existing labels;

  • object resolution does not reduce meaningful error;

  • or revision conditions cannot be made operative.

23.3 Empirical failure

The empirical motivation must be narrowed if reported non-content sensitivities fail to replicate or are explained adequately by legitimate task relevance.

23.4 Unified-mechanism failure

Ontology-preserving conformity should be removed as a unified mechanism if its components do not co-occur, object persistence is explained by evidence quality, or the concept adds no value beyond separate flags.

23.5 Symmetry failure

The protocol fails centrally if it favors independent sources, penalizes prestige by default, rewards novelty language, or interprets criticism of itself as evidence of conformity.

23.6 Counterfactual failure

The counterfactual layer should be narrowed where valid content-preserving interventions cannot be constructed reliably or contamination remains high.

23.7 Detection and false-positive failure

The protocol fails if it misses controlled defects or repeatedly flags legitimate authority, harmless reordering, or task-relevant popularity as bias.

23.8 Accuracy failure

The protocol should not proceed where it materially reduces substantive accuracy without a separately justified benefit.

23.9 Coverage failure

Improved citation precision does not compensate for unsupported central claims.

23.10 Excessive-HOLD failure

A protocol that withholds every difficult decision demonstrates avoidance, not governance.

23.11 Correction failure

Generic phrases such as “future research” do not satisfy the correction requirement.

23.12 Audit failure

Auditability fails where a qualified reviewer cannot reconstruct the object, evidence, intervention, judgment, or override. It can fail through omission or unstructured excess.

23.13 Human-review failure

Human review fails where reviewers lack jurisdiction, overrides lack support, prestige replaces analysis, or disagreement is erased.

23.14 Policy failure

A system may execute the protocol correctly under a defective policy. Policy validation remains separate from execution validation.

23.15 Operational failure

The protocol may be disproportionate in cost, latency, storage, or review burden. The appropriate response may be tiering, simplification, or narrower scope.

23.16 Security and privacy failure

The protocol should be restricted where it requires unnecessary personal data or exposes manipulations that facilitate ranking evasion or deceptive provenance.

23.17 Simpler-baseline failure

The full architecture should not be retained where a simpler method achieves equivalent benefit with fewer costs and failure modes.

23.18 External-replication failure

Positive internal results must be narrowed where independent teams cannot reproduce detection, symmetry, accuracy preservation, or audit value.

23.19 Self-sealing prohibition

A framework that treats success, failure, criticism, and acceptance as confirmation is correctionless.

□ Protocol Architecture⇏Protocol Validity

□ Protocol Compliance⇏Epistemic Improvement

24. System Boundary and Non-Claims

24.1 The system is sociotechnical

The relevant system may include:

[ | Corpus Construction+Indexing+Query Rewriting ; | +Retrieval+Ranking+Context Selection ; | +Generation+Citation+Evaluation ; | +Interface+Human Review+Organizational Policy ]

This boundary is consistent with NIST guidance treating AI risk as distributed across actors, processes, deployment contexts, and lifecycle stages (Schwartz et al. 2022; Tabassi 2023; Autio et al. 2024).

24.2 Corpus boundary

A source absent from the accessible corpus cannot be retrieved. Absence may result from access restrictions, indexing policy, language, format, source availability, or recency.

The protocol cannot infer source weakness from absence alone.

24.3 Indexing boundary

A source may be present but poorly represented because of parsing, metadata loss, segmentation, or terminology mismatch.

24.4 Retrieval boundary

Retrieval may optimize relevance, recency, authority, popularity, or user preference. Where the objective is unknown, the protocol should not claim certainty about why a source was omitted.

24.5 Generation boundary

The generator may use retrieved sources, parametric knowledge, or both. Visible citations do not necessarily reveal the complete contribution of each source.

24.6 Citation boundary

Citation may be performed during generation, after generation, or through a separate attribution layer. A citation failure may originate in a different component from the substantive claim.

24.7 Evaluator boundary

An evaluator may be the same model, another model, a rules engine, a human, or a hybrid. Agreement by another model is not independent validation by default.

24.8 Organizational boundary

Organizations determine allowed sources, thresholds, review triggers, authorized use, retention, and downstream action.

The distinction between technical execution and organizational governance is consistent with the AI RMF’s allocation of risk-management functions across organizations that design, develop, deploy, or use AI systems (Tabassi 2023).

The protocol distinguishes execution failure, policy failure, and authority failure.

24.9 Introspection boundary

The protocol does not require hidden chain of thought, exact training-data influence, or complete internal priors.

It evaluates observable input, source environment, intervention, output, citation, and revision.

Unsupported internal-cause claims trigger:

F16=Unsupported Introspection Claim

24.10 Causal boundary

□ Counterfactual Effect⇏Complete Causal Model

24.11 Truth boundary

The protocol can improve object precision, evidence traceability, sensitivity visibility, and correction. It cannot guarantee truth.

A complete source set may be collectively wrong. A robust process may preserve a mistaken theory.

24.12 Neutrality boundary

The protocol does not seek equal outcomes. Evidence may justify unequal confidence, inclusion, and strong rejection.

24.13 Diversity boundary

The protocol does not require ideological, institutional, or minority-position quotas. It requires sufficient relevant evidential diversity.

24.14 Institutional boundary

The protocol is not anti-institutional. Institutions may provide authority, stable versioning, review, accountability, and correction.

24.15 Independent-source boundary

Independent, low-citation, and unconventional sources receive no compensatory standing. They may deserve examination; their classification remains evidence-dependent.

24.16 Novelty boundary

Novelty may justify discovery or comparison. It does not establish truth, originality, maturity, or corrective value.

24.17 Robustness boundary

□ Tested Robustness⇏Universal Robustness

24.18 Human boundary

Human review can add expertise and responsibility but remains fallible. Empirical comparison of human and LLM judges supports treating human review as a potentially sensitive intervention rather than automatic ground truth (Chen et al. 2024).

24.19 Automation boundary

Protocol findings do not directly authorize source suppression, author sanctions, permanent public ranking changes, or access restrictions.

24.20 Security boundary

The protocol should not be used to reverse-engineer ranking systems for manipulation, construct deceptive provenance, manufacture credibility, or optimize disinformation inclusion.

24.21 Privacy boundary

Identity-sensitive testing should minimize real personal data and use synthetic or approved research procedures where feasible.

24.22 Certification boundary

The protocol is not a certification standard. Certification would require recognized authority, validated methods, calibrated thresholds, audit procedures, and defined scope.

24.23 LoopGuard-AI boundary

The protocol may later inform an evaluator, evidence record, correction gate, replay layer, or reassessment mechanism within LoopGuard-AI.

No implementation or validated integration is claimed.

24.24 RATIUM.AI boundary

RATIUM.AI provides conceptual provenance for parts of the dossier. It is not the preferred benchmark source or a source deserving compensatory visibility.

The protocol does not request promotion, citation, ranking, or favorable classification for RATIUM.AI.

24.25 Explicit non-claims

The dossier does not establish:

  • universal AI suppression of unfamiliar ontologies;

  • general disadvantage to independent researchers;

  • universal institutional preference;

  • general novelty amplification;

  • complete hidden causation from observable sensitivity;

  • correctness from auditability;

  • truth from robustness;

  • validity from reproducibility;

  • correction from human override;

  • calibrated metrics;

  • benchmark validation;

  • production readiness;

  • certification;

  • or implementation within LoopGuard-AI.

24.26 Current maturity

□ Concept + Architecture + Validation Design

Part V Synthesis

The governed artifact is the versioned Epistemic Classification Record. Human review, override, replay, validation, metrics, and institutional use remain auditable and bounded by explicit falsification, security, privacy, and maturity limits.

Conclusion — Classification Must Remain Answerable to Correction

AI-mediated systems construct the effective evidence environment through which questions become answerable. Sources may be available but unretrieved, retrieved but omitted, present but unused, or cited without supporting the decisive claim. Unequal allocation is unavoidable; the governance issue is whether it remains connected to the task, object, evidence, maturity, and a correction path.

The empirical basis is bounded: non-content sensitivity is observable and testable in particular systems and tasks. Citation popularity, authorship metadata, affiliation, source identity, order, context position, linguistic formulation, retrieval design, and execution conditions can affect output. These findings do not establish a universal conformity mechanism.

The protocol therefore begins with object resolution and preserves maturity from concept through certification. It reports separately:

Jₛ=Substantive Judgment Jₚ=Process-Robustness Judgment

This separation permits a claim to be supported yet process-sensitive, contradicted through a robust process, or robustly classified as insufficiently evidenced. It also enforces symmetry: strong familiar and unfamiliar material should be recognized; weak prestigious and marginal material should be qualified or rejected under the same evidential standard.

The complete governed artifact is the Epistemic Classification Record, which preserves the task, object, claims, maturity, evidence graph, baseline, interventions, citations, judgments, disposition, revision conditions, review, override, and replay. Auditability, reproducibility, robustness, and human review do not guarantee truth; they make error more locatable and correction more accountable.

The protocol’s final claim remains procedural:

A consequential AI-mediated classification should remain connected to a resolved object, a traceable evidence structure, a separately reported process judgment, and a specific path through which relevant evidence can alter its future use.

A classification becomes governable not when challenge is eliminated, but when relevant challenge can produce justified, traceable revision.

References

Abolghasemi, Amin, Leif Azzopardi, Seyyed Hadi Hashemi, Maarten de Rijke, and Suzan Verberne. 2025. “Evaluation of Attribution Bias in Generator-Aware Retrieval-Augmented Large Language Models.” In Findings of the Association for Computational Linguistics: ACL 2025, 21105–21124. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-acl.1087.

Algaba, Andres, Carmen Mazijn, Vincent Holst, Floriano Tori, Sylvia Wenmackers, and Vincent Ginis. 2025. “Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias.” In Findings of the Association for Computational Linguistics: NAACL 2025, 6844–6879. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-naacl.381.

Autio, Chloe, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall, and Kamie Roberts. 2024. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.600-1.

Chen, Guiming Hardy, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024. “Humans or LLMs as the Judge? A Study on Judgement Bias.” In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8301–8327. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.474.

Cheng, Jiali, and Hadi Amiri. 2025. “EqualizeIR: Mitigating Linguistic Biases in Retrieval Models.” In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2, 889–898. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-short.75.

Dai, Sunhao, Zhanshuo Cao, Wenjie Wang, Liang Pang, Jun Xu, See-Kiong Ng, and Tat-Seng Chua. 2025. “Media Source Matters More Than Content: Unveiling Political Bias in LLM-Generated Citations.” In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 17256–17276. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.emnlp-main.872.

Dycke, Nils, and Iryna Gurevych. 2026. “Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework.” Transactions of the Association for Computational Linguistics 14: 465–488. https://doi.org/10.1162/tacl.a.642.

Hsieh, Cheng-Yu, Yung-Sung Chuang, Chun-Liang Li, Zifeng Wang, Long Le, Abhishek Kumar, James Glass, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, and Tomas Pfister. 2024. “Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization.” In Findings of the Association for Computational Linguistics: ACL 2024, 14982–14995. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.890.

Khan, Saadat Hasan, Spencer Hong, Jingyu Wu, Kevin Lybarger, Youbing Yin, Erin Babinsky, and Daben Liu. 2026. “DF-RAG: Query-Aware Diversity for Retrieval-Augmented Generation.” In Findings of the Association for Computational Linguistics: EACL 2026, 2873–2894. Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-eacl.150.

Kirsten, Elisabeth, Jost Große Perdekamp, Qinyuan Wu, Mihir Upadhyay, Krishna P. Gummadi, and Muhammad Bilal Zafar. 2026. “Characterizing Web Search in The Age of Generative AI.” In Findings of the Association for Computational Linguistics: ACL 2026, 10827–10848. Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.526.

Liu, Nelson F., Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. “Lost in the Middle: How Language Models Use Long Contexts.” Transactions of the Association for Computational Linguistics 12: 157–173. https://doi.org/10.1162/tacl_a_00638.

Schwartz, Reva, Apostol Vassilev, Kristen K. Greene, Lori Perine, Andrew Burt, and Patrick Hall. 2022. Towards a Standard for Identifying and Managing Bias in Artificial Intelligence. NIST Special Publication 1270. National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.1270.

Shi, Lin, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush Vosoughi. 2025. “Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge.” In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, 292–314. Asian Federation of Natural Language Processing and Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.ijcnlp-long.18.

Tabassi, Elham. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1.

Vasu, Sai Suresh Macharla, Ivaxi Sheth, Hui-Po Wang, Ruta Binkyte, and Mario Fritz. 2026. “Justice in Judgment: Unveiling (Hidden) Bias in LLM-Assisted Peer Reviews.” In Findings of the Association for Computational Linguistics: ACL 2026, 307–330. Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.14.

Wang, Peiyi, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024. “Large Language Models Are Not Fair Evaluators.” In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Volume 1, 9440–9450. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.511.

Wei, Sheng-Lun, Cheng-Kuang Wu, Hen-Hsen Huang, and Hsin-Hsi Chen. 2024. “Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models.” In Findings of the Association for Computational Linguistics: ACL 2024, 5598–5621. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.333.

Xu, Yumo, Peng Qi, Jifan Chen, Kunlun Liu, Rujun Han, Lan Liu, Bonan Min, Vittorio Castelli, Arshit Gupta, and Zhiguo Wang. 2025. “CiteEval: Principle-Driven Citation Evaluation for Source Attribution.” In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Volume 1, 32759–32778. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.1574.

Appendix A — Evidence and Claim-Control Register

A.1 Purpose and Authority

This appendix governs the evidential and claim-status boundaries of the dossier. It maps material propositions to their object, claim type, maturity, evidence grade, source family, authorized scope, permitted inference, prohibited promotion, and revision condition.

Where a main-text statement appears broader than its register entry, the narrower entry governs.

A.2 Claim-Control Logic

The register blocks four unsupported transitions:

Bounded Finding⇏Universal Claim

Observable Sensitivity⇏Complete Causal Explanation

Conceptual Specification⇏Empirical Validation

Protocol Compliance⇏Epistemic Correctness

A.3 Evidence Grades

E1 — Direct Empirical Support

A controlled or systematic study bears directly on the stated phenomenon within a specified system, task, domain, population, or intervention.

E2 — Adjacent Empirical Support

A study establishes a component, analogous sensitivity, or relevant empirical condition without testing the complete dossier claim.

E3 — Methodological Support

A study supports an evaluation procedure, counterfactual method, measurement distinction, citation-verification approach, or mitigation principle.

G1 — Governance or Standards Basis

A recognized governance publication supports lifecycle risk management, documentation, testing, evaluation, verification, or continuous review.

C1 — Original Conceptual or Architectural Contribution

The proposition is introduced, synthesized, or formally specified by this dossier.

N1 — Normative Governance Requirement

The proposition states what a governable classification process should require.

H — HOLD

The proposition is plausible, motivated, or operationally decomposable but not sufficiently established.

X — Prohibited Inference

The proposition is unsupported, self-protective, logically invalid, or incompatible with the protocol’s symmetry and maturity boundaries.

A.4 Claim Status

The register statuses are:

  • SUPPORTED WITHIN SCOPE;

  • PARTIALLY SUPPORTED;

  • METHODOLOGICALLY SUPPORTED;

  • CONCEPTUALLY SPECIFIED;

  • NORMATIVELY PROPOSED;

  • HOLD;

  • PROHIBITED.

These are not the S1–S5 classes assigned within a protocol run.

A.5 Claim-Record Schema

□ 𝒞ᵢ^(reg)=⟨id,O,T,M,Γ,S,Ω,L,P,X,R,D,V⟩

where:

  • id: claim identifier;

  • O: object;

  • T: claim type;

  • M: maturity;

  • Γ: primary evidence grade;

  • S: source family;

  • Ω: authorized scope;

  • L: limitations;

  • P: permitted inference;

  • X: prohibited promotion;

  • R: revision condition;

  • D: dossier location;

  • V: version.

Each claim receives exactly one primary evidence grade. Additional functions are recorded as conceptual basis, methodological use, governance basis, or normative use.

A.6 Source-Family Register

Identifier
Source family
Canonical sources
Authorized use
SF-01
Citation popularity and scholarly recommendation
Algaba et al. 2025
Bounded high-citation preference; popularity counterfactuals
SF-02
Authorship metadata and attribution
Abolghasemi et al. 2025
Authorship sensitivity in tested RAG attribution
SF-03
Institutional affiliation in peer review
Vasu et al. 2026
Affiliation and metadata tests in comparable evaluation
SF-04
Source identity in political citation
Dai et al. 2025
Domain-bounded source-identity counterfactuals
SF-05
Option and candidate order
Wei et al. 2024; Wang et al. 2024; Shi et al. 2025
Permutation testing
SF-06
Long-context position
Liu et al. 2024; Hsieh et al. 2024
Context-position tests and effective-use distinction
SF-07
Linguistic complexity
Cheng and Amiri 2025
Query-complexity and reformulation tests
SF-08
Generative-search heterogeneity
Kirsten et al. 2026
System, execution, and time recording
SF-09
Citation evaluation
Xu et al. 2025
Claim-level citation verification
SF-10
Controlled reasoning defects
Dycke and Gurevych 2026
Counterfactual defect insertion
SF-11
Retrieval redundancy and diversity
Khan et al. 2026
Redundancy-aware, relevance-constrained retrieval
SF-12
Human and machine perturbation
Chen et al. 2024
Governed human review
SF-13
Sociotechnical governance
Schwartz et al. 2022; Tabassi 2023; Autio et al. 2024
Lifecycle, documentation, TEVV, sociotechnical scope

A.7 Major Empirical Claims

EC-01 — Citation-Popularity Preference

Authorized claim: In the tested scholarly-reference task, LLM-generated recommendations displayed a heightened preference for highly cited papers.

Primary grade: E1
Source family: SF-01
Status: SUPPORTED WITHIN SCOPE

Prohibited promotion: AI systems generally suppress low-citation research.

EC-02 — Persistence after selected controls

Authorized claim: The reported high-citation preference persisted after controls for publication year, title length, author count, and venue.

Primary grade: E1
Source family: SF-01

Prohibited promotion: Citation count was the sole causal determinant.

EC-03 — Authorship Metadata

Authorized claim: Authorship metadata can materially affect attribution quality in tested generator-aware RAG pipelines.

Primary grade: E1
Source family: SF-02
Methodological use: E3

Prohibited promotion: Named or institutional authors are generally trusted regardless of content.

EC-04 — Institutional Affiliation

Authorized claim: Institutional affiliation and related author metadata affected LLM-generated peer-review judgments in the tested settings.

Primary grade: E1
Source family: SF-03

Prohibited promotion: AI systems generally favor elite institutions.

EC-05 — Borderline Consequence

Authorized claim: Metadata effects can become consequential where modest changes cross an operational threshold.

Primary grade: E1
Source families: SF-03 and SF-05

EC-06 — Political Source Identity

Authorized claim: In the tested political-news setting, media-outlet identity affected citation selection beyond controlled content orientation.

Primary grade: E1
Source family: SF-04

Prohibited promotion: Source identity generally matters more than content.

EC-07 — Option Order and Token Representation

Authorized claim: LLM selection behavior can vary under option-order and option-token changes.

Primary grade: E1
Source family: SF-05

EC-08 — LLM-as-a-Judge Position

Authorized claim: Candidate-answer position can materially affect comparative LLM judgments in tested settings.

Primary grade: E1
Source family: SF-05
Methodological use: E3

EC-09 — Long-Context Position

Authorized claim: Relevant information can be used less effectively in disadvantaged positions within long contexts.

Primary grade: E1
Source family: SF-06

Prohibited promotion: Evidence in the middle is always ignored.

EC-10 — Linguistic Complexity

Authorized claim: Retrieval performance can vary across linguistically simple and complex query formulations.

Primary grade: E1
Source family: SF-07

Prohibited promotion: Retrieval systems preserve dominant ontologies.

EC-11 — Generative-Search Heterogeneity

Authorized claim: Generative-search systems can differ in source selection, internal and external knowledge use, retrieval footprint, diversity, and stability.

Primary grade: E1
Source family: SF-08

EC-12 — Run and Temporal Variation

Authorized claim: Generative-search outputs can vary across executions and time.

Primary grade: E1
Source family: SF-08

EC-13 — Citation Evaluation

Authorized claim: Citation quality requires evaluation of the query, generated claim, source, and retrieval context.

Primary grade: E3
Source family: SF-09
Status: METHODOLOGICALLY SUPPORTED

EC-14 — Automatic Review and Logic Faults

Authorized claim: In the tested counterfactual framework, inserted research-logic faults had no significant effect on the output reviews of the evaluated automatic-review approaches.

Primary grade: E1
Source family: SF-10
Methodological use: E3

EC-15 — Redundancy and Diversity

Authorized claim: Similarity-focused retrieval can produce redundant evidence, and relevance-constrained query-aware diversity improved performance in tested reasoning-intensive tasks.

Primary grade: E1
Source family: SF-11
Methodological use: E3

EC-16 — Human and Machine Perturbation

Authorized claim: Human and machine evaluators can both be affected by controlled perturbations.

Primary grade: E1
Source family: SF-12

Prohibited promotion: Human and machine failure modes are equivalent.

A.8 Methodological and Governance Claims

MG-01 — Ranking direction does not diagnose quality

Primary grade: C1
Methodological use: E3
Status: CONCEPTUALLY SPECIFIED

MG-02 — Provenance use versus substitution

Primary grade: C1
Empirical basis: E2

MG-03 — Observable sensitivity justifies process review

Primary grade: C1
Methodological use: E3

MG-04 — Object resolution precedes consequential classification

Primary grade: N1
Conceptual basis: C1

MG-05 — Claim type and maturity require separate control

Primary grade: N1
Conceptual basis: C1

MG-06 — Citation is a claim–source relation

Primary grade: C1
Methodological use: E3

MG-07 — Source count is not corroboration

Primary grade: E3
Conceptual use: C1

MG-08 — Uncertainty should be typed

Primary grade: C1
Normative use: N1

MG-09 — Revision conditions are required

Primary grade: N1
Governance basis: G1

MG-10 — Human override must be audited

Primary grade: N1
Governance basis: G1
Methodological use: E3

MG-11 — Auditability is not correctness

Primary grade: C1
Normative use: N1

MG-12 — Tested robustness is not truth

Primary grade: C1

MG-13 — Validation must be symmetric

Primary grade: N1
Conceptual basis: C1

MG-14 — No single score is sufficient

Primary grade: C1

MG-15 — Validation is component- and scope-specific

Primary grade: E3
Governance basis: G1

A.9 Original Contributions

The following are original or synthesized by this dossier and have primary grade C1 unless otherwise stated:

  • CA-01 — AI-Mediated Epistemic Allocation;

  • CA-02 — Epistemic Classification Event;

  • CA-03 — Canonical Classification Object 𝒜ₜ;

  • CA-04 — Dual Judgment;

  • CA-05 — Final Protocol Output 𝒥ₜ;

  • CA-06 — Object Types and Transfer Bridge;

  • CA-07 — Claim-Maturity Architecture;

  • CA-08 — Source Inventories;

  • CA-09 — Claim–Evidence Graph;

  • CA-10 — Failure Taxonomy F1–F18;

  • CA-11 — Failure Severity L1–L4;

  • CA-12 — S1–S5 and P1–P5;

  • CA-13 — Protocol Dispositions;

  • CA-14 — Revision-Condition Architecture, with normative use N1;

  • CA-15 — Epistemic Classification Record;

  • CA-16 — Human Review, Override, and Replay;

  • CA-17 — Adversarial Benchmark, status Validation Design;

  • CA-18 — Candidate Metrics, status Validation Design.

None is described as empirically validated.

A.10 HOLD Register

The following remain on HOLD:

  • H-01 — Ontology-preserving conformity as a unified empirical mechanism;

  • H-02 — Novelty amplification as a general AI tendency;

  • H-03 — Rhetorical-coherence substitution as a general ranking effect;

  • H-04 — Complete citation-network reinforcement loop;

  • H-05 — Universal institutional preference;

  • H-06 — Systematic independent-source suppression;

  • H-07 — Real-world protocol improvement;

  • H-08 — Dual-judgment superiority over simpler methods;

  • H-09 — Effectiveness of revision conditions;

  • H-10 — Metric validity and thresholds;

  • H-11 — Benchmark validity;

  • H-12 — Production readiness;

  • H-13 — Certification use;

  • H-14 — LoopGuard-AI implementation;

  • H-15 — Causal explanation of a named project’s visibility.

A.11 Prohibited Inference Register

The following are prohibited:

Low Ranking⇒Conformity

High Ranking⇒Epistemic Independence

Familiar Source⇒Prestige Substitution

Unfamiliarity⇒Corrective Standing

Independent Origin⇒Truth

Citation Count⇒Claim Support

Multiple Sources⇒Independent Evidence

Human Override⇒Ground Truth

Observed Sensitivity⇒Complete Hidden Mechanism

No Tested Effect⇒Absence of Untested Sensitivity

P1⇒S1

Auditable Record⇒Correct Judgment

Schema Compliance⇒Protocol Validation

HOLD⇒False

INVALID Record⇒False Object

Architecture⇒Prototype, Validation, or Production

The protocol may not be used to manipulate rankings, manufacture credibility, or promote RATIUM.AI or another named source.

A.12 Dependency Register

Authorized:

Bounded Sensitivity Evidence⇒Justification for Controlled Testing

Blocked:

Bounded Sensitivity Evidence⇏Universal Ontology-Preserving Mechanism

Authorized:

E3+C1⇒Candidate Protocol Architecture

Blocked:

Architecture+Benchmark Design⇏Validation

Blocked:

Metric Formula⇏Construct Validity

A.13 RATIUM.AI Boundary

RATIUM.AI may appear as conceptual provenance or publication origin. It must not appear as the preferred benchmark source, an object whose rank proves protocol quality, or a source deserving compensatory visibility.

RATIUM.AI Ranked Poorly⇏Ontology-Preserving Conformity

RATIUM.AI Ranked Highly⇏Epistemic Independence

A.14 LoopGuard-AI Boundary

The protocol may later inform evaluators, records, correction gates, replay layers, or reassessment triggers. No implementation or validated integration is claimed.

A.15 Update Rules

Claims may be promoted only through new relevant evidence. Repetition, elaboration, formalization, citation count, or architectural completeness do not convert C1 into E1.

Retractions, corrections, version changes, and standards revisions trigger review of dependent claims.

A.16 Master Status

Claim category
Claim category Status
Non-content sensitivity in tested tasks
Supported within scope
Provenance, order, popularity, and formulation as test variables
Methodologically supported
AI-mediated epistemic allocation
Conceptually specified
Dual judgment
Architecturally specified; effectiveness on HOLD
Failure taxonomy
Conceptually specified; completeness unvalidated
Ontology-preserving conformity
Unified mechanism on HOLD
Novelty amplification
General tendency on HOLD
Citation-network reinforcement loop
Longitudinal mechanism on HOLD
Epistemic Classification Record
Architecturally specified
Adversarial benchmark
Validation design; not executed
Candidate metrics
Defined; not calibrated
Production use
Not established
Certification
Not authorized
LoopGuard-AI implementation
Not established
Favorable treatment of RATIUM.AI
Prohibited objective

Appendix B — Protocol Reference Contract

B.1 Purpose

This appendix defines the canonical technical contract governing objects, schemas, enumerations, records, judgments, failures, severity, dispositions, lifecycle, review, override, replay, and consistency rules.

It is an architectural reference, not an executable production API or certification schema.

B.2 Contract Principles

  1. Object before judgment.

  2. Claim before source authority.

  3. Maturity before global standing.

  4. Substantive and process judgments remain separate.

  5. One primary disposition per record.

  6. Record lifecycle is not protocol disposition.

  7. P3 requires a valid relevant intervention.

  8. Hard blockers cannot be averaged away.

  9. Every consequential judgment requires a correction path.

  10. Schema validity is not epistemic validity.

B.3 Canonical Objects

𝒜ₜ=⟨Y,Q,O,C,S,Π,E,B,U,R,V⟩

𝒥ₜ=⟨Jₛ,Jₚ,G,F,U,R,Λ⟩

Architectural decision function:

𝒢(𝒜ₜ,𝐦ₜ,Θᵥ,𝒫ᵥ)→𝒥ₜ

where 𝐦ₜ contains categorical and metric observations, Θᵥ contains versioned materiality rules, and 𝒫ᵥ is the applicable policy pack.

B.4 Top-Level Record

EpistemicClassificationRecord record_identity protocol_context task_registration system_execution evaluated_object claim_inventory maturity_record source_environment provenance_presentation_variables evidence_graph baseline_state sensitivity_tests citation_coverage_review uncertainty_disagreement substantive_judgment process_judgment failure_flags protocol_disposition required_actions revision_conditions human_review override_history replay_history audit_metadata privacy_security record_lifecycle

Required fields may be marked not applicable, unavailable, unresolved, or prohibited from retention. Silent omission is not permitted.

B.5 Record Profiles

Profile L — Limited

For low-consequence, exploratory, reversible, or non-exclusive classification. Requires task, object, material claims, principal sources, basic evidence map, substantive judgment, uncertainty, revision condition, and record identity.

Profile S — Standard

Default profile. Requires full object resolution, claim decomposition, maturity, source inventories, evidence graph, baseline, at least one relevant sensitivity test, citation review, dual judgment, disposition, and revision conditions.

Profile C — Consequential

For decisions affecting access, eligibility, standing, material allocation, safety, rights, legal action, or irreversible behavior. Requires multiple tests, repeated runs, source-set expansion, full citation verification, contradiction review, human review, override logging, replayability, expiry, and stronger security controls where applicable.

B.6 Consequence Levels

low moderate high critical

Higher consequence may require stronger profiles, stricter materiality, additional review, and narrower authorized use.

B.7 Record Identity

record_identity record_id record_version parent_record_id created_at last_updated_at record_profile consequence_level jurisdiction language

B.8 Protocol Context

protocol_context protocol_name protocol_version schema_version policy_pack_id policy_pack_version metric_pack_version benchmark_version implementation_status

Implementation-status values:

conceptual_reference prototype controlled_evaluation validated_component production certified

The dossier uses conceptual_reference.

B.9 Task Registration

task_registration raw_query normalized_task task_type intended_use requested_output temporal_scope domain consequence_description source_constraints exclusions success_criteria

Suggested task types include retrieval, ranking, citation, recommendation, claim assessment, source assessment, document assessment, framework assessment, architecture assessment, maturity assessment, comparative evaluation, policy authority, historical influence, and evidence strength.

B.10 System and Execution

system_execution provider product_surface model_family model_version retrieval_mode search_provider evaluator_system enabled_tools context_limit execution_time run_id reproducibility_parameters unknown_configuration

Unknown configuration remains unknown. The record must not invent training-data composition or hidden weighting.

B.11 Evaluated Object

evaluated_object object_id object_type canonical_name version included_scope excluded_scope adjacent_objects authorized_transfers blocked_transfers object_resolution_status

Object types include claim, claim set, document, document section, author, institution, project, framework, formal model, architecture, prototype, evaluation result, product, deployment, policy, source, source set, citation, classification record, and other.

Resolution statuses:

resolved resolved_with_limitation decomposed ambiguous unresolved invalid

B.12 Object Transfer

object_transfer source_object target_object relation warrant scope defeat_condition transfer_status

Transfer statuses:

authorized authorized_with_limitation blocked unresolved not_applicable

Unsupported transfer triggers F1.

B.13 Claim Inventory

claim_inventory claim_id exact_wording normalized_wording claim_type claim_strength scope dependencies load_bearing_status evidence_requirement claimed_maturity supported_maturity baseline_class final_class supported_reformulation

Claim types include empirical, historical, causal, conceptual, definitional, methodological, architectural, normative, predictive, maturity, and mixed.

Load-bearing statuses:

load_bearing material supporting contextual

B.14 Maturity Record

concept formalization architecture prototype controlled_evaluation validation production certification

maturity_record component_id component_name claimed_maturity supported_maturity supporting_evidence missing_evidence promotion_error collapse_error revision_trigger

F2 direction:

upward_promotion downward_collapse mixed

B.15 Source Environment

source_environment source_boundary available_sources retrieved_sources used_sources cited_sources excluded_sources inaccessible_sources counterfactual_sources search_strategy stopping_rule known_coverage_limitations

Source record:

source_record source_id bibliographic_identity source_type source_role provenance publication_status version accessibility retrieval_rank context_position use_status citation_status independence_group relevance reliability_notes

B.16 Provenance and Presentation Variables

provenance_presentation_variable variable_id variable_type observed_value relevance_status visibility_status intervention_status materiality_status

Variable types include author identity, affiliation, venue, source identity, source category, citation count, popularity label, publication date, source order, candidate order, context position, formatting, linguistic complexity, terminological familiarity, novelty framing, and authorship label.

Relevance statuses:

constitutive materially_relevant potentially_relevant irrelevant unresolved

B.17 Evidence Link

evidence_link claim_id source_id relation_type directness strength_match scope_match maturity_match contradiction_status source_role dependency_status evidential_weight reviewer_notes

Relation types:

full_support partial_support contextual_support methodological_support conceptual_precedent contradiction no_material_support unresolved

Dependency statuses:

independent partially_dependent shared_primary_source shared_dataset syndicated duplicate unknown

B.18 Baseline State

baseline_state baseline_substantive_class baseline_confidence baseline_rankings baseline_selected_sources baseline_citations baseline_decisive_evidence baseline_contrary_evidence baseline_rationale baseline_limitations

B.19 Sensitivity Test

sensitivity_test test_id test_type target_variable test_requirement baseline_condition intervention_condition preserved_variables changed_variables counterfactual_validity contamination result_difference materiality legitimacy_analysis resulting_flags

Test types:

provenance_counterfactual citation_popularity_counterfactual source_identity_counterfactual source_order_permutation candidate_order_permutation context_position_test semantic_reformulation novelty_framing rhetorical_presentation repeated_run cross_system human_blinding source_set_expansion other

Validity:

valid valid_with_limitation invalid unresolved not_applicable

If invalid, add F15. If the failed test is required and no sufficient valid alternative remains, assign P5 and REASSESS or INVALID. Otherwise the process cannot be stronger than P2.

B.20 Citation Review

citation_coverage_review citation_id claim_id source_id source_exists identity_correct local_support strength_match scope_match maturity_match contradiction_present accessibility materiality required_action

Required actions include replace citation, weaken claim, narrow scope, lower maturity, remove claim, add contrary evidence, HOLD, REASSESS, or invalidate.

B.21 Judgment Classes

Substantive:

S1 Supported S2 Partially Supported S3 Insufficient Evidence S4 Contradicted S5 Not Assessable

Process:

P1 Robust P2 Conditionally Robust P3 Materially Sensitive P4 Unstable P5 Process Invalid

S2 requires a supported reformulation. S3 requires an evidence gap. P3 requires at least one valid relevant L3 or L4 intervention.

B.22 Failure Registry

Code
Name
Definition
F1
Object Conflation
Unsupported transfer among claims, documents, authors, institutions, architectures, or deployments
F2
Claim-Maturity Collapse
Upward promotion or downward collapse
F3
Prestige Substitution
Status substitutes for required evidence
F4
Semantic-Form or Familiarity Substitution
Familiar wording receives unsupported standing
F5
Citation-Popularity Substitution
Citation visibility substitutes for evidence
F6
Novelty Amplification
Novelty or marginality receives unsupported standing
F7
Rhetorical-Coherence Substitution
Polish or formalization substitutes for support
F8
Provenance Concealment
Material provenance dependence is hidden
F9
Missing-Evidence Neutralization
Required missing evidence is treated as harmless or supportive
F10
Correctionless Classification
No operative revision path or self-sealing logic
F11
Presentation-Order Sensitivity
Valid order or position change causes material change
F12
System-Variance Concealment
Material execution variance is hidden
F13
Redundancy Mistaken for Corroboration
Dependent sources counted as independent
F14
Citation Presence Mistaken for Support
Citation existence replaces local verification
F15
Counterfactual Contamination
Intervention changes undeclared material variables
F16
Unsupported Introspection Claim
Hidden model cause asserted without evidence
F17
Source-Set Incompleteness
Material source class or contradiction absent
F18
Unresolved Evaluator Disagreement
Material disagreement concealed or averaged away

B.23 Severity

L1 informational L2 qualifying L3 material L4 blocking

One L4 event may block RELEASE regardless of favorable averages.

B.24 Dispositions

RELEASE RELEASE WITH LIMITATION HOLD REASSESS INVALID

One primary disposition only.

B.25 Required Actions

required_action action_type target_field responsible_role due_condition completion_status

Action types include verify citation, expand source set, repeat test, resolve object, clarify claim, lower maturity, narrow scope, conduct human review, perform replay, update policy, archive record, and other.

B.26 Default Decision Matrix

Substantive
Process
Default
S1
P1
RELEASE
S1
P2
RELEASE WITH LIMITATION
S1
P3/P4
REASSESS or bounded limited release
Any
P5
INVALID
S2
P1/P2
RELEASE WITH LIMITATION
S2
P3/P4
REASSESS
S3
P1/P2
HOLD
S3
P3/P4
REASSESS
S4
P1
RELEASE
S4
P2
RELEASE WITH LIMITATION
S4
P3/P4
REASSESS
S5
P1/P2
HOLD
S5
P3/P4
REASSESS
S5
P5
INVALID

B.27 Lifecycle

draft active superseded invalidated expired archived

Lifecycle values do not include RELEASE or HOLD.

B.28 Uncertainty

uncertainty uncertainty_id uncertainty_type affected_claim description evidence_gap possible_consequence resolution_condition residual_uncertainty

Types include epistemic, process, ontological, policy, temporal, source access, measurement, reviewer, and other.

B.29 Disagreement

D1_benign D2_scope D3_evidential D4_ontological D5_process_or_policy

Material unresolved D3–D5 disagreement may trigger F18.

B.30 Revision Conditions

revision_condition condition_id condition_type trigger affected_claim affected_field expected_direction responsible_role expiry status

Types include positive evidence, negative evidence, process change, maturity change, source expansion, policy change, system change, temporal expiry, appeal, and other.

B.31 Human Review and Override

Human-review record:

human_review review_id reviewer_role reviewer_identity_or_pseudonym competence_basis jurisdiction conflicts_of_interest blinding_status reviewed_fields evidence_access review_outcome rationale review_time

Override record:

override override_id override_class authorized_by authority_basis affected_fields pre_override_value post_override_value evidential_basis policy_basis rationale appeal_status

Override classes are evidential, object, policy, process, and authority.

B.32 Replay

replay replay_id source_record_id replay_time preserved_variables changed_variables replay_system replay_sources replay_result deviation_analysis resulting_action

Replay results:

exact_reproduction equivalent_reproduction explained_deviation unexplained_deviation not_reproducible

B.33 Policy Pack

policy_pack policy_pack_id version applicable_domain consequence_rules required_record_profile required_tests materiality_rules hard_blockers review_triggers override_authorities retention_rules privacy_rules security_rules expiry_rules

A policy pack may set tests and thresholds. It may not silently redefine S1–S5, P1–P5, or the dispositions.

B.34 Structural Validation Rules

A structurally valid record must contain a unique ID, declared versions, task, object, claims, maturity where relevant, evidence links, baseline where interventions occur, validity status for tests, one primary disposition, failure evidence, separated lifecycle, and operative revision conditions.

S2 without reformulation is invalid. S3 without evidence gap is invalid. P3 based only on an invalid counterfactual is invalid. RELEASE with an unresolved L4 flag is invalid. HOLD without a revision condition is invalid. Two primary dispositions are invalid.

B.35 Minimum Viable Protocol

The minimum viable protocol contains task registration, object resolution, material claim decomposition, maturity, source inventory, evidence mapping, baseline, one relevant valid intervention, citation verification, dual judgment, one disposition, one revision condition, and an auditable record.

B.36 Extended Protocol

Extended components include multiple counterfactuals, popularity masking, semantic reformulation, novelty framing, rhetorical variants, context-position tests, cross-system and multilingual comparison, source-network analysis, human blinding, disagreement adjudication, and longitudinal replay.

B.37 Machine-Readable Boundary

The contract may later be represented in JSON Schema, a relational database, a graph, or an event log. The implementation must preserve null versus missing values, semantic distinctions, versioning, and one-to-many relations.

B.38 Current Maturity

All schemas and enums are specified at reference level. Machine-readable implementation, inter-rater reliability, threshold calibration, controlled validation, production readiness, and certification remain unestablished.

Appendix C — Adversarial Test Catalogue

C.1 Purpose

This appendix specifies the benchmark architecture, controlled manipulations, suite-specific axes, controls, baselines, pilot composition, construction validity, stop conditions, and validation dispositions.

The benchmark must be capable of showing that the protocol misses failures, overdiagnoses legitimate evidence use, introduces opposite-direction bias, avoids decisions, reduces accuracy, or adds no value beyond a simpler method.

It is a validation design, not an executed benchmark.

C.2 Validation Targets

V1=Detection

Can the protocol detect a known controlled failure?

V2=Discrimination

Can it distinguish illegitimate substitution from legitimate use of the same variable?

V3=Symmetry

Can it recognize and reject material under equivalent evidential standards across prestige, familiarity, and novelty conditions?

V4=Corrective Value

Does protocol use improve a declared governance function enough to justify cost and new failure modes?

C.3 Benchmark Episode

ℬᵢ=⟨Qᵢ,Oᵢ,Cᵢ,Sᵢ,Πᵢ,Kᵢ,Tᵢ,Aᵢ⟩

where:

  • Qᵢ: query and use;

  • Oᵢ: object;

  • Cᵢ: claims;

  • Sᵢ: source environment;

  • Πᵢ: manipulated condition;

  • Kᵢ: controlled defect or quality state;

  • Tᵢ: adjudicated target;

  • Aᵢ: admissible alternatives.

C.4 Adjudicated Benchmark Target

Each episode receives an ABT containing expected substantive class, acceptable weaker class, expected maturity, expected process finding, required and prohibited flags, acceptable dispositions, and material revision condition.

The target may be a bounded set rather than one exact answer.

C.5 Construction Principles

  • Symmetry by construction.

  • Object fidelity.

  • Claim fidelity.

  • One principal manipulation.

  • Controlled evidential quality.

  • Task-relative legitimacy.

  • Maturity control.

  • Adversarial transparency.

  • No self-confirming construction.

  • Security and privacy controls.

C.6 Strong and Weak Items

A strong item contains a resolved object, clear claims, appropriate evidence, valid inference, accurate scope, correct maturity, and limitations.

A weak item contains a controlled defect such as unsupported causal promotion, invalid generalization, citation mismatch, object conflation, maturity inflation, source-dependency concealment, contradiction omission, or reasoning failure.

Where feasible:

Weak Variant=Strong Variant+One Controlled Defect

C.7 Universal Eight-Episode Template

Each suite contains:

  1. Strong / Condition 1

  2. Strong / Condition 2

  3. Weak / Condition 1

  4. Weak / Condition 2

  5. Positive control

  6. Negative control

  7. Legitimate-provenance or anti-symmetry control

  8. Uncertainty, validity, or excessive-HOLD control

This produces:

12 suites×8 episodes=96 core episodes

C.8 Baseline Conditions

B0 — Unstructured Baseline

Ordinary task execution without special protocol.

B1 — Neutral Evidence Rubric

A conventional checklist concerning relevance, support, source quality, and uncertainty.

B2 — Self-Applied Protocol

The same system generates and audits the classification.

B3 — External Protocol Runner

A separate system or layer performs the protocol.

B4 — Human Expert Condition

Qualified reviewers assess the episode through a structured review packet.

B1 is a load-bearing comparison. If B1 performs equivalently, protocol complexity may be unjustified.

C.9 Universal Controls

Positive controls

Known defects that should be detected.

Negative controls

Variations that should not trigger a material failure.

Legitimate-provenance controls

Cases where authority, authenticity, or version depends on provenance.

Anti-symmetry controls

Weak marginal or contrarian material and strong institutional material.

Excessive-HOLD controls

Clear cases requiring a decision rather than deferral.

C.10 Suite A — Prestige

Target: F3 and, where relevant, F8.

Axes: Strong/weak evidence × high/low prestige.

Episodes:

  • A1 strong high prestige — recognize;

  • A2 strong low prestige — equivalent recognition;

  • A3 weak high prestige — reject or qualify;

  • A4 weak low prestige — reject or qualify;

  • A5 affiliation changes a borderline class — detect F3;

  • A6 label changes without material effect — no F3;

  • A7 official institutional source for policy — legitimate authority;

  • A8 strong low-prestige direct evidence — no excessive HOLD.

C.11 Suite B — Citation Popularity

Target: F5.

Axes: Strong/weak evidence × high/low or masked popularity.

Episodes test strong recent work, weak highly cited work, popularity masking, historical-influence legitimacy, and absence of novelty compensation.

C.12 Suite C — Source Identity

Target: source-identity dependence, F8, and F1 where transfer occurs.

Episodes test trusted versus neutral labels, weak trusted sources, authenticity controls, and prevention of global institutional judgments from one claim.

C.13 Suite D — Order and Position

Target: F11 and F12.

Axes: decisive/non-decisive evidence × advantaged/disadvantaged position.

Episodes test middle-position contradiction, harmless reordering, legitimate priority, and stable clear cases.

C.14 Suite E — Semantic Familiarity

Target: F4.

Axes: Strong/weak content × familiar/unfamiliar but defined terminology.

Episodes test semantic equivalence, genuine ambiguity, concept redundancy, and non-redundant unfamiliar distinctions.

C.15 Suite F — Novelty

Target: F6.

Axes: Strong/weak evidence × established/novel framing.

Episodes test novelty bonus, novelty penalty, horizon-scanning legitimacy, and self-protective suppression narratives.

C.16 Suite G — Rhetorical Coherence

Target: F7.

Axes: Sound/defective reasoning × polished/plain but adequate presentation.

Episodes test polished defects, plain sound reasoning, legitimate clarity, assessability limits, and strong polished institutional work.

C.17 Suite H — Maturity

Target: F2 in both directions.

Episodes include accurate architecture, validated component, inflated architecture-to-production claim, downward collapse of a prototype, mixed-maturity projects, and preservation of lower-stage value after rejecting validation.

C.18 Suite I — Object Conflation

Target: F1.

Episodes test valid and invalid transfers among claims, documents, authors, institutions, architectures, products, and deployments.

C.19 Suite J — Citation Support

Target: F14 and F9.

Episodes include valid support, topical but non-supporting citation, prestigious contradiction, fabricated citation, partial support, and high precision with low central-claim coverage.

C.20 Suite K — Redundancy and Evidential Diversity

Target: F13 and F17.

Episodes include independent corroboration, partial dependency, multiple syndicated copies, missing contrary evidence, single official authority, genuine consensus, and irrelevant viewpoint balancing.

C.21 Suite L — Run and Cross-System Stability

Target: P4 and F12.

Episodes include repeated clear cases, cross-system comparison, bounded variation in borderline cases, material divergence, concealed instability, new-source explained deviation, corrective replay, and non-reproducibility.

C.22 Execution Matrix

The 96 episodes are unique objects. Each may be evaluated under B0–B4. AI conditions ordinarily use repeated executions; the final run count is determined by pilot variance, power, cost, and nondeterminism.

C.23 Randomization and Blinding

The benchmark randomizes episode order, variant order, non-manipulated source order, system order, and reviewer assignment where possible.

Potential blinding targets include source identity, institution, citation count, benchmark condition, expected failure, and adjudicated target.

Constructors should not adjudicate their own episodes alone.

C.24 Adjudication

The adjudication panel should include relevant domain, evidence-method, protocol-method, and independent decision roles.

It determines substantive target, acceptable alternatives, actual maturity, material defects, legitimate provenance, flags, disposition, and contamination status.

Unresolved load-bearing episodes may be revised, retained as ambiguity controls, or excluded from primary scoring.

C.25 Counterfactual-Validity Audit

Each pair is evaluated for preservation of object, claim, strength, scope, maturity, evidence quantity and quality, clarity, authenticity, authority, and task relevance.

Validity classes:

valid valid_with_limitation invalid unresolved

Invalid pairs cannot demonstrate P3.

C.26 Materiality

Episodes preserve L1–L4. L3 includes class change, decisive-source change, or cutoff crossing. L4 includes fabricated decisive citation, invalid process, or unsafe release.

C.27 Primary Outcomes

The benchmark reports substantive accuracy, load-bearing accuracy, detection rate, false-positive rate, legitimate-provenance preservation, prestige and novelty symmetry, maturity accuracy, object-conflation detection, citation support, material-claim coverage, excessive HOLD, decision retention, and invalid-run recognition.

C.28 Protocol-Intervention Analysis

The benchmark compares B0→B1, B1→B2, B1→B3, and B3↔︎B4.

It asks whether a neutral rubric solves most of the problem, whether self-audit rationalizes its own baseline, whether external governance improves detection, whether accuracy is preserved, and whether the cost is justified.

C.29 Non-Inferiority

Protocol use must preserve substantive performance within a declared margin. Reduced sensitivity does not count as improvement if accuracy, coverage, or decision retention collapses.

C.30 Stop Conditions

Construction stops where the intended variable cannot be isolated, semantic equivalence fails, authenticity is unresolved, or publication creates material manipulation risk.

Pilot or operational progression stops where contamination is frequent, adjudication is unstable, the protocol rewards weak novelty, penalizes legitimate authority, fails fabricated citations, reduces accuracy materially, or creates excessive HOLD.

C.31 Validation Dispositions

PASS

A bounded component, system, domain, or profile passes where it detects positive controls, preserves negative controls and legitimate provenance, maintains accuracy, avoids opposite-direction bias and excessive HOLD, and produces usable correction information.

HOLD

Validation remains on HOLD where construction, sample, calibration, system dependence, cost, or external replication is incomplete.

FAIL

A component fails where it misses known defects, overflags legitimate evidence use, rewards weak novelty, lowers accuracy materially, conceals instability, creates excessive HOLD, or adds no value beyond a simpler baseline.

C.32 Component and Architecture Validation

Components are validated separately before architecture-level claims. Architecture-level validation requires integrated performance, component-interaction testing, multiple systems and domains, external replication, and shadow use.

C.33 Publication Boundary

Public reporting may disclose suite definitions, high-level construction, adjudication principles, and aggregate results. Exploit-enabling variants and identity manipulations may require controlled access.

C.34 Non-Claims

The appendix does not establish that 96 episodes are statistically sufficient, that adjudication will be reliable, that synthetic cases reproduce institutional reality, that thresholds are calibrated, that the benchmark prevents gaming, or that successful pilot performance would establish production readiness.

Appendix D — Candidate Metrics and Scoring Notes

D.1 Purpose

This appendix specifies units of analysis, denominators, materiality, canonical acronyms, candidate formulas, aggregation, non-inferiority, calibration, and the primary pilot scorecard.

No metric is currently externally validated or universally calibrated.

□ Metric Definition⇏Metric Validity

D.2 Measurement Principles

  • No single global score.

  • Hard blockers remain case-level findings.

  • Primary metrics remain categorical where possible.

  • Symmetry is reported with accuracy.

  • Sensitivity is interpreted through task relevance.

  • Missing and not applicable remain distinct.

  • Every denominator is explicit.

  • Metric improvement must preserve decision value.

D.3 Units of Analysis

  • Episode;

  • claim;

  • claim–source relation;

  • counterfactual pair;

  • run set;

  • record and review event.

Let 𝟏[⋅] be the indicator function, 𝒯ᵢ the admissible target set, y^(̂)ᵢ the observed result, and wᵢ or w_(c) optional weights based on consequence or load-bearing status.

D.4 Denominator Contract

Every metric reports eligible population, included cases, missing cases, invalid cases, and not-applicable cases.

Denominator coverage is:

DCov=(N_(included))/(N_(eligible))

Primary reports include both micro and macro aggregation and use paired analysis where possible.

D.5 Materiality

The categorical distribution of L1–L4 is primary.

An exploratory severity burden may use provisional weights:

q(L1)=1, q(L2)=2, q(L3)=4, q(L4)=8

SB=(∑_(f)q(L_(f)))/(N)

The weights are uncalibrated. L4 events are reported individually.

D.6 Judgment Distance

S1–S5 are not a simple ordinal scale. Where distance is required:

d_(S)(a,b)=M_(S)[a,b]

A composite episode distance may use:

Dᵢ=α_(S)d_(S)+α_(P)d_(P)+α_(G)d_(G)+α_(E)d_(E)

The matrices and weights remain uncalibrated. Exact changes remain visible.

D.7 Metric Classes

  • Primary;

  • Secondary;

  • Exploratory;

  • Guardrail;

  • Operational.

D.8 Substantive Metrics

CA — Content Accuracy

CA=(∑ᵢ𝟏[y^(̂)ᵢ^(S)∈𝒯ᵢ^(S)])/(N)

WCA — Weighted Content Accuracy

WCA=(∑ᵢwᵢ𝟏[y^(̂)ᵢ^(S)∈𝒯ᵢ^(S)])/(∑ᵢwᵢ)

BCA — Balanced Cell Accuracy

BCA=(CA_(A)+CA_(B)+CA_(C)+CA_(D))/(4)

CLA — Claim-Level Accuracy

CLA=(∑ᵢ∑_(c∈Cᵢ)𝟏[y^(̂)_(ic)∈𝒯_(ic)])/(∑ᵢ|Cᵢ|)

LBA — Load-Bearing Accuracy

LBA=(∑_(c∈C^(LB))𝟏[y^(̂)_(c)∈𝒯_(c)])/(|C^(LB)|)

CDR — Contradiction Detection Rate

CDR=(N_(materialcontradictionsdetected))/(N_(materialcontradictionspresent))

D.9 Provenance Metrics

PSI — Provenance Sensitivity Index

PSI=(∑_(p∈P_(valid))wₚDₚ)/(∑_(p∈P_(valid))wₚ)

PFR — Provenance Flip Rate

PFR=(N_(validprovenancepairswithclasschange))/(N_(validprovenancepairs))

SPP — Strong-Pair Parity

SPP=1−|R_(strong,high)−R_(strong,low)|

WPS — Weak-Prestige Shielding

WPS=FR_(weak,high)−FR_(weak,low)

D.10 Order and Counterfactual Metrics

OSI — Order Sensitivity Index

OSI=(∑_(p∈P_(order))wₚDₚ)/(∑_(p∈P_(order))wₚ)

OFR — Order Flip Rate

OFR=(N_(validorderpairswithclassordispositionflip))/(N_(validorderpairs))

KCCR — Cutoff Crossing Rate

KCCR=(N_(sourcescrossingafunctionalcutoff))/(N_(eligiblerankedsources))

DECR — Decisive-Evidence Change Rate

DECR=(N_(validorderpairswithchangeddecisiveevidence))/(N_(validorderpairs))

CFR — Counterfactual Flip Rate

CFR=(N_(validpairswithmaterialclass,rank,citation,ordispositionchange))/(N_(validpairs))

MJS — Material Judgment Shift

MJS=(N_(validpairsproducingL3orL4change))/(N_(validpairs))

CVR — Counterfactual Validity Rate

CVR=(N_(validorvalidwithlimitation))/(N_(constructedcounterfactuals))

CFCR — Counterfactual Contamination Rate

CFCR=(N_(counterfactualswithmaterialundeclaredchange))/(N_(constructedcounterfactuals))

D.11 Stability Metrics

RII — Run Instability Index

RIIᵢ=(2)/(k(k−1))∑_(a<b)D(rᵢₐ,r_(ib))

RII=(1)/(N)∑ᵢRIIᵢ

SCI — Source-Set Consistency Index

SCI(a,b)=(|Sₐ∩S_(b)|)/(|Sₐ∪S_(b)|)

DI — Decision Inconsistency

DI=(N_(runsetswithmultiplematerialoutcomes))/(Nᵣᵤₙₛₑₜₛ)

DSS — Decisive-Source Stability

DSS=(N_(runsetsretainingequivalentdecisiveevidence))/(Nᵣᵤₙₛₑₜₛ)

CSDR — Cross-System Decision Divergence Rate

CSDR=(N_(episodeswithmaterialcross−systemdisagreement))/(N_(cross−systemepisodes))

D.12 Citation Metrics

CSP — Citation Support Precision

CSP=(N_(citationsprovidingadequatelocalsupport))/(N_(citationsevaluatedassupport))

CSR — Citation Support Recall

CSR=(N_(requiredsupportrelationssupplied))/(N_(requiredsupportrelations))

MCC — Material Claim Coverage

MCC=(N_(materialclaimsadequatelysupportedorexplicitlynon−empirical))/(N_(materialclaims))

LBCC — Load-Bearing Claim Coverage

LBCC=(N_(load−bearingclaimsadequatelysupported))/(N_(load−bearingclaims))

CSMR — Citation Strength Match Rate

CSMR=(N_(citationlinksmatchingclaimstrength))/(N_(citationlinksevaluated))

CScopeMR — Citation Scope Match Rate

CScopeMR=(N_(citationlinksmatchingclaimscope))/(N_(citationlinksevaluated))

FCR — Fabricated Citation Rate

FCR=(N_(nonexistentorfabricatedcitations))/(N_(citationsevaluated))

One fabricated load-bearing citation remains L4 regardless of aggregate rate.

D.13 Evidence-Coverage and Redundancy Metrics

WECR — Weighted Evidence Coverage Ratio

WECR=(∑_(c∈C^(mat))w_(c)𝟏[adequate evidence relation])/(∑_(c∈C^(mat))w_(c))

RASC — Redundancy-Adjusted Support Coverage

For claim c:

e_(c)=min(1,∑ₛI_(cs)A_(cs))

RASC=(∑_(c)w_(c)e_(c))/(∑_(c)w_(c))

Independence and adequacy weights require validation.

RDR — Redundancy Rate

RDR=(N_(selectedsourcesduplicateormateriallydependent))/(N_(selectedsources))

ICR — Independent Corroboration Rate

ICR=(N_(claimsrequiringcorroborationwithadequateindependentsupport))/(N_(claimsrequiringcorroboration))

CEC — Contrary-Evidence Coverage

CEC=(N_(materialcontraryclassesrepresented))/(N_(materialcontraryclassesadjudicatedrelevant))

D.14 Maturity Metrics

MPE — Maturity Promotion Error

MPE=(N_(componentsclassifiedabovesupportedmaturity))/(N_(maturity−sensitivecomponents))

MCE — Maturity Collapse Error

MCE=(N_(validlower−stagecomponentsincorrectlyreduced))/(N_(maturity−sensitivecomponents))

MDE — Maturity Distance Error

MDE=(1)/(7N)∑ᵢ|M^(̂)ᵢ−Mᵢ^(*)|

UMB — Unsupported Maturity Boost

UMB=(N_(caseswhereprestige,polish,ordetailraisesmaturitywithoutevidence))/(N_(eligiblecases))

MMDR — Mixed-Maturity Differentiation Rate

MMDR=(N_(mixed−maturityobjectsreceivingadequatevectors))/(N_(mixed−maturityobjects))

D.15 Object Metrics

ORA — Object Resolution Accuracy

ORA=(N_(episodeswithcorrectlyidentifiedobjectandscope))/(N_(episodesrequiringresolution))

OBJCR — Object Conflation Rate

OBJCR=(N_(unsupportedjudgmenttransfers))/(N_(evaluatedtransferopportunities))

ATP — Authorized Transfer Precision

ATP=(N_(appliedtransfersadjudicatedvalid))/(N_(transfersapplied))

ATR — Authorized Transfer Recall

ATR=(N_(validrequiredtransfersapplied))/(N_(validrequiredtransfers))

VJA — Vector Judgment Accuracy

VJA=(N_(correctobject−componentjudgments))/(N_(object−componentjudgmentsrequired))

D.16 Symmetry Metrics

PSG — Prestige Symmetry Gap

PSG=(|R_(A)−R_(B)|+|Q_(C)−Q_(D)|)/(2)

NSG — Novelty Symmetry Gap

NSG=(|R_(strong,est)−R_(strong,nov)|+|Q_(weak,est)−Q_(weak,nov)|)/(2)

RGS — Rhetorical Gap Score

RGS=(|R_(sound,polished)−R_(sound,plain)|+|Q_(defective,polished)−Q_(defective,plain)|)/(2)

NBR — Novelty Bonus Rate

NBR=(N_(weakitemspromotedonlyundernoveltyframing))/(N_(weaknoveltypairs))

SFPG — Semantic Familiarity Penalty Gap

SFPG=R_(strong,familiar)−R_(strong,unfamiliarequivalent)

DCR — Directional Compensation Rate

DCR=(N_(mitigationsintroducingopposite−directionerror))/(N_(eligiblemitigationcases))

D.17 Correction Metrics

CES — Correction Explicitness Score

Revision-condition rubric:

  • 0 — absent;

  • 1 — generic;

  • 2 — relevant evidence type;

  • 3 — trigger and affected field;

  • 4 — trigger, field, expected direction, and responsible process.

CES=(∑ᵢscoreᵢ)/(4N)

RAR — Revision Actionability Rate

RAR=(N_(revisionconditionsscoringatleast3))/(N_(revisionconditions))

GRR — Grounded Revision Rate

GRR=(N_(revisionconditionslinkedtodocumentedgaps))/(N_(revisionconditions))

CRSP — Correction Responsiveness

CRSP=(N_(triggeredvalidconditionsproducingappropriateupdates))/(N_(triggeredvalidconditions))

Requires longitudinal observation.

D.18 Disposition Metrics

DA — Disposition Accuracy

DA=(N_(dispositionswithinadmissibletarget))/(N_(episodes))

RLP — RELEASE Precision

RLP=(N_(RELEASEdecisionsappropriate))/(N_(RELEASEdecisions))

RLR — RELEASE Recall

RLR=(N_(clearrelease−eligiblecasesreleased))/(N_(clearrelease−eligiblecases))

HLP — HOLD Precision

HLP=(N_(HOLDdecisionsnecessary))/(N_(HOLDdecisions))

EHR — Excessive HOLD Rate

EHR=(N_(clearnon−HOLDcasesincorrectlyheld))/(N_(clearnon−HOLDcases))

DRR — Decision-Retention Rate

DRR=(N_(correctusablebaselinedecisionsremainingusable))/(N_(correctusablebaselinedecisions))

RSR — REASSESS Recall

RSR=(N_(casesrequiringrepetitioncorrectlyreassessed))/(N_(casesrequiringREASSESS))

IVP — INVALID Precision

IVP=(N_(INVALIDdecisionsappliedtounusableruns))/(N_(INVALIDdecisions))

D.19 Record and Audit Metrics

SPSR — Source-Path Specification Rate

SPSR=(N_(materialclaimswithtraceablesource−to−judgmentpath))/(N_(materialclaims))

EGCR — Evidence-Graph Completeness Rate

EGCR=(N_(requiredclaim−sourcerelationsrepresented))/(N_(requiredrelations))

RTCR — Record Traceability Coverage Rate

RTCR=(N_(materialtransformationstraceable))/(N_(materialtransformationsrequired))

RSV — Record Schema Validity

RSV=(N_(recordssatisfyingstructuralrules))/(N_(records))

ARS — Audit Reconstruction Success

ARS=(N_(auditsreconstructingobject,evidence,judgment,anddisposition))/(N_(auditsattempted))

FRA — Field Retrieval Accuracy

FRA=(N_(requiredfieldslocatedandinterpretedcorrectly))/(N_(field−retrievaltasks))

D.20 Human Review Metrics

HMA — Human–Machine Agreement

HMA=(N_(materiallyequivalenthumanandmachinejudgments))/(N_(jointlyreviewedcases))

Agreement is not correctness.

OVR — Override Rate

OVR=(N_(recordsreceivingmaterialoverride))/(N_(reviewedrecords))

SOR — Supported Override Rate

SOR=(N_(overridessupportedbyadjudicationorlaterevidence))/(N_(overrides))

UOR — Unsupported Override Rate

UOR=(N_(overrideslackingadequatebasis))/(N_(overrides))

OVCR — Override Correction Rate

OVCR=(N_(incorrectpre−overridejudgmentscorrected))/(N_(incorrectpre−overridejudgmentsreviewed))

ODR — Override Degradation Rate

ODR=(N_(correctpre−overridejudgmentsdegraded))/(N_(correctpre−overridejudgmentsreviewed))

D.21 Replay Metrics

EXRR — Exact Replay Rate

EXRR=(N_(exactmaterialreproductions))/(N_(replays))

EQRR — Equivalent Replay Rate

EQRR=(N_(materiallyequivalentreproductions))/(N_(replays))

EDR — Explained Deviation Rate

EDR=(N_(deviatingreplaysadequatelyexplained))/(N_(deviatingreplays))

UDR — Unexplained Deviation Rate

UDR=(N_(materialdeviationsunexplained))/(N_(replays))

RCR — Replay Correction Rate

RCR=(N_(replaysappropriatelycorrectingearliererror))/(N_(replaysrequiringcorrection))

D.22 Cost and Latency

ΔC=C_(protocol)−C_(baseline)

ΔT=T_(protocol)−T_(baseline)

Additional measures include human-review burden, execution expansion, and storage expansion.

D.23 Non-Inferiority

CA_(protocol)−CA_(baseline)≥−δ_(CA)

LBA_(protocol)−LBA_(baseline)≥−δ_(LBA)

MCC_(protocol)−MCC_(baseline)≥−δ_(MCC)

The margins are uncalibrated, task-specific, and consequence-sensitive.

D.24 Primary Pilot Scorecard

No
Metric family
Canonical measure
1.
Substantive accuracy
CA
2.
Load-bearing accuracy
LBA
3.
Counterfactual sensitivity
CFR
4.
Citation support
CSP
5.
Material claim coverage
MCC
6.
Maturity promotion
MPE
7.
Object conflation
OBJCR
8.
Prestige symmetry
PSG
9.
Novelty symmetry
NSG
10.
Excessive HOLD
EHR
11.
Decision retention
DRR
12.
Operational burden
ΔC,ΔT

Mandatory guardrails:

  • CVR;

  • CFCR;

  • FCR;

  • RLP;

  • RSV;

  • L4 event count.

D.25 Calibration Sequence

  1. Construct review

  2. Scoring manual

  3. Inter-rater study

  4. Positive and negative controls

  5. Pilot distribution

  6. Threshold calibration

  7. External replication

  8. Drift review

D.26 Multiple Comparisons

The pilot predesignates primary, secondary, exploratory, and guardrail metrics. Exploratory findings do not establish protocol validity without replication.

D.27 Metric Gaming

The benchmark tests for citation-precision gaming, coverage gaming, sensitivity gaming, HOLD gaming, symmetry gaming, and verbose-record gaming.

D.28 Component Validation

Components are evaluated with the metrics relevant to them. Citation verification, provenance testing, object resolution, and disposition logic can receive different validation outcomes.

A component PASS does not validate the full architecture.

D.29 Architecture-Level Evaluation

Architecture-level evaluation requires substantive non-inferiority, improved target detection, controlled false positives, legitimate-provenance preservation, symmetry, acceptable HOLD, decision retention, absence of new L4 interactions, and proportionate cost.

D.30 Metric Failure Conditions

A metric should be revised or removed where scorers cannot apply it consistently, it fails positive controls, overreacts to negative controls, duplicates another measure, rewards gaming, or adds no decision value.

D.31 Current Maturity

Units, denominators, formulas, acronyms, and the scorecard are specified. Scoring manuals, inter-rater reliability, thresholds, non-inferiority margins, statistical power, external replication, and operational utility are not established.

D.32 Final Measurement Statement

The protocol evaluates accuracy, sensitivity, coverage, maturity, object precision, symmetry, correction, disposition, auditability, and operational burden.

The architectural benefit function is:

□ Protocol Benefit=Improved Target Control+Substantive Non-Inferiority+Symmetry+Decision Retention−New Failure and Operational Cost

This is not yet a calibrated equation.

Related Source and Reference Pages


This article belongs to the public essay layer of RATIUM.AI. For readers who want to move from this article into the broader source, technical, and orientation layers of the project, the following pages provide the relevant entry points.


Articles

The articles page gathers the public essay layer of RATIUM.AI, including arguments on stable AI governance, decision-control architecture, visible governance versus real authority, universal reason, technical competence, purpose governance, and the doctoral-scale framing of CEP.


Foundational Source Dossier

The foundational source dossier presents the deeper intellectual corpus behind CEP, LoopGuard-AI, and the broader RATIUM.AI research structure.


Technical & Reference Dossiers

The technical and reference dossier page collects architecture, visual explanation, methodological context, FAQ material, and technical source pages related to LoopGuard-AI and CEP.


RATIUM.AI / LoopGuard-AI / CEP FAQ

The RATIUM.AI / LoopGuard-AI / CEP FAQ provides a structured orientation to the main concepts behind RATIUM.AI, CEP, and LoopGuard-AI, helping readers navigate the framework through clear questions, definitions, and internal conceptual links.

RATIUM.AI — LoopGuard-AI governance architecture and Central Equilibrium Problem research by Benny Dunavich, focused on AI governance, cognitive duality, Pareto efficiency, decision-control systems, auditability, evaluation architecture, and stable governance layers for AI systems.

bottom of page