October 5, 2026
Chain of Custody for Research Integrity: Institutional Proof
Detection tells you that you have a problem. Proof tells you who, what, and when. This is the case for signed, content-addressed custody of the scholarly record — held inside the institution's own trust boundary from ingestion onward.

The Cleanup Economy: Your Institution Is Audited on Proof It Cannot Produce
By DecentraSec Team
Research integrity is not failing for lack of better detectors. It is failing because institutions hold no tamper-evident proof over their own scholarly record at the point of ingestion. Retractions, paper-mill clusters, AI-distorted review, and cross-border data exposure are downstream symptoms of one upstream absence: no signed, content-addressed, institutionally held chain of custody that binds a submission to an accountable human, proves the artifact reviewed is the artifact accepted, and keeps AI-assisted decisions attributable inside the institution's own trust boundary.
By proof, this article means evidence a challenger can recompute: a content-derived artifact digest, an append-only actor-action lineage, and an institutional signature over both. By screening, it means a detector's classification of content. The difference is the difference between an audit and an accusation.
I. The Cleanup Economy: We Fund Detection and Call It Integrity
The Retraction Watch Database lists more than 67,000 retractions. The Wiley/Hindawi episode alone produced 11,300+ retractions, closed 19 journals, and cost an estimated $35–40 million in disclosed revenue; no tooling measures the reputational and accreditation cost.
The community knows the diagnosis. In PLOS Biology (DOI: 10.1371/journal.pbio.3002870), 72% of biomedical researchers report a reproducibility crisis, 67% say their institution values new research over replication, and only 16% report an institutional procedure to improve reproducibility. That last number is the operational failure behind the first two: the institution is not equipped to produce a chain of custody before a paper becomes a problem.
A plagiarism detector, an image-screening service, and a text classifier all operate on a finished artifact outside your control. They emit classifications, similarity scores, or probabilities. They do not bind a named actor to an exact artifact revision at the moment of decision, and they do not preserve the policy state under which the decision was made. None answers the four questions a funder asks: who submitted this, what exactly was reviewed, who decided, and can you show me the chain?
Screening is a downstream control. It tells you that you have a problem. It cannot tell you when, who, or what changed.
II. Five Structural Gaps No Detector Closes
- Attribution. Nothing binds a submission to an accountable human at ingestion. You cannot answer "who is responsible" with a record — only with an inference. Binding means an authenticated institutional identity issues a signed assertion over the submission digest, not that an email address appears in a form field.
- Fidelity. The artifact reviewed is not provably the artifact accepted. Revisions and silent post-acceptance edits break the chain exactly where disputes occur. A content-derived digest per revision makes substitution detectable; a platform title or version string does not.
- Attributability. AI-assisted decisions remain unattributable. At ICLR 2024, researchers estimated that at least 15.8% of reviews were AI-assisted, with AI-assisted reviews skewing more positive and raising near-threshold acceptance probability by 4.9 percentage points (DOI: 10.48550/arXiv.2405.02150). A 2026 detection study classifies roughly 20% of ICLR 2025 and 12% of Nature Communications reviews as AI-generated (DOI: 10.48550/arXiv.2602.00319). Across 111 venues, reviewer-facing AI policies remain inconsistent (DOI: 10.48550/arXiv.2608.03581). The failure is not that AI exists; it is that the disclosure is not a machine-readable field attached to the review record, so institutions cannot reconstruct the decision state. Turning that stated policy into a reconstructable decision record means enforcing AI disclosure as machine-checkable policy at ingestion, not auditing it after publication.
- Custody and jurisdiction. Storage location is not custody. A server in one country can be controlled by an entity in another and subject to a third country's disclosure rules. Custody means legal and organizational control: who may access, alter, delete, or disclose the record, and under whose law. Data-sovereignty and Indigenous data-governance frameworks make the same demand: prove the controlling entity and jurisdiction, not the data-center address.
- Replication. Replication is a publication act, not a machine-checkable one. Versioned code, model weights, environment manifests, and datasets arrive as a supplementary PDF. Without digest-addressable versions, "same data" is an assertion, not a verifiable property.
For a principal investigator, gaps 1 and 2 are not compliance burdens — they are priority-of-discovery protection. A content-addressed lineage is a timestamped, portable claim to what you had, when you had it, and what you changed.
A survey of 2,944 authors across 69 countries found firm opposition to autonomous AI decision-making or content generation in publishing, even where administrative and language assistance is accepted (DOI: 10.48550/arXiv.2606.27447). Researchers are not asking for less accountability — they are asking for accountability that is legible and human-anchored.
III. Why a Third-Party Log Is Not a Proof Record
The obvious objection: our publisher platform already logs everything. A platform log records that a session occurred; it does not bind an accountable human to a specific artifact revision. The same log line — "review submitted" — can describe two different files; only an artifact digest disambiguates them. The log is not institutionally signed, not exportable as a verifiable bundle, and not guaranteed to survive migration, vendor change, or journal closure. Recall the Wiley/Hindawi case: 19 journals closed. Where did their review logs go, and who can produce them now?
AI-based screening of cancer research publications has shown the forensic value of detection at scale, surfacing paper-mill-style patterns even in high-impact journals (bioRxiv DOI: 10.1101/2025.08.29.673016). That finding demonstrates the central point: screening is a powerful forensic tool and a poor preventive control. Use both. Own one.
If your evidence of integrity lives in a vendor's database, your integrity carries a vendor dependency — and a vendor carries a jurisdiction.
IV. The Institutional Trust Boundary: Proof You Operate, Not a Service You Buy
Define the boundary plainly: the perimeter within which your institution generates, holds, and discloses tamper-evident proof without depending on an external party's cooperation. Verification stops being a service you buy and becomes a capability you operate.
The controls: content-addressed storage and signed lineage per artifact; author and reviewer identity binding at submission; AI-use policy enforced as code; review logs linked to named humans and exact artifact digests; jurisdiction-bound custodianship; versioned, machine-verifiable datasets, code, model weights, and environment manifests.
Three primitives carry the argument.
Mathematical Validation is the deterministic layer of the proof. Every manuscript, review, dataset, model weights file, and environment manifest is addressed by a content-derived digest. Each revision record carries the digests of its inputs and predecessor, so any change to a prior artifact propagates as a broken reference. Verification is local recomputation: a challenger hashes the artifact, compares the digest with the signed manifest, and follows the revision pointers. No third party is asked to interpret the record.
Decentralized Provenance is the accountability layer. Each append-only record binds an actor, an action, an artifact digest, a policy state, and a timestamp. The institution signs the record; independent nodes operated by participating institutions witness and retain copies. Decentralization here means provenance that stays inside the institution's own trust boundary — federated notarization among participating institutions, not a public ledger and not a single vendor's database. The proof is exported as a bundle a third party can verify without access to any vendor system: digests, signatures, policy state, and custody assertions.
Algorithmic Integrity is the output. AI-use declarations, blinding rules, and conflict gates are encoded as structured, machine-checkable fields attached to the decision record — not prose in a guidance PDF. The review decision therefore carries the exact artifact digest and the disclosed tool state that preceded the decision. Integrity becomes a property an auditor can recompute rather than a claim a platform publishes.
Why this beats centralization on your own terms: the proof is portable, auditable, and independent of any third party's willingness to disclose. That is the sentence a Dean repeats to a funder. You are already accountable for the record; this is the first configuration in which you also hold it.
V. The Policy Wave Is Already Here: HEC 2025 and the Next 90 Days
Pakistan's Higher Education Commission Journals and Publications Policy 2024, amended effective 5 November 2025, reopened local journal accreditation — an explicit demand for auditable credibility rather than metric compliance. The same requirement arrives through data-sovereignty frameworks and Indigenous data-governance principles, which converge on proving custody and jurisdiction, not storage location.
Five moves, achievable without new headcount:
- Inventory the proof gap. For your ten highest-stakes outputs of the past year, can you produce the submission-to-acceptance chain without emailing a publisher? Write the answer down.
- Pick one bounded scope. One institutional journal, or one department's pipeline. Institutional change fails on scope, not ambition.
- Bind identity at ingestion. The highest-leverage control, and the cheapest.
- Enforce AI disclosure as policy-as-code. A gate, not a guidance PDF.
- Make one artifact machine-verifiable end to end. One is a demonstration; one is also the template.
The success metric to put before your council is not fewer retractions. It is time-to-proof — the measured interval from an auditor's request to a complete, signed, recomputable chain of custody for any given output. Institutions ready to baseline that number can apply for a funded pilot; see the program terms at the close of this article.
Institutional Pilot Grant & Early Adopter Subsidy
This is funded evaluation and co-development, not procurement. Access is capacity-limited by onboarding, not price-limited.
Track A — Institutional Pilot Grant (Deans, ORIC Directors, Research Integrity Officers, society-journal editors). A funded pilot over one semester to one academic year on a bounded scope: deployment within your jurisdiction, identity binding at ingestion, AI-disclosure enforcement as policy-as-code, and an accreditation-ready proof bundle you own and can submit directly to HEC, funders, or your accreditor. Deliverables include a published case study and a baseline-to-post time-to-proof measurement. The ask is a governance decision: nominate the scope and the accountable lead.
Track B — Early Adopter Subsidy (Tier-1 PIs, lab directors, core facilities). Subsidized provisioning for groups that want signed lineage over one active project now — manuscripts, datasets, model weights, environment manifests — with priority-of-discovery protection as the immediate benefit and a direct path into your institution's eventual pilot.
You cannot buy integrity after the fact. You can hold proof before it. The institutions that move first will not simply have cleaner records — they will be the ones whose records other people must accept on their terms.
ScholarMark by DecentraSec is building the pre-submission infrastructure that academic publishing has never had — AI-powered integrity checks, paid peer review via the GEAR Network, and immutable provenance-based authorship seals. Start here →
References
Retraction Watch Database · CASRAI/Wiley earnings reporting · PLOS Biology, DOI 10.1371/journal.pbio.3002870 · The AI Review Lottery, DOI 10.48550/arXiv.2405.02150 · Detecting AI-Generated Content in Academic Peer Reviews, DOI 10.48550/arXiv.2602.00319 · AI-Assisted Peer Review Across Research Communities, DOI 10.48550/arXiv.2608.03581 · A&A community survey on the future of scientific publishing, DOI 10.48550/arXiv.2606.27447 · Revealing the Paper Mill Iceberg, bioRxiv DOI 10.1101/2025.08.29.673016 · HEC Journals and Publications Policy 2024 (amended 5 November 2025).
02 // RELATED RESEARCH · ARCHIVE DISPATCHES
Related Papers & Dispatches.
Peer-reviewed analyses, cryptanalysis papers, and zero-trust engineering dispatches.

Research Integrity Infrastructure: Pre-Ingestion Checks
Research integrity failures are infrastructure failures. This post outlines pre-ingestion verification, appeal-aware records, and sovereign provenance for academic institutions.

The Verification Gap: Research Integrity Is Infrastructure
Institutions do not lack AI detectors — they lack a verifiable chain of custody. Why the 21% AI-review finding is an audit failure, how tamper-evident provenance makes retraction clusters structurally impossible, and how deans, ORIC directors, and regulators can pilot verifiable research infrastructure this accreditation cycle.

Peer Review Crisis? It's a Research Infrastructure Problem
More than eight million papers now enter a peer review system that can't verify who validated what. The fix is verifiable triage: identity-backed reviewers, screening at ingest, and audit-ready provenance—installed at the institutional layer, not bolted onto editorial software.


