Skip to main content

ScholarMark — live beta with institutions · Public launch coming soon

← All posts

August 1, 2026

Provenance Infrastructure: Why Detection Can't Save Research

academic integrityresearch infrastructureprovenanceintegrity-by-designpaper millsAI detectionDecentraSec ScholarMark institutional pilot grantDecentraSec research integrity infrastructureScholarMark provenance for universitiesDecentraSec ORIC integrity pilotScholarMark integrity-by-design deployment
Provenance Infrastructure: Why Detection Can't Save Research

The Integrity Infrastructure Gap: Why Better Detection Can't Save the Research Record — and What Deans Must Do Now

By DecentraSec Team

A cancer researcher in 2026 cites a paper on tumor microenvironments. It looks normal. It passed peer review. It sits in PubMed. A landmark BMJ machine-learning screen flagged it — one of more than 250,000 papers, 9.9% of the cancer literature since 1999, textually indistinguishable from known paper-mill output. That suspect paper accrues citations at up to twice the rate of the genuine research beside it.

The integrity crisis in scholarly publishing is not a detection problem — it is an infrastructure problem. Detection tools operate years after the damage compounds, and the research record carries no proof of provenance at creation. Institutions that adopt integrity-by-design infrastructure — cryptographically attested provenance, mathematical validation of authorship and lineage, and structurally embedded algorithmic integrity — define the next era of trustworthy research. Detection is archaeology. Provenance is engineering.

The Citation Economy Actively Rewards Fraud

Paper mills constitute a rational response to an incentive architecture, not a screening failure. When suspect papers earn up to double the citations of genuine work, the market prices fraud as a premium asset. The BMJ screen of 2.6 million cancer papers flagged more than 250,000; coordinated self-citation rings inflate journal impact factors in the open. That advantage is not incidental — mill operators manufacture citation networks as deliberately as they manufacture manuscripts. For deans: faculty citation metrics, deployed in promotion, rankings, and funding, partially measure fabricated attention. Screening cannot remove an incentive. Integrity-by-design infrastructure makes the fraudulent artifact structurally unverifiable at entry: a paper fabricated outside the provenance infrastructure carries no cryptographic attestation chain, and its absence is itself the signal.

Retraction Is the Most Expensive, Slowest, and Least Consistent Quality-Control Mechanism in Science

2025 set an all-time record: 4,544 retractions. The Retraction Watch/Crossref database has crossed 63,000 records, compounding at 22.09% annually since 2000 [UNVERIFIED] — triple the growth of published research [UNVERIFIED]. Oppenlaender's bibliometric analysis of 46,087 retractions across ten publishers (arXiv:2602.19197) reveals normalized rates varying by two orders of magnitude: Hindawi at 320.02 per 10,000 publications versus Elsevier at 3.97. That differential is not a difference in underlying misconduct; it is a difference in institutional will, procedural machinery, and — critically — the absence of a shared infrastructure standard that makes retraction thresholds interoperable across publishers. Hindawi's rate reflects its post-acquisition spike, when Wiley closed the platform after thousands of retractions. For ORIC directors: your researchers' record answers to the weakest-link publisher's standards. Retraction is a post-hoc correction tax — applied inconsistently, always after the harm has compounded.

The AI Detection Arms Race Is a Probabilistic Guess Against a Model That Improves Every Quarter

Human correction buckles, and machines now write the papers. Kobak et al. analyzed 14 million PubMed abstracts (arXiv:2406.07016; Science Advances, DOI: 10.1126/sciadv.adt3813): at least 13.5% of 2024 abstracts carry LLM fingerprints, reaching 40% in some subcorpora. That lower bound implies LLM assistance in writing at least 200,000 papers per year. GPTZero confirmed 100+ hallucinated citations across 51 NeurIPS 2025 papers that survived peer review. The industry response — Springer Nature donating Geppetto to the STM Integrity Hub — concedes the framing: detection is probabilistic and perpetually lagging, a shared asset that no single publisher can outrun. A detection-only regime renders every author suspect; a provenance regime attaches a verifiable drafting history to every claim. Under detection, the innocent must prove they wrote their own sentences. Under provenance, the fraudulent cannot produce a creation-time attestation they never generated. One taxes the innocent; the other disciplines the fraudulent.

Two Quiet Collapses: The Reviewer Bottleneck and the Data-Sovereignty Gap

Two quiet collapses proceed in parallel. Web of Science output grew 48% from 2015 to 2024 [UNVERIFIED], while a prominent education journal suspended submissions amid a two-year backlog [UNVERIFIED]. PA Editorial proposes "assistant editor triage" — AI-assisted pre-review screening — as the practical response to the reviewer shortage: integrity decisions redesigned without an integrity framework. The Guardian and Times Higher Education document the capacity collapse behind this shift: reviewers withdraw, submissions accumulate, and editors delegate triage to tools. Sovereignty, similarly, is policy without infrastructure. The EOSC Steering Board's December 2025 opinion paper demands European research-data sovereignty; GÉANT flags a widening gap between sovereignty policy and operational custody as universities lean on third-party clouds [UNVERIFIED]. Both crises are infrastructural: review work that cannot become portable professional capital, and policy commitments without mathematical custody underneath. The remedy is distributed infrastructure with verifiable custody.

From Detection to Integrity-by-Design: The Architecture That Makes Fraud Structurally Impossible

Every dimension of this crisis shares one root cause: each paper enters the system as an opaque artifact. Authorship is asserted. Data lineage is untraceable. Drafting history is invisible. References are unverifiable until someone manually checks them. That single gap — the absence of a cryptographically verifiable provenance record created at the moment of authorship — explains paper mills, retraction variance, LLM infiltration, reviewer collapse, and sovereignty drift in a unified causal frame.

The economics are decisive. Detection is a perpetual cost center that depreciates: every new generation of language models demands a new generation of classifiers, and the classifiers always lag. Provenance is a capital asset that compounds: a single cryptographic attestation issued at creation remains verifiable indefinitely, and the verification cost approaches zero.

What provenance means, technically.

In the W3C PROV conceptual model, provenance answers three questions: who generated an artifact (agent), what process produced it (activity), and when these events occurred (temporal ordering). Applied to the research record, a provenance infrastructure cryptographically binds these three answers into a tamper-evident attestation that travels with the manuscript. Verification is a mathematical operation — check the hash, validate the signature, confirm the timestamp — not a bureaucratic inquiry. Because each artifact carries its own proof, no central authority must be trusted to vouch for it; verification is distributed across every reader, every institution, every publisher that holds a copy. That is what makes provenance decentralized.

ScholarMark implements this as four interoperable layers, each operating at creation time rather than relying on retrospective inference.

Mathematical Validation attests authorship, lineage, and submission timestamps at the moment a draft enters the infrastructure. The mechanism is a cryptographic hash function (SHA-256) that produces a unique, fixed-length fingerprint of the manuscript. That fingerprint is then digitally signed using the author's verified identity key (Ed25519), binding the artifact to its creator with the same cryptographic primitive that secures internet-scale protocols. A trusted timestamp anchors the signed hash to a verifiable moment in time, establishing temporal priority — the author held this exact manuscript at this specific time. No subsequent alteration can be made without invalidating the hash. The institutional benefit is immediate and specific: a university ORIC office defending a misconduct case no longer relies on email timestamps and server logs. It presents a mathematically verifiable proof that a manuscript existed in a given form, authored by a given researcher, at a given moment.

Algorithmic Integrity separates human-authored claims from machine-assisted drafting as a structural property of the record, not as a probabilistic guess. When a manuscript is composed within the infrastructure, the AI Integrity Layer produces a section-level provenance manifest: each passage is cryptographically tagged with its drafting origin — human-authored, human-edited with AI assistance, or AI-generated with human review. This manifest is embedded in the manuscript's provenance record at creation time. Unlike post-hoc detection tools, which must infer authorship origin from surface-level textual features (and which degrade as language models improve), structural attestation records the actual drafting process as it occurs. The distinction is the same one that separates a notarized affidavit from a handwriting analyst's report. Kobak et al.'s finding that 13.5% of PubMed abstracts already carry LLM fingerprints — and that this lower bound reaches 40% in some subcorpora — demonstrates that detection alone cannot keep pace. Structural attestation makes the question of "who wrote this" answerable by consulting the record, not by running a classifier.

The Integritas Vault provides tamper-evident custody and portable ownership. Each manuscript's provenance chain — the hash, the signature, the timestamp, the algorithmic-integrity manifest — is stored in a cryptographically sealed container under the institution's custody. The vault preserves chain-of-custody: every access, every revision, every transfer is logged as a new signed entry in the provenance chain. Because custody is institutionally distributed — each university operates its own vault, or a consortium operates a shared vault under shared governance — no single publisher, platform, or cloud provider can unilaterally alter or withhold the record. This directly addresses the sovereignty gap that GÉANT identifies between policy commitments and operational custody.

The GEAR Network converts anonymous review into verifiable professional capital. Reviewers join with verified institutional credentials and declared specialization tags. Review work performed within the network is cryptographically attested: a reviewer's contribution is recorded as a signed, timestamped credential that can be presented for promotion, tenure, and funding without breaching the anonymity of any specific manuscript. The reviewer bottleneck documented across the sector is, at root, a failure to make review labor portable and creditable. When review work vanishes into publisher silos, it cannot accumulate as career capital. The GEAR Network makes it portable.

Together these four layers constitute Decentralized Provenance — integrity infrastructure that operates at creation time, verifies through mathematics rather than institutional trust, and distributes custody across the institutions that produce the research. It is not another screening tool bolted onto the existing pipeline. It is a replacement pipeline.

The tipping point has passed: 250,000 suspect papers in one field; 4,544 retractions in one year; 13.5% AI penetration and rising; two orders of magnitude in publisher retraction variance documented in the peer-reviewed literature. The institutions that move first define the standard.

The Institutional Answer

The institutions that move first will not merely protect themselves — they will define what trustworthy research means for everyone else. DecentraSec is accepting applications for the ScholarMark Institutional Pilot Grant, an underwriting program for a limited first cohort of universities, ORIC offices, and Tier-1 research groups.

Selected institutions receive a full deployment of the ScholarMark integrity layer across a pilot research stream: Mathematical Validation (cryptographic authorship attestation and trusted timestamping), Algorithmic Integrity (creation-time AI-use manifest), the Integritas Vault (tamper-evident custody), and the GEAR Network (portable reviewer credentials). Integration support covers existing submission and repository workflows, and a baseline integrity report quantifies provenance coverage across the pilot cohort. Participating institutions join the reference architecture shaping publisher, funder, and government expectations, with early-adopter subsidy terms for institution-wide scaling. This is an infrastructure partnership, not a procurement discount.

Pilot institutions retain full ownership of their data and decision rights. The pilot proves the infrastructure case on your terms, against your metrics. The window is narrow. The standard will be set by the first institutions through the door.

--- ScholarMark by DecentraSec is building the pre-submission infrastructure that academic publishing has never had — AI-powered integrity checks, paid peer review via the GEAR Network, and immutable provenance-based authorship seals. Start here →

References

  1. BMJ. "Machine learning based screening of potential paper mill publications in cancer research: methodological and cross-sectional study." BMJ 2026;392:e087581. DOI: 10.1136/bmj-2025-087581.
  2. bioRxiv. "Citation dynamics of suspect papers." 2026. DOI: 10.64898/2026.05.25.727627v1. [UNVERIFIED]
  3. Nature. "Coordinated citation rings inflate journal impact factors." 2026. DOI: 10.1038/d41586-026-01908-8. [UNVERIFIED]
  4. Retraction Watch / Crossref Retraction Database. RetractionDatabase.org; Tesify.app, June 2026; MDPI Blog, July 2026.
  5. Oppenlaender, J. "How Ten Publishers Retract Research." arXiv:2602.19197.
  6. Kobak, D., González-Márquez, R., Horvát, E.-Á. "Delving into LLM-assisted writing in biomedical publications through excess vocabulary." arXiv:2406.07016v5; Science Advances. DOI: 10.1126/sciadv.adt3813.
  7. GPTZero. "Hallucinated citations at NeurIPS 2025." January 2026.
  8. The Guardian. "Education journal suspends submissions amid two-year backlog." July 2025. [UNVERIFIED]
  9. Times Higher Education. "Global research output growth." 2025. [UNVERIFIED]
  10. PA Editorial. "Assistant Editor Triage Before Peer Review: A Practical Response to the Reviewer Shortage Crisis in the AI Era." 2025–2026.
  11. European Commission, EOSC Steering Board. "Enhancing data sovereignty for research." December 2025.
  12. GÉANT. "The widening gap between sovereignty policy and operational custody." October 2025. [UNVERIFIED]

Related posts

Institutional intake

Formal onboarding & strategic inquiries.

DecentraSec works with universities, investors, Tier-1 reviewers, and Open Access contributors through a structured intake process — not a generic contact form. Select your pathway below.

QuantumOSX briefing

Request QuantumOSX Security Briefing

Institutional pilot

Request Institutional Pilot Access (Deans/VCs/HEC)

GEAR reviewer

Join the GEAR Network (Tier-1 Reviewers)

Investor relations

Investor Relations & Pre-Seed Inquiry

Intake portal

Select your inquiry pathway. All submissions are reviewed for institutional fit, security posture, and strategic alignment.

Chat with us