Skip to main content

ScholarMark — live beta with institutions · Public launch coming soon

Join waitlist
All Research & Dispatches

September 19, 2026

Research Integrity Infrastructure: Verify Before Publication

Scholarly integrity does not fail on ethics — it fails at intake. A structural look at pre-publication attestation, deterministic citation validation, and persistent provenance, and what a scoped institutional pilot actually returns.

#Research Integrity#Scholarly Publishing#Pre-Publication Verification#Citation Validation#Reproducibility#Peer Review#AI Governance#Institutional Infrastructure#Provenance#Research Evaluation#DecentraSec institutional pilot grant#DecentraSec AI Integrity Layer#DecentraSec Integritas Vault#DecentraSec GEAR Network#ScholarMark research integrity verification#ScholarMark attested-record standard#DecentraSec United States research universities#DecentraSec European Union research institutions#DecentraSec United Kingdom research offices#DecentraSec Pakistan HEC evaluation
Research Integrity Infrastructure: Verify Before Publication
EDITORIAL // RESEARCH

Scholarly Integrity as Critical Infrastructure: Why Verification Must Move Before Publication

By DecentraSec Team

The integrity crisis in scholarly communication is not a failure of researcher ethics. It is a failure of infrastructure. The default intake points in the research record still trust centralized, self-reported, human-inspected data, and no volume of reviewer goodwill can repair a system whose verification remains manual. arXiv alone now holds close to three million manuscripts, and they do not fit through a reading process. The remedy is architectural: verification must move from post-hoc correction to pre-publication attestation that intake enforces and Decentralized Provenance retains. Institutions that adopt attestation first will not merely avoid retractions — they will hold the only verifiable record in the room.

Three controls carry the argument, and none of them is a storage slogan.

Decentralized Provenance means an append-only, independently replicated record of attestations held by the institutions that create and consume them, not by a single archive. The decentralization is institutional, not ornamental: because multiple parties hold copies and compare them, a silent rewrite becomes a visible disagreement. What changed, who approved it, and when remains reconstructable without asking a central party to vouch for its own database.

Mathematical Validation means treating a citation as a structured tuple — identifier, author string, title, venue, year — and evaluating each field as a predicate against an authoritative registry such as Crossref, DataCite, DBLP, or arXiv. The validator returns resolution metadata or a contradiction; it never guesses. At the decision boundary the check is deterministic, and it produces a receipt, not a probability.

Algorithmic Integrity means every automated or AI-assisted judgment is decomposed into named dimensions — novelty, method soundness, data adequacy, statistical validity — and logged with the model, inputs, prompt, and supporting evidence. It is not the absence of model bias; it is the presence of an attributable record that makes bias detectable.

The Intake Failure: Fabricated Citations Now Operate at Machine Scale

In 2025, roughly one in twenty accepted papers at NeurIPS and USENIX Security — venues that rejected the overwhelming majority of submissions — carried at least two references to research that does not exist (arXiv:2607.00738; doi:10.48550/arXiv.2607.00738). Each of those papers cleared three or more expert reviewers. GPTZero's independent audit of 4,841 papers separately confirmed more than a hundred phantom citations across fifty-one accepted NeurIPS papers.

Those reviewers were not careless. They were human, and the record had stopped being human-scale. arXiv now holds close to three million manuscripts; in 2026 the repository moved to first-time submitter endorsement and announced one-year bans for hallucinated references. A separate audit of 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN, and PubMed Central estimated roughly 146,900 hallucinated citations in 2025 alone (arXiv:2605.07723; doi:10.48550/arXiv.2605.07723). GhostCite's review of 2.2 million citations across 56,381 papers found 1.07% of papers carrying at least one invalid citation — an 80.9% surge during 2025 — and reported model hallucination rates from 14.23% to 94.93% (arXiv:2602.06718; doi:10.48550/arXiv.2602.06718). The range is the operational point: no reviewer can intuit which submission is affected.

A reference is a checkable claim about existence, authorship, and venue. The default intake path still does not check it as one. Reviewers read it. Reading does not scale; resolution does. Resolution is Mathematical Validation: normalize the reference, query the authoritative registry, compare the returned metadata with the claimed fields, and record the result as a receipt. The validator can return "resolved," "contradiction," or "unresolved," and each state has a defined institutional action. This is not a substitute for substantive review; it removes a class of checkable errors so reviewers spend attention on whether the cited work actually supports the claim. Policy must therefore bind at intake, not after publication. The AI Integrity Layer is built to enforce exactly that.

The Reproducibility Debt: It Was Never a Culture Problem

Machine-learning research fails to reproduce chiefly because of unpublished data and code and sensitivity to training conditions — not primarily because researchers refuse to share (arXiv:2307.10320; doi:10.48550/arXiv.2307.10320). The packaging practice exists: Docker-based artifact preparation has been documented for years (arXiv:2308.14122; doi:10.48550/arXiv.2308.14122). What packaging alone does not provide is proof. A 2026 study of 5,298 Docker builds found that Docker does not guarantee reproducibility under any tested definition, and that no single set of Dockerfile rules produces reproducible images (arXiv:2601.12811).

That is the diagnosis. Reproducibility is not a virtue to be encouraged; it is a provenance gap. A result cannot be reliably reproduced when its compute environment, dependency graph, data version, and random seed never entered a record that survives the researcher leaving the institution. Institutional repositories preserve the paper. Almost none preserve the artifact state that produced it: image digest, pinned dependencies, build command, data lineage, and evaluation harness. Artifact attestations are useful only if they outlive the originating lab; Decentralized Provenance is what makes the record survive institutional turnover.

Where a data-management or reproducibility requirement is satisfied by a plan rather than a machine-checkable artifact, the institution holds attestation of intent, not evidence of lineage. That is audit exposure, and it compounds with the half-life of your storage links. The remedy is an append-only record of artifact state, replicated across the institutions that create and consume it — the Integritas Vault.

AI Peer Review: Fluent, Biased, and Unauditable

More than half of researchers now report using AI in review tasks (Frontiers, n=1,645; Nature d41586-025-04066-5). That is the baseline, not a forecast. An evaluation of ICLR 2026 and Nature Communications reviews found models producing overly positive, weakly grounded, systematically biased assessments (arXiv:2608.03581; doi:10.48550/arXiv.2608.03581). Positivity bias is the dangerous direction: it admits weak work.

Auditability fails at a deeper level. A single aggregate recommendation collapses novelty, method soundness, data adequacy, and statistical validity into one number no one can interrogate. When a model produced or assisted that number, the institution holds no attributable record of how. If your faculty review for indexed venues — and for promotion, they must — your institution's name sits behind decisions it cannot reconstruct six months later. Algorithmic Integrity requires that institutions decompose quality into named dimensions and log each as attributable evidence — model, prompt, inputs, and the text fragment supporting the judgment — rather than reduce the decision to a score. The score may remain; the evidence behind it must exist.

Venue Identity: Whitelists Are Stale, and Journals Are Being Acquired

Nature's 2025 investigation documented 36 legitimate, indexed journals acquired by recently formed firms that then raised fees and churned papers (d41586-025-01198-6). The journal's legitimacy was real. Its ownership changed. Whitelists kept it listed because the fields they track — title, ISSN, index status — did not change. COPE's September 2025 retraction guidelines now explicitly target paper mills and third-party interference, while Sage closed one investigation with a final batch of 678 retractions, bringing the total above 1,500.

Journal legitimacy has lived as a static list that a third party maintains. Venue identity and ownership instead require persistent, attested records that update when control changes — a snapshot cannot track a moving target. Those records must be replicated across the institutions that consume them; otherwise the whitelist is just another centralized database with a stale cache. A quietly acquired venue on your approved list makes every publication routed there a liability, and the funder conversation will be retrospective.

National Evaluation Regimes: Volume Without Verifiability

Pakistan's system produces 30,000–35,000 Scopus-indexed papers annually, with retraction rates above OECD benchmarks and persistent predatory-publishing patterns (DOI 10.1093/reseval/rvae053) — despite HJRS tiering driving promotion decisions. The incentive geometry is general, not local: national evaluation regimes reward self-reported volume, so volume is what gets produced. Institutions assume legitimacy at reporting and verify it, if ever, after the fact.

These failures resolve to one principle. Verification moves to pre-publication attestation, enforced at intake. An institution that attests at intake produces a record a national evaluation body can audit rather than accept on trust. That is the institutional form of Decentralized Provenance: reported volume becomes independently verifiable because the underlying attestations are replicated and comparable, and the institution becomes the reference case rather than the risk case.

What an Institutional Research-Integrity Pilot Looks Like

Run a scoped, time-boxed pilot on a defined cohort: one faculty, one department, one intake cycle. Let your own manuscripts produce the baseline audit. The pilot returns a measured pre-publication defect rate — invalid citations intercepted, artifacts attested, review decisions logged, venue statuses confirmed. That is a number a Dean can take to a Senate or a funding body. It is not a vendor's claim; it is the institution's own intake data under the three controls above.

DecentraSec is issuing an Institutional Pilot Grant: fixed scope, no cost, limited cohort selected for diversity of institutional size, discipline mix, and regional evaluation regime. Participants become the reference institutions for the attested-record standard. Institutions committing during the pilot window qualify for an early-adopter subsidy covering onboarding of the AI Integrity Layer, Integritas Vault, and GEAR Network as a single institutional deployment, funded from research-integrity or compliance budget lines.

Verification before publication is not a culture reform. It is Integrity Infrastructure — and it belongs in the same budget as the record it protects.


ScholarMark by DecentraSec is building the pre-submission infrastructure that academic publishing has never had — AI-powered integrity checks, paid peer review via the GEAR Network, and immutable provenance-based authorship seals. Start here →

References

  • Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences, arXiv:2607.00738, doi:10.48550/arXiv.2607.00738.
  • LLM hallucinations in the wild: Large-scale evidence from non-existent citations, arXiv:2605.07723, doi:10.48550/arXiv.2605.07723.
  • GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models, arXiv:2602.06718, doi:10.48550/arXiv.2602.06718.
  • arXiv submission policy, 2026 — submitter endorsement and hallucinated-reference sanctions.
  • Semmelrock et al., Reproducibility in Machine Learning-Driven Research, arXiv:2307.10320, doi:10.48550/arXiv.2307.10320.
  • Canesche et al., Preparing Reproducible Scientific Artifacts using Docker, arXiv:2308.14122, doi:10.48550/arXiv.2308.14122.
  • Malka et al., Docker Does Not Guarantee Reproducibility, arXiv:2601.12811, doi:10.48550/arXiv.2601.12811.
  • Fichtl et al., AI-Assisted Peer Review Across Research Communities, arXiv:2608.03581, doi:10.48550/arXiv.2608.03581.
  • Nature d41586-025-04066-5 — AI adoption in peer review.
  • Nature d41586-025-01198-6 — acquisition of indexed journals.
  • COPE — retraction guidelines, September 2025.
  • Research Evaluation, DOI 10.1093/reseval/rvae053 — publication incentives and retraction rates.

02 // RELATED RESEARCH · ARCHIVE DISPATCHES

Related Papers & Dispatches.

Peer-reviewed analyses, cryptanalysis papers, and zero-trust engineering dispatches.

05 // NEWS & MILESTONES · COMPANY DISPATCHES

Latest Updates & Strategic Milestones.

News, institutional pilot rollouts, and engineering milestones from the DecentraSec core team.

06 // RESEARCH & DISPATCHES · EDITORIAL ARCHIVE

From The Engineering & Research Team.

Deep-dive analyses, cryptanalysis papers, post-quantum protocols, and zero-trust systems design.

07 // INSTITUTIONAL INTAKE · STRATEGIC ONBOARDING

Formal Onboarding & Strategic Inquiries.

DecentraSec works with universities, institutional investors, Tier-1 peer reviewers, and Open Access contributors through a structured intake process.

SLA: 24–48 HOUR REVIEW WINDOW

INTAKE VERIFICATION PROTOCOL

All submissions are encrypted at rest and routed directly to DecentraSec core engineering officers under a strict mutual non-disclosure protocol.

DIRECT HEADQUARTERS

A-201, Block-12, Gulistan-e-Jauhar, Karachi, Pakistan

INSTITUTIONAL INQUIRIES

partners@decentrasec.com

DIRECT INSTITUTIONAL LINE

+92 310 1288813
SECURE INTAKE PORTALSLA: 24–48 HOUR REVIEW WINDOW

Select your institutional pathway. Requests are reviewed for strategic alignment, security posture, and technical capacity.

Chat with us