The Open Scholarly Operating System: An Open-Source Architecture for Persistent Identification, Preservation, Normalization, Classification, and AI-Assisted Research

John Swygert

August 27, 2026

DOI: [to be assigned]

A Secretary Suite Project

Abstract

This paper proposes an open-source scholarly operating system: a public, interoperable architecture that treats scholarship not merely as a collection of files and links, but as a persistent, structured, relational record of intellectual objects. The system combines immutable work identification, authorized preservation of original artifacts, standardized machine-readable renderings, revisable multi-axis classification, provenance and version control, citation and evidence graphs, and an AI research agent capable of searching, comparing, explaining, and tracing scholarly claims. The proposal is motivated by a transition already visible in scholarly discovery. Google Scholar Labs now decomposes detailed research questions into topics, aspects, and relationships, searches across them, evaluates candidate papers, and supports follow-up questions; Scholar Quick Read further summarizes answers, methods, and caveats for accessible papers. These developments demonstrate that AI-assisted scholarly navigation is no longer speculative. The remaining opportunity is architectural: to make identification, preservation, normalization, classification, provenance, and machine-assisted research interoperable components of an open scholarly commons. Google Scholar is unusually well positioned to implement many of these capabilities, but the architecture should not depend upon any single company. If no incumbent builds the complete system, this paper recommends that Secretary Suite, an independent consortium, or another open-source community take the project forward. The objective is not to abolish journals, DOIs, repositories, libraries, or existing discovery systems. It is to connect them through a durable scholarly layer in which abundance becomes navigable rather than exclusionary.

1. From Scholarly Search to Scholarly Infrastructure

The scholarly problem has changed. For much of the print era, access was constrained by physical scarcity, distribution, cataloguing, and institutional reach. The networked era produced the opposite difficulty: an expanding abundance of papers, versions, repositories, preprints, datasets, software, commentary, corrections, and citations distributed across heterogeneous systems. The central problem is increasingly not whether scholarship exists, but whether it can be reliably identified, preserved, connected, interpreted, and revisited.

Search engines substantially reduced the discovery problem, but discovery alone does not create a durable scholarly record. A search result can point to a paper whose URL later changes. The same work can appear in several versions. Metadata can be incomplete. A citation can refer to a preprint while another refers to the version of record. A paper can be corrected or retracted while copies of the earlier state remain in circulation. These are identity, provenance, preservation, and relationship problems.

An open scholarly operating system should therefore sit above and between existing repositories, publishers, identifier registries, libraries, and discovery engines. It should not require those systems to disappear. Its role is to make their objects legible to one another and to researchers.

2. The Transition Has Already Begun

Google Scholar provides an important demonstration of the direction of travel. Scholar Labs, introduced in 2025 and expanded in 2026, accepts detailed research questions, identifies their key topics, aspects, and relationships, searches Scholar across those components, evaluates results, explains why individual papers may be useful, and permits follow-up questions. Google described the launch as a new direction for Scholar. In June 2026, Google reported major increases in speed, depth, and daily search capacity for Labs. [1,2]

On August 25, 2026, Google Scholar announced Quick Read, which presents query-focused summaries organized around answers, approach, and considerations for papers the user can access. [3] These features move scholarly discovery beyond a ranked list of titles and toward reasoning over the relationship between a research question and the contents, methods, and limitations of papers.

This paper takes that trajectory seriously but does not assume that Google, or any other company, will build every necessary layer. The proposal therefore treats the complete architecture as an open specification.

3. The Scholarly Object Is Not the File

A scholarly work may exist as an author’s manuscript, preprint, accepted manuscript, publisher HTML page, PDF, repository copy, translation, corrected edition, or later revision. The file is an embodiment of the work; it is not necessarily the work’s complete identity.

The system should distinguish the intellectual work, a particular version or state of that work, and a particular file or representation of that version. Relationships among these levels should be explicit and machine-readable.

4. Two Coupled Identifiers, Not One Overloaded Number

The architecture should use an immutable Scholarly Object Identifier (SOI) for identity and a separate Scholarly Classification Address (SCA) for intellectual placement. The first never changes merely because interpretation changes. The second is deliberately revisable.

This separation avoids a weakness in classification-bearing identifiers. A paper initially categorized in one field may later become important to several others. Identity should remain stable while classification can evolve. The classification address should also be plural: digital scholarship does not need to occupy one shelf.

5. The Scholarly Classification Address

The SCA should be inspired by the navigational virtues of library classification without inheriting the physical constraint that one book must stand in one place. Its purpose is not merely to label a field. It should expose structured relationships that help humans and machines navigate the literature.

Classification dimensions could include discipline, subdiscipline, method, data type, object of study, theoretical framework, geographic scope, temporal scope, evidentiary status, and replication status. Human curators, authors, publishers, and machine classifiers could all contribute, with provenance retained rather than collapsed into a single unexplained label.

AI makes this system practical at a scale that manual cataloguing alone cannot reach. Machine classification should nevertheless remain inspectable and correctable.

6. Preservation: Keep the Original

Discovery without preservation is fragile. The operating system should permit authors, publishers, repositories, and rights holders to authorize preservation of an immutable archival copy. The preserved original should remain exactly as supplied, including its pagination, typography, figures, errors, and other documentary characteristics.

Preservation should not be confused with claiming ownership. The architecture should record the rights basis under which an artifact is stored and served. Where public redistribution is not authorized, the system can preserve metadata, checksums, relationships, and lawful links without presenting a public copy.

Content-addressed checksums should allow the system to demonstrate that an archived artifact has not silently changed. New states should be registered as versions rather than overwriting history.

7. Normalization: Preserve One Artifact, Render Another

In addition to the original artifact, the system should generate a normalized scholarly representation. This is where consistency becomes an information technology rather than a cosmetic preference.

The normalized representation could expose title, authors, affiliations, abstract, sections, equations, figures, tables, references, datasets, software, funding, conflicts, corrections, and version history through a predictable structure regardless of the typography of the source. The normalized edition should always link back to the preserved original and clearly identify itself as a machine-generated or machine-assisted representation.

This would allow every paper to be read through a consistent interface without destroying the evidentiary value of the original document.

8. Provenance Must Be First-Class

Every transformation should have provenance. If an AI system extracts an equation, repairs a malformed reference, translates an abstract, identifies a dataset, or generates a normalized rendering, the system should record what transformed what, when, with which software or model, and under which rules.

Crossref’s 2026 position on persistent identifiers reinforces the broader point: identifiers alone are insufficient without rich linked metadata, services, interoperability, sustainability, and governance. [4] The proposed system adopts that principle and extends it to the full scholarly object lifecycle.

9. From Citation Graph to Evidence Graph

A citation tells us that one work refers to another. It does not tell us why. The next-generation scholarly graph should represent richer relationships: supports, contradicts, replicates, fails to replicate, extends, reuses method, reuses dataset, corrects, retracts, critiques, translates, derives from, or independently converges upon.

These relationships should be claims with sources, confidence, and provenance rather than unquestionable machine judgments. This transforms literature review from a flat list into a traversable evidence network.

10. The AI Scholarly Agent

The natural user interface is both conventional search and an AI agent. Users should not be forced to choose between them. A simple search box remains ideal for known-item retrieval; an agent is better for multi-stage questions.

The agent should be capable of decomposing a research question, searching multiple indexes and repositories, following citations, comparing versions, distinguishing primary evidence from commentary, identifying methodological differences, exposing contradictory findings, and returning the provenance of every substantive conclusion.

Google Scholar Labs already demonstrates part of this interaction model. [1,2] An open system should generalize it so that the agent can operate across interoperable scholarly infrastructure rather than being tied to a single proprietary index.

11. Interface: Search, Browse, or Ask

The front end should be deliberately familiar. One landing page can provide three complementary modes: Search, Browse, and Ask. Search behaves like a traditional scholarly search engine. Browse exposes the classification graph. Ask opens the AI research agent.

A dedicated mobile and desktop application could maintain reading lists, annotations, alerts, project rooms, citation collections, saved evidence trails, and offline normalized copies where licensing permits. The same identity and provenance layer should underlie the website, app, API, and agent.

12. Abundance Is a Feature If Navigation Improves

Traditional scholarly filtering partly developed because human attention is scarce. That scarcity remains, but AI changes the amount of material a researcher can inspect intelligently. Weak, preliminary, unconventional, or later-disconfirmed work can still reveal abandoned approaches, negative evidence, overlooked observations, or conceptual combinations that ignite new research.

The response to abundance need not be exclusion. It can be better classification, provenance, ranking, evidence tracing, methodological inspection, and user-controlled filtering. Peer review, editorial judgment, replication, institutional expertise, and specialist databases remain valuable; the operating system adds another conduit rather than requiring older conduits to vanish.

13. Open by Architecture, Not Merely by License

An open-source scholarly operating system should avoid becoming a new single point of dependency. Its schemas, identifier specifications, export formats, APIs, provenance model, and core software should be openly documented. Data should be portable.

Existing initiatives demonstrate that open scholarly infrastructure is already a serious institutional objective. OpenAlex has publicly recommitted to the Principles of Open Scholarly Infrastructure, while Crossref emphasizes open metadata, interoperability, and sustainable governance. [4,5] A recent computational-architecture proposal for open scholarly infrastructure likewise emphasizes openness, sustainability, interoperability, reproducibility, orchestration, scalability, and portability. [6]

The present proposal differs in scope: it specifies the user-facing and knowledge-object architecture that can sit across such infrastructure, joining preservation, normalization, classification, provenance, evidence relationships, and AI-assisted research.

14. Coexistence With Existing Scholarly Infrastructure

The system should ingest and preserve existing identifiers rather than pretending they never existed. A scholarly object record can contain a DOI, author identifiers such as ORCID, institutional identifiers such as ROR, repository handles, publisher identifiers, ISBNs, dataset identifiers, software identifiers, and the proposed open SOI.

Multiple identifiers are not necessarily duplication. They represent different governance systems and purposes. The operating system’s job is to resolve and relate them. Likewise, the normalized representation should never masquerade as the publisher’s version of record.

15. Minimum Viable Open Scholarly Operating System

  • An immutable scholarly-object identifier independent of disciplinary classification.
  • A revisable, multi-address scholarly classification system with provenance.
  • Version and manifestation records that distinguish work, version, and file.
  • Authorized archival preservation with cryptographic integrity checks.
  • A standardized, accessible, machine-readable scholarly rendering that never replaces the original.
  • Open metadata and APIs connecting existing identifiers for works, people, institutions, repositories, datasets, and software.
  • A provenance ledger for every transformation, extraction, correction, and machine inference.
  • A citation graph extended toward evidence, replication, contradiction, correction, and derivation relationships.
  • Traditional search, graph browsing, and an AI scholarly agent in the same interface.
  • Exportability, transparent governance, open schemas, and a credible continuity plan.

16. Implementation Path

The system does not need to begin by ingesting all scholarship. A credible first implementation could focus on openly licensed papers and author-authorized deposits. It could mint SOIs, preserve originals, create normalized representations, classify works across multiple dimensions, connect existing identifiers, and expose the corpus through search and an AI agent.

A second stage could add claim extraction and evidence relationships. A third could federate with external repositories and indexes so that institutions can run compatible nodes while participating in a shared graph.

Google Scholar is a natural candidate to build much of this architecture because it already possesses large-scale scholarly discovery and an emerging AI research interface. [1-3] But this paper intentionally does not make implementation contingent upon Google.

If Google does not build the complete architecture, the specification should remain available for implementation by Secretary Suite, a nonprofit consortium, universities, libraries, open-infrastructure organizations, or another capable open-source community. The important object is the architecture, not the corporate logo attached to it.

17. Governance and the New Gatekeeper Problem

Any system that becomes the dominant interface to scholarship acquires gatekeeping power even if it began by reducing older forms of gatekeeping. Search ranking, inclusion rules, entity resolution, AI summaries, and classification can all influence what researchers see.

The solution is not to pretend that mediation can disappear. It is to make mediation inspectable. Inclusion criteria, machine transformations, provenance, correction channels, export mechanisms, and governance should be sufficiently transparent that researchers are not trapped inside an unexplained epistemic layer.

An open implementation also creates competitive pressure. If one interface becomes untrustworthy, the underlying open graph and identifiers should remain usable by another.

18. Conclusion

The next major scholarly platform should not merely find papers. It should know which intellectual object a paper represents, preserve authorized originals, distinguish versions, render scholarship consistently, classify it without forcing it onto one shelf, expose provenance, connect citations to richer evidentiary relationships, and provide an AI agent capable of navigating the resulting graph.

Google Scholar’s recent AI features demonstrate that the movement from keyword retrieval toward question-centered scholarly assistance is already underway. The opportunity is to carry that transition to its architectural conclusion while preserving interoperability and avoiding dependence on a single vendor.

The scholarly problem of the twentieth century was scarcity of access. The scholarly problem of the twenty-first is abundance without sufficient architecture. An open scholarly operating system would treat that abundance not as something to suppress, but as a body of knowledge to identify, preserve, connect, interrogate, and make navigable.

References

[1] Google Scholar Blog. “Scholar Labs: An AI Powered Scholar Search.” November 18, 2025.

[2] Google Scholar Blog. “Scholar Labs update: Search 10x faster, 3x deeper.” June 3, 2026.

[3] Google Scholar Blog. “Quick Read: see how a paper answers your question.” August 25, 2026.

[4] Crossref. “Persistent identifiers in research infrastructure policy: the need for a holistic approach.” Crossref Position Paper, July 20, 2026. DOI: 10.13003/q4vu-l2mw.

[5] OpenAlex. “Recommitting to the Principles of Open Scholarly Infrastructure (POSI).” March 30, 2026.

[6] Heibi, Ivan; Petrella, Mario; Di Iorio, Angelo; Peroni, Silvio. “Towards a Definition of the Computational Architecture of Open Scholarly Infrastructures.” arXiv, August 24, 2026.

*Copyright © John Swygert 2026; TSTOEAO.com; IvoryTowerJournal.com; SecretarySuite.com; TSTOEAO Room: https://chatgpt.com/g/g-6a6d017c73bc8191a6f5de01f7beab5; Ivory Tower Publishing.

Leave a Reply

Scroll to Top

Discover more from Ivory Tower Journal - ISSN: 3070-9342

Subscribe now to keep reading and get access to the full archive.

Continue reading