The Ora Knowledge Foundation is an initiative in formation. It is not incorporated and is not accepting donations. The Knowledge Library described here is a design proposal: no Foundation library, expert committees, judicial-review process, processing pipeline, distributed node network, or cryptographic signing system currently operates.

Why this is its own paper

The Second Brain paper covers the personal vault — the user’s own substrate, accumulated through the user’s own work. The proposed Knowledge Library would extend the same commitment to civilizational scale. It is a separate paper because the intended operational form is different: the personal vault is curated by the user; a future Knowledge Library would be curated against public specifications, processed through pipelines designed to operate at scale, and hosted on decentralized infrastructure.

The two layers are intended to compose. A future retrieval system could pull from the personal vault first (Level 1 provenance, the user’s own authored or curated content) and then from a Knowledge Library where the personal vault does not have the answer (Level 3 provenance, meaning material vetted under the proposed process). The personal vault is the user’s intelligence amplified by their own substrate; the proposed Knowledge Library would extend that substrate into domains the user has not curated personally.

The four-layer operating model

The Knowledge Library is designed around a four-layer model that would pair subject-matter expertise with faithful automated execution. The constitutional governance model from earlier drafts of the project — four layers with separation of powers — was retired as the Foundation’s overall governance model, but the same pattern remains the proposal for the library specifically.

Layer 1 — Constitutional principles. Six commitments would govern library work:

  • All output is provenance-weighted and source-traceable.
  • All published datasets remain freely available and in the public domain.
  • No editorial bias is introduced by processing — the algorithms execute specifications, they do not make editorial judgments.
  • Processing specifications are public documents, open to inspection by anyone.
  • The automation serves the specification, never the reverse — if the algorithm cannot faithfully execute a specification, the algorithm is fixed or the specification is amended by its committee, not silently ignored.
  • No user interaction data is monetized, sold, or shared.

These principles would apply from the first operating version regardless of organizational size, because they are meant to govern any future automation whether or not anyone is watching.

Layer 2 — Expert committees as specification authors. If the library is built and staffed, subject-matter experts would author the specification documents for each knowledge domain. A future committee for each domain would define what qualifies as a source, how provenance is evaluated and weighted, how content is identified, processed, and atomized, what output and cross-referencing standards apply, and what edge cases require human review rather than algorithmic resolution. No such committees or published domain specifications currently exist.

Layer 3 — Human review for edge cases. Under the proposal, content that does not clearly meet or violate a specification would be flagged for review. Reviewers would interpret rather than write specifications. Their decisions could inform later amendments by the relevant committee and resolve disputes between domains. No judiciary or equivalent review body currently operates.

Layer 4 — Algorithms as faithful executors. Future automated processing pipelines would execute the specifications faithfully, deterministically, and auditably. They would not make editorial decisions. When content fell outside a specification, the pipeline would flag it for human review rather than guess. Every algorithmic decision would be traceable to the specification that authorized it.

An initial implementation could have the founder perform the human roles until qualified volunteers and organizational capacity exist. Formal committee staffing and any operating launch remain future decisions. The principles would need to be fixed before automation begins.

The universal pipeline

Every future knowledge domain would follow the same proposed processing pattern. The pattern would be shared infrastructure; only the specifications would differ.

Source identification — locate qualifying content as defined by the domain’s specification. Provenance verification — evaluate source reliability against the domain’s provenance hierarchy. Processing — run content through the document processing pipeline (any document → atomic notes → structured output). Cross-referencing — link atomic notes to related content within and across domains. Indexing — embedding and metadata tagging for retrieval-augmented generation. Publication — output published to standard distribution channels and as downloadable datasets for local use.

Under the proposal, each domain’s specification document — authored by its future committee — would govern steps 1 through 3. Steps 4 through 6 would be infrastructure-level and domain-agnostic.

This is the proposed Knower at civilizational scale. The personal Ora vault is Level 1 provenance — what the user has authored or curated. Future library domains could supply Level 3 material vetted under the proposed specifications. A future retrieval engine would pull from both layers and weight the user’s own content above external or lower-tier material, preserving the AHI commitment that the user’s kept corpus is never silenced.

Decentralized infrastructure

The library would be hosted on decentralized public-domain infrastructure rather than concentrated on Foundation servers. This is a proposed architectural commitment, not a current deployment.

A knowledge library hosted on a single Foundation server is a single point of enclosure failure. If the Foundation is captured, defunded, sued out of existence, or simply dissolved, the library disappears with it. A library hosted on distributed infrastructure persists regardless of what happens to the Foundation, which is the public-domain commitment made operational at the data layer.

Three established patterns could compose into the architecture:

Content-addressed storage. Library content would be addressed by cryptographic hash rather than by location. The same content would have the same address on any node that hosted it. IPFS-class infrastructure is the mature reference implementation; the Foundation initiative would not need to invent new infrastructure.

Distributed hosting through volunteer nodes. A future library could be hosted across independent nodes — partner organizations, volunteer operators, mirror sites at universities and libraries, contemplative-tradition digital archives, and other willing hosts. The Foundation initiative operates no library nodes and coordinates no such network today. Internet Archive’s distributed-mirror practice is the closest peer-group precedent for the design.

Cryptographic provenance verification through the P1–P6 hierarchy. The existing Ora provenance hierarchy could be extended through cryptographic signing of canonical documents at each level. A user retrieving a document from the future library could then verify its provenance level without trusting the node that served it. No cryptographic signing service currently implements this design.

If this architecture is built, the Foundation’s proposed role would be signing authority and specification authorship rather than exclusive host: Foundation as authority, network as infrastructure. This separation is intended to let the library persist independently while preserving verifiable provenance.

The architecture would itself be a defense mechanism. Enclosure attempts would have to compromise enough of the network to render a verified version inaccessible, which is much harder than compromising a single server.

Phase 1 domains

A possible initial rollout would cover public-domain material that is already digitized and available in formats that can be processed without negotiation. Phase 1 would aim to produce a Level 3 provenance base; it has not begun.

Encyclopedia. Automated ingestion, verification, and atomization of encyclopedic knowledge. Source material could include existing open encyclopedic content, university and institutional publications, and verified public-domain reference works. The processing framework would extract verifiable claims, attribute them to sources, cross-reference related entries, and publish the result as a retrieval-optimized dataset. Unlike the source encyclopedias, the intended output would be atomic, cross-referenced, provenance-weighted knowledge units designed for retrieval rather than articles designed for reading.

Source library. This domain would cover public-domain texts, government documents, primary sources, and historical records. It would preserve full text while producing atomized, indexable components through structural analysis, metadata extraction, and cross-referencing.

Dictionary / lexicon. This domain would cover controlled vocabularies, definitions, terminology, and etymologies. It would provide consistent definitions, standardized terminology, and cross-domain disambiguation, particularly for prompts, framework specifications, and retrieval.

Phase 2 — News and current events

Phase 2 would apply the proposed library pipeline to current events: the same provenance hierarchy, document processing, and atomic-note output, but running on a continuous feed rather than archival content. If built, a continuously updated, provenance-verified feed formatted for retrieval-augmented generation could narrow the AI training-cutoff gap.

A possible initial scope is US national news. This domain could support political writing under the founder’s pen name and test the methodology in a high-stakes domain. Geographic and topical expansion would follow only after a US national pipeline and specification were proven.

Journalistic standards would apply: source verification, multi-source corroboration requirements, and provenance tracking on all claims. Under the proposed adversarial pipeline, two independent agents would evaluate source reliability before a story entered the knowledge base.

Main Street Independent is an existing publication informed by the same broad methodology. It is not currently a Foundation-operated library program, and its output does not currently flow through the proposed Phase 2 pipeline. A future relationship would require an explicit implementation and accurate public description.

Phase 3 and beyond

These domains could follow a working news service, driven by educational interests, public institutional data, and eventually the harder cases.

Textbooks. This domain could cover open educational resources — textbooks, instructional materials, and how-to guides. Its processing framework would need to preserve pedagogical structure rather than merely extract facts.

Courses. This domain could cover structured learning paths, curricula, and assessment frameworks while preserving prerequisite relationships, learning objectives, and assessment criteria.

Government data. This domain could cover Federal Reserve papers, census data, regulatory filings, public economic data, congressional records, and agency reports. Their varied formats and update cycles would require sub-specifications by data type.

Deferred domains

The proposal recognizes some domains as important but defers them because they present challenges beyond the standard pipeline.

Legal databases. Case law, statutes, regulations, legal commentary. Deferred because of jurisdictional complexity (federal, fifty states, municipal, international), copyright on legal commentary and annotations (the law itself is public domain; most useful compilations are not), and the specialized legal citation system. High value when built — legal research is expensive and access is inequitable.

Medical literature. Clinical research, treatment protocols, drug information, public health data. Deferred because of liability concerns, regulatory complexity (FDA, HIPAA implications for certain data types), and specialized verification requirements (peer review status, retraction tracking, conflict-of-interest disclosure). Any specification for this domain would require medical professionals on its future committee.

Museum and cultural heritage collections. Art, artifacts, cultural objects, archival collections. Deferred because of rights management complexity — physical objects may be in the public domain while photographs of them are copyrighted; institutional access agreements vary widely; metadata standards differ across institutions.

Patent databases. Patent filings, claims, prosecution histories, prior art. Deferred because of specialized formatting (patent claims have a specific legal syntax that affects processing), the distinction between granted patents and applications, and the international scope. Valuable for prior-art defense work supporting the public-domain defense mission.

Literature and film. Fiction, poetry, drama, screenwriting, film, television, and other narrative and artistic works. This domain is categorically different from the other domains in the proposal and is deferred for reasons fundamentally distinct from the regulatory and rights concerns that defer the domains above.

Every other domain in this plan is oriented around extracting verifiable knowledge. The processing pattern is: find source, verify provenance, atomize into facts, cross-reference, index. Fiction does not work this way. A novel’s value is not its facts. The thing that makes a great novel important is not reducible to atomic notes about its plot points — it is the experience of reading it, the way it restructures how the reader thinks. That is not knowledge extraction. That is cultural transmission. The processing framework for literature and film would need to operate on different principles entirely. What you extract from fiction is not facts — it is structure: themes, narrative architecture, character relationships, rhetorical strategies, intertextual connections. The atomic unit of fiction is not a verifiable claim — it is an interpretive lens.

This domain is deferred not because it is unimportant but because the processing framework itself has not been designed. It is a genuinely different intellectual problem, not just the same pipeline pointed at harder sources.

What the Knowledge Library is for

The proposed Knowledge Library is not a search engine. It would be a substrate for retrieval-augmented generation when a user’s personal vault does not already have the answer. The user’s first source would remain their own work; a future library could supply additional, provenance-weighted context.

The library is also conceived as defensive infrastructure. As cognitive automation becomes more integral to how people think, work, and communicate, the question of who controls its source material becomes load-bearing. A free public-domain knowledge library hosted on decentralized infrastructure could reduce dependence on vendor licensing, search ranking, or a content provider’s continued cooperation.

The library is a long-arc proposal. It would not need to be complete before becoming useful: Phase 1 could establish an encyclopedia, source library, and dictionary; Phase 2 could narrow the training-cutoff gap; later phases could expand coverage as committees form and specifications mature. The architecture is intended to support that compounding without requiring the Foundation’s organizational continuity.

The summary

The Knowledge Library is a proposal to scale the personal-vault commitment to a civilizational substrate. Its four-layer model would pair subject-matter expertise with faithful automated execution, and decentralized hosting would be intended to preserve the corpus beyond any single organization. A phased rollout could begin with a useful public-domain base and expand as committees form and specifications mature. Under the design, the Foundation would be steward rather than owner: it would not hold copyright in the corpus, and anyone could mirror, fork, or build alternatives. None of the library, signing, specification, committee, or distribution machinery described in this paper currently operates.