OKF Builder Reference / Open Knowledge Format v0.2, explained for builders OKF v0.2
OKF OpenKnowledgeFormat

Content updated

Implementation guide · OKF v0.2

How to use OKF with RAG: a working example.

Use OKF with RAG by retrieving relevant concept files and passing their text and source paths to a language model. Start with the local example below, then choose retrieval by bundle size and test it against questions from your own users.

Run an OKF retrieval example locally

This Python 3.9+ demo searches the downloadable sample bundle using shared words, selects up to three concepts, and prints an evidence prompt. It needs no packages, account, or API key. It demonstrates the retrieval step of RAG; it does not call a language model or claim benchmark performance.

Download the reference bundle ZIP and Python retrieval script into a new folder. Inspect the script, then run:

Run in your download folder

shell

unzip okf-v0-2-reference-bundle.zip
python3 okf-retrieve.py okf-v0.2-reference "recognized revenue" > evidence-prompt.txt
cat evidence-prompt.txt

Inspect the retrieved evidence

The output contains concept paths, matched-term counts, and full Markdown bodies. For this query, inspect the revenue concepts and their policy links. The files use fictional business data. Frontmatter is searched as text; the demo does not validate trust claims, apply freshness rules, follow links, or execute computations.

Generate an answer from the evidence

Paste evidence-prompt.txt into your chosen language model and ask it to explain recognized revenue using only those sources. A supported answer should describe delivered orders whose return window has closed and cite the relevant concept path. Check the answer against the files yourself. No generated answer or quality score is supplied by the script.

If a result references another file needed to answer the question, inspect that file and add it to the evidence. Do not treat a keyword match as proof that the question can be answered. With private bundles, use only a model you are authorized to send that material to.

Check a query with no matching evidence

No-match check

shell

python3 okf-retrieve.py okf-v0.2-reference zxqnonexistent

Expected output: “No matching concepts. Do not infer an answer from missing evidence.” This checks the empty-result path; it does not measure whether the retriever can detect every unanswerable natural-language question.

Measure retrieval before changing the index

Use the retrieval benchmark kit to record results from your own retriever. Its included example-results file is a scoring fixture, not measured output from this demo. Evaluate retrieval separately from the correctness and citations of generated answers.

Read the OKF vs RAG comparison for the conceptual distinction, or start with a minimal OKF example.

Decision guide

Start with the smallest system that answers the query.

Small bundle

Traverse first.

Read index.md, follow direct links, then load the selected concept. Filename search and ordinary text search are usually enough.

Good default: repository files

Growing bundle

Filter, then search.

Scan frontmatter into a compact catalog. Filter by type, status, freshness, or tags before lexical search. Preserve links and backlinks for neighborhood expansion.

Good default: metadata plus full text

Large or mixed corpus

Build an external index.

Use lexical, vector, or hybrid retrieval based on measured query quality. Keep each result tied to its source concept and use graph links for a second retrieval hop.

Good default: benchmark two candidates

Governed serving

Keep the bundle portable when the serving layer grows.

Google Cloud’s official Knowledge Catalog example publishes OKF concepts as searchable catalog entries, maps v0.2 signal fields to structured aspects, and applies IAM to agent retrieval. That can solve organization-wide discovery and governance without changing the OKF files or adding new conformance requirements.

Treat the catalog as a synchronized serving layer: keep the portable bundle as the reviewable source, publish updates through CI, and retrieve full concept bodies and structured trust signals through the catalog APIs.

Agent reading order

Spend context in stages.

Do not load the whole bundle before you know which concepts matter.

  1. 01

    Read the index

    Use index.md as the bundle map and identify likely concept paths.

  2. 02

    Inspect metadata

    Check type, description, status, stale_after, sources, and verification before trusting a candidate.

  3. 03

    Load selected content

    Read the relevant concept or section, then follow only the links needed to resolve the question.

  4. 04

    Return evidence

    Keep the concept path and source IDs with the answer so another consumer can inspect the same trail.

Ranking policy

Trust signals should affect rank, not erase history.

Prefer stable, current, human-reviewed concepts when relevance is otherwise equal. Mark stale or deprecated results clearly and lower their rank when a current replacement exists.

Do not silently delete an unverified or deprecated concept from retrieval. It may explain an old decision, provide a migration trail, or be the only relevant evidence. Let the consumer see why it ranked lower.

Measure it

A benchmark you can reproduce.

  1. 1. Write 20 to 50 real questions. Include exact-name lookup, multi-hop questions, stale-policy traps, and questions with no answer.
  2. 2. Label acceptable concepts. Record which files contain enough evidence, not a single preferred wording.
  3. 3. Compare candidates. Measure recall at a fixed result count, evidence quality, latency, and tokens loaded for lexical, vector, and hybrid approaches.
  4. 4. Test lifecycle behavior. Confirm that current material outranks stale or deprecated material without making historical knowledge undiscoverable.
  5. 5. Publish the corpus and settings. A retrieval claim has little value without the bundle, questions, index settings, and scoring method.

No universal threshold makes embeddings necessary or unnecessary. Add complexity when your own evaluation shows a meaningful improvement.

Mathias Onea

Mathias Onea

Senior Engineer, Product Builder, and Founder

Systems, product software, and practical execution for teams that need clear decisions, durable implementation, and agent-ready knowledge structures.

Focus
Knowledge systems, Laravel platforms, automation, and technical SEO infrastructure.
Related work
Founder-led software work through Craftwell and independent open-source projects.
Profile
mathiasonea.com