Use OKF with RAG by retrieving relevant concept files and passing their text and source paths to a language model. Start with the local example below, then choose retrieval by bundle size and test it against questions from your own users.
Run an OKF retrieval example locally
This Python 3.9+ demo searches the downloadable sample bundle using shared words, selects up to three concepts, and prints an evidence prompt. It needs no packages, account, or API key. It demonstrates the retrieval step of RAG; it does not call a language model or claim benchmark performance.
The output contains concept paths, matched-term counts, and full Markdown bodies. For this query, inspect the revenue concepts and their policy links. The files use fictional business data. Frontmatter is searched as text; the demo does not validate trust claims, apply freshness rules, follow links, or execute computations.
Generate an answer from the evidence
Paste evidence-prompt.txt into your chosen language model and ask it to explain recognized revenue using only those sources. A supported answer should describe delivered orders whose return window has closed and cite the relevant concept path. Check the answer against the files yourself. No generated answer or quality score is supplied by the script.
If a result references another file needed to answer the question, inspect that file and add it to the evidence. Do not treat a keyword match as proof that the question can be answered. With private bundles, use only a model you are authorized to send that material to.
Expected output: “No matching concepts. Do not infer an answer from missing evidence.” This checks the empty-result path; it does not measure whether the retriever can detect every unanswerable natural-language question.
Measure retrieval before changing the index
Use the retrieval benchmark kit to record results from your own retriever. Its included example-results file is a scoring fixture, not measured output from this demo. Evaluate retrieval separately from the correctness and citations of generated answers.
Start with the smallest system that answers the query.
Small bundle
Traverse first.
Read index.md, follow direct links, then load the selected concept. Filename search and ordinary text search are usually enough.
Good default: repository files
Growing bundle
Filter, then search.
Scan frontmatter into a compact catalog. Filter by type, status, freshness, or tags before lexical search. Preserve links and backlinks for neighborhood expansion.
Good default: metadata plus full text
Large or mixed corpus
Build an external index.
Use lexical, vector, or hybrid retrieval based on measured query quality. Keep each result tied to its source concept and use graph links for a second retrieval hop.
Good default: benchmark two candidates
Governed serving
Keep the bundle portable when the serving layer grows.
Google Cloud’s official Knowledge Catalog example publishes OKF concepts as searchable catalog entries, maps v0.2 signal fields to structured aspects, and applies IAM to agent retrieval. That can solve organization-wide discovery and governance without changing the OKF files or adding new conformance requirements.
Treat the catalog as a synchronized serving layer: keep the portable bundle as the reviewable source, publish updates through CI, and retrieve full concept bodies and structured trust signals through the catalog APIs.
Do not load the whole bundle before you know which concepts matter.
01
Read the index
Use index.md as the bundle map and identify likely concept paths.
02
Inspect metadata
Check type, description, status, stale_after, sources, and verification before trusting a candidate.
03
Load selected content
Read the relevant concept or section, then follow only the links needed to resolve the question.
04
Return evidence
Keep the concept path and source IDs with the answer so another consumer can inspect the same trail.
Ranking policy
Trust signals should affect rank, not erase history.
Prefer stable, current, human-reviewed concepts when relevance is otherwise equal. Mark stale or deprecated results clearly and lower their rank when a current replacement exists.
Do not silently delete an unverified or deprecated concept from retrieval. It may explain an old decision, provide a migration trail, or be the only relevant evidence. Let the consumer see why it ranked lower.
Measure it
A benchmark you can reproduce.
1. Write 20 to 50 real questions. Include exact-name lookup, multi-hop questions, stale-policy traps, and questions with no answer.
2. Label acceptable concepts. Record which files contain enough evidence, not a single preferred wording.
3. Compare candidates. Measure recall at a fixed result count, evidence quality, latency, and tokens loaded for lexical, vector, and hybrid approaches.
4. Test lifecycle behavior. Confirm that current material outranks stale or deprecated material without making historical knowledge undiscoverable.
5. Publish the corpus and settings. A retrieval claim has little value without the bundle, questions, index settings, and scoring method.
No universal threshold makes embeddings necessary or unnecessary. Add complexity when your own evaluation shows a meaningful improvement.