Skip to content
Government knowledge bases

RAG and AI knowledge bases for government

A knowledge base that answers from your organisation's own documents, cites the passage it used and keeps restricted material restricted.

01 / How it works

What a RAG knowledge base is

Retrieval-augmented generation (RAG) pairs a search index over your documents with a language model. For each question it retrieves the passages that matter, then writes an answer from them and cites them. Staff get an answer in plain language and the document behind it in one click.

01. Your documents

Policies, procedures, reports and web pages are loaded into an index, along with their access rules.

02. Retrieval

Each question pulls the most relevant passages that the person asking is allowed to see.

03. A cited answer

The answer is written from those passages, with numbered citations that open at the source.

02 / What we build in

A knowledge base an agency can stand behind

Government knowledge bases get reviewed. These are the parts that let one hold up when someone asks where an answer came from and who could see it.

Permission-aware retrieval

Entitlements are resolved before retrieval, so a user's question only reaches documents that user is allowed to read.

Citations and confidence

Numbered citations open each source at the passage used, and each answer carries a confidence rating.

Audit log

Every question and retrieval is logged against the requesting user, ready for records and review.

Sign-in and roles

Sign-in through your organisation's single sign-on, with roles for readers, curators and administrators.

Loading at estate scale

CMS content, network drives, PDFs and scanned back catalogues. Getting content in cleanly is most of the work.

Measured answer quality

A scored set of real questions is run before launch and after changes, so quality is tracked over time.

03 / Common uses

Where a knowledge base earns its keep

Staff knowledge base

Policy, procedure and guidance that staff can question in plain language, with the source one click away.

Research portal

Published reports and research made searchable and answerable for a sector or the public. CorpusKit is built for this.

Website search and answers

Search and cited answers on your public website, drawn from content your CMS already publishes.

04 / Approach

From pilot to production

Four steps, each with something your team can review before the next one starts. We are building the Knowledge Sharing Platform for the Victorian Collaborative Centre for Mental Health and Wellbeing on Progress Agentic RAG.

01. Scope

Agree the collection, the users, the access rules and the obligations that apply to your organisation.

02. Load

Ingest and structure the documents, map permissions, and set up labels and topics.

03. Pilot

Put a working knowledge base in front of real users, measured against questions your team actually asks.

04. Run

Support, re-sync new material and keep measuring answer quality as the collection grows.

06 / Questions

What agencies ask about knowledge bases

It is a knowledge base that answers questions in plain language from your own documents. Retrieval-augmented generation (RAG) finds the passages relevant to a question, then a language model writes an answer from those passages and cites them.

A general chatbot answers from what its model learned in training. A RAG knowledge base answers from your documents, shows the passages it used, and can say when the documents do not cover a question.

Yes. We resolve each user's entitlements before retrieval, so the model only sees passages that user is allowed to read. Every retrieval is logged against the requesting identity.

PDFs, Office documents and web pages are the usual starting point, including scheduled re-syncs of your public website. Scanned documents need their text extracted first, and we handle that as part of loading.

Each claim carries a numbered citation that opens the source at the passage used. Answers also carry a confidence rating, and a thin answer says so.

Before the build starts we document where documents are stored, where retrieval and generation run, and anything that leaves Australia. On Progress Agentic RAG, knowledge boxes can be created in an Australian region.

Progress Agentic RAG where the requirement suits it, with licences supplied through BlueChip IT. Platform selection happens in discovery, against your assurance requirements.

Commonwealth bodies work to the Digital Transformation Agency's policy for the responsible use of AI in government and its technical standard. Victorian bodies work to the VPS generative AI guideline, the VPS AI Assurance Framework, OVIC privacy guidance and the PROV recordkeeping policy. We design against these from the start.

Yes. The CorpusKit demo at demo.corpuskit.org is a public knowledge base you can question now. For your own documents, we run a pilot over a real collection with real permissions.