Enterprise knowledge search with RAG (retrieval-augmented generation) lets staff ask a question in plain language and get an answer written from the company's own documents, with links to the passages it used. Done properly, it only searches content the person asking may already open, keeps its index current, and says so when it cannot find an answer.
Most organisations have the answers somewhere: in SharePoint libraries, shared drives, wikis, old tickets and policy PDFs. The problem is finding them. This guide is about the enterprise search use case specifically: connecting sources, respecting permissions, keeping answers accurate and rolling out department by department. For the wider engineering picture of RAG, agents and guardrails, see our guide to LLM app development.
What is enterprise knowledge search with RAG, in plain words?
It is a search box that answers instead of listing files. The system finds the most relevant passages in your documents, gives them to a language model, and the model writes an answer from those passages with citations.
It works in three stages:
- Indexing. Connectors read documents from your systems, split them into passages, record metadata and permissions, and store them in a search index.
- Retrieval. When someone asks a question, the system finds the passages most likely to contain the answer, filtered to what that person may see.
- Answering. A language model writes an answer using only those passages, and cites them.
Compared with keyword search, it handles questions phrased differently from the document wording and pulls an answer together from several places. Compared with asking a general chatbot, it answers from your content, and every claim can be checked against a cited source.
Which sources can it connect to?
Most document and knowledge systems can be connected through their APIs. What matters is whether a connector can read content, metadata and permissions, and detect changes.
| Source category | Examples | What the connector needs | Watch out for |
|---|---|---|---|
| Document libraries | SharePoint and OneDrive, Google Drive, file shares | Files, folder paths, owners, sharing permissions, change tracking | Inherited permissions, sharing links, old versions and duplicates |
| Wikis and intranets | Confluence, intranet pages | Pages, spaces, labels, space permissions and page restrictions | Outdated pages that nobody owns |
| Ticketing and helpdesk | IT and customer support systems | Resolved tickets, knowledge articles, categories | Personal data in ticket threads; unresolved or wrong answers |
| Business systems | ERP, CRM, HR systems | Usually live queries through an API rather than indexing | Numbers change constantly; field-level access rules |
| Email and chat | Mailboxes, team channels | Messages, participants, retention rules | Highly personal; usually a later phase, if at all |
Two examples of what the platforms expose. Microsoft Graph can list "the effective sharing permissions" on a SharePoint or OneDrive item, including permissions inherited from parent folders, and its delta query tracks changes over time, with options to flag permission changes. The Google Drive API lists a file's or shared drive's permissions. Confluence Cloud's REST API returns the restrictions on a piece of content, which work alongside space permissions.
For live business data, such as stock levels or a customer's balance, searching an index is the wrong tool. Connect the assistant to the system through an API so it reads the current value under the user's own access rights. Our systems and API integration page covers that work.
Why must retrieval respect source permissions at query time?
Because the index usually reads documents with a broad service account, it can see more than any single employee. Without permission filtering, the assistant will quote a restricted document to someone who could never open it.
That is a data leak. It is also hard to spot, because the answer looks helpful. The OWASP Top 10 for LLM Applications (2025) covers this under LLM08, Vector and Embedding Weaknesses, and recommends "permission-aware vector and embedding stores"; LLM02, Sensitive Information Disclosure, covers the wider risk.
| Approach | How it works | Strengths | Weaknesses |
|---|---|---|---|
| Store permissions in the index | Each passage carries the users and groups allowed to see it; every query is filtered by the asking user's identity and groups | Fast; works at scale | Only as current as the last permission sync |
| Check with the source at query time | After retrieval, the system confirms with the source that the user can open each document | Always current | Slower; depends on source APIs being available and fast enough |
| Hybrid | Filter in the index, then re-check the documents used in the answer | Balances speed and accuracy | More to build and test |
Rules worth adopting whichever approach you choose:
- Deny by default. If a document's permissions cannot be read, do not index it.
- Resolve groups properly, including nested groups, and keep group membership in sync.
- Sync permission changes quickly, and make removals as fast as additions.
- Decide how to treat broad sharing links, such as "anyone in the organisation with the link". Many organisations choose not to index them by default.
- Test with real accounts from different departments and seniority levels before launch.
How should documents be chunked and tagged?
Split documents along their own structure and attach metadata to every passage. Retrieval quality depends more on this step than on the choice of language model.
- Chunk by structure: headings, sections and numbered clauses, not fixed character counts that cut a sentence or table in half.
- Keep tables and lists intact, or convert them into a form that keeps rows together.
- Carry context into each passage: document title, section heading and path, so a passage still makes sense out of context.
- Tag metadata: source, owner, department, document type, status (draft or approved), language, effective date and last modified date.
- Use keyword and meaning-based search together. Part numbers, policy codes and names are matched better by keywords; loosely phrased questions are matched better by meaning. Combining both is a common pattern.
- Exclude noise: superseded versions, drafts, duplicates and templates, unless people need them.
How do you keep the index fresh?
Sync changes continuously rather than re-indexing everything on a schedule, and make deletions and permission changes travel as fast as new content.
- Use change-tracking features where the source offers them, such as Microsoft Graph delta queries for drives.
- Process deletions and permission removals first. A stale answer is annoying; an answer from a deleted or restricted document is a breach.
- Run a periodic full reconciliation as a safety net, comparing the index with the source.
- Show the document date in every citation, so users can judge freshness themselves.
- Give each source an owner who is told when content is flagged as wrong or outdated.
How should answers cite their sources?
Every answer should link to the specific passages it used, so the reader can check it in one click.
- Cite at the level of the passage or section, not just the file.
- Show the document title, date and source system beside each citation.
- When sources conflict, show both and say they disagree, rather than picking one silently.
- When nothing relevant is found, say so and suggest who owns the topic. "I could not find this in the documents you have access to" is a correct answer.
- Keep answers short and let the citations carry the detail.
How do you know the answers are accurate?
Build a set of real questions with known answers and known source documents for each department, and measure retrieval and answering separately.
- Retrieval: did the right passage appear among the results? If not, the problem is chunking, metadata or the connector, not the model.
- Answer correctness: is the answer right and complete?
- Groundedness: is every claim supported by a cited passage?
- Refusal: for questions the documents cannot answer, does it say so?
- Permissions: does a user from one department get nothing from another department's restricted content?
Our guide on how to evaluate an AI pilot explains how to build the test set, set thresholds in advance and run a go/no-go decision.
Where does your data go?
That depends on where the model and the index run. Decide this with your security team before choosing tools.
| Option | What it means | Questions to ask |
|---|---|---|
| Hosted model API | Your question and the retrieved passages are sent to a model provider's service | Are inputs used for training? How long are prompts and outputs retained? Which regions? Which subprocessors? |
| Model service in your own cloud account | A cloud provider's model offering inside your tenancy and chosen region | Do the cloud agreements you already have cover it? What logging is on by default? |
| Self-hosted model | An open-weight model running on infrastructure you control | Do you have the hardware and skills to run and update it? Is its quality sufficient for your test set? |
Whichever you choose, the search index is now a copy of your documents. Protect it like the source systems: encryption, access controls and backups. OWASP's LLM08 entry notes that embeddings can sometimes be reversed to recover source text, so treat the vector store as sensitive. Set retention for query logs, because the questions people ask can be sensitive in themselves, and decide who may read those logs. A well-designed system is designed to support your data protection obligations, but it does not remove them. Our security and data protection page explains how we approach this.
What can enterprise knowledge search not do?
- It cannot fix poor content. If two policies contradict each other, the assistant will surface the contradiction, not resolve it.
- It cannot answer from what is not indexed. Knowledge that lives only in people's heads stays there.
- It is not a reporting tool. Totals across thousands of records belong in reports and dashboards, not in an answer assembled from passages.
- It does not replace records management. Retention, legal holds and official records stay in the systems built for them.
- It can still be wrong. Citations make errors checkable; they do not make them impossible.
- It cannot enforce permissions the source does not record. If access is controlled informally ("everyone knows not to open that folder"), fix the permissions first.
How should you roll it out by department?
Start with one department whose content is in good shape and whose permissions are clear, prove it there, then expand one source and one department at a time.
| Phase | Scope | Exit criteria |
|---|---|---|
| 1. Content and permission audit | Pick the first department; list sources, owners and access rules; remove obvious clutter | Owners named; permissions understood and tidied |
| 2. Pilot | One or two sources, a small group of users, a department test question set | Test set thresholds met; permission tests passed |
| 3. Department launch | Whole department, feedback button on every answer | Usage and feedback reviewed; flagged content fixed by owners |
| 4. Next department | Repeat audit and pilot for the next content area | Same criteria, with its own test questions |
| 5. Cross-department search | Enable search across departments for users with access to both | Permission tests across departments passed |
Internal policies, IT help content and standard operating procedures are often good starting points because the content is written down and widely accessible. Departments with highly restricted content, such as HR case files or legal matters, usually come later, if at all.
When is RAG search not the right approach?
- A small, stable document set. A well-organised intranet with a good FAQ may serve better at lower cost.
- The answers need live transactional data. Integrate with the business system directly instead.
- Permissions are disorganised. Fix access first; a RAG system will expose every mistake.
- No error is acceptable. Where any wrong answer has serious consequences, keep a human expert in the loop or use structured decision tools.
Enterprise knowledge search checklist
- First department chosen; sources and owners listed
- Permissions audited; broad sharing links reviewed
- Connectors read content, metadata, permissions and changes
- Permission filtering on every query; deny by default
- Group membership, including nested groups, kept in sync
- Chunking follows document structure; tables kept intact
- Metadata includes status, owner, department and dates
- Change sync running; deletions and permission removals prioritised
- Citations at passage level, with dates
- "Not found" behaviour tested
- Department test question set built; thresholds agreed in advance
- Permission tests with real accounts passed
- Model hosting option, retention and logging agreed with security
- Feedback loop to content owners in place
Building knowledge search with Timeline Digital
We build knowledge search that respects the permissions you already have and cites every answer. Our knowledge base and AI search page describes the service, and our AI development page covers the wider work. Every project starts with a free pilot of 2 to 3 key modules before the full project, typically one department's sources with its own test questions.