Skip to main content
Knowledge base + AI

AI knowledge base chatbot: answers from your documents, with sources

An AI knowledge base chatbot is an assistant that answers questions from your company’s own documents (the manuals, policies, past quotes, tickets and files spread across SharePoint, Google Drive, PDFs and a few people’s heads) and shows where each answer came from. It is built for small and mid-sized companies where new staff and customers ask the same questions every day, and the answer exists somewhere but nobody can find it fast. The technique behind it is retrieval-augmented generation: your documents are indexed, the relevant passages are retrieved for each question, and a language model writes the answer from those passages with citations. No model is trained on your data. I build these as an independent AI systems engineer (Sr DevOps, ten years of production infrastructure) on your own Postgres with pgvector and your own cloud, with per-user permissions, scheduled syncs and an evaluation set that measures whether the answers are right. Three weeks, a fixed price agreed up front, a 60-day operational guarantee.

What does "chat with your documents" actually mean?

The system reads your documents once, cuts them into passages, and stores each passage with a numeric fingerprint of its meaning (an embedding) in a database. When someone asks a question, it finds the passages closest in meaning, hands them to a language model together with the question, and the model writes an answer using only what it was handed. Each answer links back to the passages it used, so the reader can check. That is retrieval-augmented generation, RAG, and it is what every serious "chat with our documents" product does under the hood.

Two consequences follow. First, nothing is trained: your documents are not fed into a model that then remembers them. They sit in a database you control and are looked up at question time, so when you update a document the next answer reflects it. Second, answer quality is mostly retrieval quality, not model cleverness. If the right passage is not found, the best model available will guess or say it does not know. Most of the engineering in these systems goes into finding the right passage.

  • Index: documents are split into passages and stored with embeddings in Postgres with pgvector
  • Retrieve: for each question, the closest passages are found, filtered by what the user may see
  • Answer: the model writes from those passages only and cites them
  • No training: your data stays in your database and can be updated or deleted at any time

Why is a custom GPT or an off-the-shelf chatbot not enough for internal data?

A custom GPT, a Claude Project or a plug-and-play chatbot widget is a fine way to test whether the idea is useful: upload twenty PDFs, ask questions, see what happens. They fall short exactly where internal knowledge gets sensitive. One set of documents for everyone, so the salary policy is visible to whoever has the link. Stale the moment someone edits the source file, because the upload was a copy. Rarely a citation. And the documents live on a vendor’s servers under the vendor’s terms, which your IT lead, your lawyer or your largest customer may not accept.

A built knowledge base assistant fixes each of those by design: access filtered per user before the model sees anything, sources synced on a schedule so the assistant reads what the folder contains today, a citation on every answer, and the index, logs and conversations in your Postgres, in your cloud account, under your API keys.

  • Permissions: one assistant, different answers depending on who is asking
  • Freshness: connectors re-sync on a schedule instead of relying on a manual re-upload
  • Sources: every answer cites the document, page or ticket it came from
  • Hosting: index and logs in your database and your cloud, not a third party’s

What goes into the knowledge base, and how does it stay fresh?

The knowledge base is the real deliverable, and the first week of any build goes into deciding what belongs in it. The usual sources: the shared drive or SharePoint library with manuals and procedures, the policy PDFs, the catalog or price list, past quotes and proposals, the helpdesk history, and the FAQ your best person answers from memory. Not everything should go in: drafts, superseded versions and personal folders add noise and produce contradictory answers, so the scope of folders and record types is agreed explicitly.

Freshness is a connector plus a schedule. Each source gets a connector (Google Drive and Workspace APIs, Microsoft Graph for SharePoint and OneDrive, the helpdesk or CRM API, a database) that pulls new and changed items hourly or nightly and removes what was deleted. Each document keeps its version and last-modified date in the index, so the assistant prefers the newest procedure over the one from 2021 and tells the reader which version it used.

  • Agreed source list: which folders, libraries and record types are in and which are out
  • Connectors per source, with incremental sync so re-indexing stays cheap
  • Deleted at the source means deleted in the index on the next sync
  • A manual "re-sync now" for the day a policy changes and cannot wait until night

Who can see what? Permissions, PII and where the data lives

Permissions are enforced at retrieval, not by asking the model to keep secrets. Every passage carries the access rules of its source document (folder sharing settings, SharePoint groups, directory roles), and the search runs only over what the asking user is allowed to read, so the model never sees a restricted paragraph. A new employee sees the onboarding manual. A finance manager also sees the pricing rules. A customer on the public widget sees only what you chose to publish.

Personal data gets its own treatment. Ticket histories and past quotes contain names, emails and phone numbers, and the safest default is to mask them at indexing time unless the use case needs them. Conversations are logged for quality review with a retention period you set. And all of it, index, logs and sync state, runs in your cloud account (Supabase, GCP, AWS or the Railway project you own), so "where is our data" is answered in one sentence.

  • Per-user filtering before retrieval, inherited from the source’s sharing settings
  • PII masking at indexing for tickets, quotes and CRM records
  • Conversation logs with a retention period you choose
  • Index, logs and secrets in your own accounts; model calls under your own API keys

Where does the assistant live: web widget, Slack, WhatsApp or the helpdesk?

The knowledge base is one service with an API; the places people talk to it are thin clients on top. For internal use, Slack (where new staff already ask) and a small web page behind your login are the most common. For customers, a widget on the site or a WhatsApp number through Meta’s Cloud API or a business solution provider. For support teams, the assistant sits inside the helpdesk and drafts a reply for the agent to approve, using the same knowledge base plus the customer’s ticket history.

Start with one channel. The channel is the cheap part; the knowledge base, the permissions and the evaluation set are the expensive part, and they are shared. Adding a second channel later is a few days of work, not a new project.

How do you know the answers are right?

Before the assistant reaches anyone, it is run against an evaluation set: fifty to a few hundred real questions, each with the answer a knowledgeable person would give and the document it should cite. Every change to prompts, chunking or retrieval is re-run against that set, so improvements are measured rather than felt, and the set grows as new questions arrive.

Three behaviors are designed in, not hoped for. The assistant says "I do not have that in the documents" when retrieval comes back weak, instead of producing a fluent guess; a confident wrong answer about a refund policy or a safety procedure costs more than no answer. Every answer carries citations. And every answer has a thumbs up or down that lands in a review queue: recurring misses usually point to a missing or badly written document, and fixing the source improves every future answer.

  • Evaluation set of real questions with expected answers and expected sources
  • Explicit "I do not know" when retrieval is weak, with a path to a person
  • Feedback loop: thumbs up or down, review queue, fix the source and not the symptom
  • Weekly report: questions asked, answered, unanswered, and the top gaps

What does a build look like, step by step?

Week one is scope and sources. We agree the source list and the permission model, connect the first connectors (say, the SharePoint library and the Drive folders where the manuals live, plus two years of resolved tickets) and collect the evaluation questions. Documents are parsed (PDF, Word, spreadsheets, HTML), split into passages sized so one answer fits in one or two, embedded, and stored in Postgres with pgvector inside your Supabase or cloud project, with access rules and version metadata attached.

Week two is the retrieval service and the first channel. A Python and FastAPI service exposes one endpoint: question and user identity in, answer with citations out. It runs the permission-filtered search, calls the model (Claude, OpenAI or Gemini), applies the guardrails (answer only from retrieved passages, refuse when retrieval is weak) and logs everything. The first channel is wired to it. Sync jobs run on a schedule, in n8n or as a cron, with retries and alerts when a connector fails.

Week three is evaluation and handoff. The assistant is run against the evaluation set, retrieval and prompts are tuned until the remaining misses are gaps in the documents, and a pilot group uses it with the feedback loop on. Handoff includes the code and workflows in your repository, documentation, Loom walkthroughs and a runbook for adding a source or a channel. You own the index, the service, the prompts and the data; nothing runs in my accounts.

  • Trigger: a question in Slack, the web widget, WhatsApp or the helpdesk
  • Data: your documents and records, synced on a schedule into Postgres with pgvector
  • Model: Claude, OpenAI or Gemini via API, under your keys
  • Guardrails: permission filtering, cite or refuse, PII masking, logging, versioned prompts
  • Human in the loop: customer-facing drafts approved by a person until trusted
  • Runs in: your Supabase or cloud account, on Railway or your Kubernetes, with backups and alerts

What does an AI knowledge base chatbot cost, and when is a wiki the better answer?

The reference build is the AI Intake Agent plus knowledge base, published on this site at USD 4,900 as a single fixed price, with an optional USD 1,200 a month for operation (connectors, feedback queue, new sources and channels). An internal assistant over a few document sources with one channel fits that scope. Larger scopes (several systems, strict compliance, a customer channel plus an internal one) are quoted as a fixed price after a discovery call, with the productized prices as reference points. Model usage and hosting are billed to you directly, in your name, with a cap declared in the proposal. The 60-day operational guarantee applies after handoff.

Sometimes the honest answer is that you do not need this yet. If your team is under ten people and the questions fit in twenty pages, a well-organized wiki with a good search box, kept current by one owner, beats an assistant: cheaper, and nothing to evaluate. The assistant earns its place when knowledge is spread across hundreds of documents in several systems, when the same questions arrive daily from staff or customers, when access rules differ by role, or when an answer has to be assembled from several sources. If your case is the first one, I will say so on the first call.

Frequently asked questions

Does the model get trained on our documents?

No. Your documents are indexed in a database you own and looked up at question time; the model only sees the passages retrieved for each question. Nothing is fine-tuned, and deleting a document from the index removes it from every future answer. Model calls run under your own API keys, so the vendor terms that apply are the ones you signed.

Can it answer from SharePoint and Google Drive at the same time?

Yes. Each source gets its own connector and lands in the same index, so one question can be answered from a SharePoint procedure and a Google Drive spreadsheet together, with both cited. Helpdesk tickets, a CRM and a database can be added the same way, each with its own sync schedule.

What happens when a document changes?

The connector picks up the change on the next scheduled sync (hourly or nightly, as agreed) and re-indexes only that document, keeping its version and last-modified date. There is also a manual re-sync for urgent changes. No retraining, no re-upload.

Can customers use it, or is it internal only?

Both, from the same knowledge base. Internal users get permission-filtered answers in Slack or a page behind your login; customers get a web widget or WhatsApp that only sees what you decided to publish. Customer-facing replies can run behind a human approval step until you trust them unattended.

Which model do you use, and can we switch later?

Claude, OpenAI or Gemini, chosen for the job on the first call and called through your own API keys. The retrieval layer, the permissions and the evaluation set do not depend on the model, so switching later is a configuration change that is re-run against the evaluation set, not a rebuild.

When should we not hire you for this?

If the knowledge fits in twenty pages and a wiki with one owner would solve it, I will say so and not sell you a build. I also do not take on a knowledge base when the documents do not exist yet; writing them is your team’s work first. And I am one senior engineer, not an agency, so a scope that needs several parallel workstreams gets a second engineer or a different provider.

Who owns the system after handoff?

You do. The index and logs are in your database, the service runs in your cloud account, the code and workflows are in your repository, and the API keys are yours. Handoff includes documentation and Loom walkthroughs, and the 60-day operational guarantee covers anything I built that breaks in that window.

Bring the questions your team answers every day. 15 minutes.

A free call. I'll tell you straight which processes today's AI can solve and what your infrastructure needs for them to actually run. If it fits, we move forward. If not, I point you the right way, free.