Skip to main content
Skip to content
Generative AI

RAG Systems

An assistant that answers from your own documents and says where the answer came from. Built on the model you already use, inside your systems.

01 / 04

The model does not know your company

A general model was trained on the public web up to a date. Your policy, your prices, last month's decision are not in it, and asked about them it answers anyway.

CH.01 · The problem

A general model does not know your company.

  1. It stops at its training date

    The model was trained on the public web up to a point. Your prices, your policy, last month's decision: none of it is there.

  2. It invents when it does not know

    Asked about something it never saw, a model produces a fluent answer anyway. In a support queue or a contract review that is a liability.

  3. It cannot show its source

    Without a reference, nobody can check the answer. The person who reads it has to trust it or redo the work.

Our answer

Retrieve first, then answer

A RAG system looks up the relevant passages in your knowledge base before the model writes a word. The answer is grounded in your documents, current as of the last index, and carries its citations.

CH.02 · What we build

How we build it

Three parts, each one measurable on its own before the next is trusted.

Your documents, indexedRetrieval by meaningAnswers in contextCitations on every answer
01

Knowledge processing

Your material becomes a searchable base.

  • Ingest documents from the systems where they live
  • Split text into passages that keep their meaning
  • Attach metadata: owner, date, access level
  • Embed passages for retrieval by meaning
  • Index and schedule the refresh
02

Retrieval

The right passages, for the question asked.

  • Search by meaning and by keyword together
  • Rewrite vague questions before searching
  • Rank passages by relevance and recency
  • Filter by the asker's permissions
  • Measure recall on a held-out question set
03

Answer generation

A grounded answer with its sources attached.

  • Prompt the model with the retrieved passages only
  • Cite the document and section for each claim
  • Refuse when the base has no answer
  • Review a sample of answers each week
  • Feed corrections back into the base
CH.03 · How it runs

How it runs

Four phases, each with a deliverable you can test.

  1. 01Phase 1

    Discovery

    What people ask, where the answers live, who may see what.

    • Inventory the knowledge sources
    • Collect the questions people actually ask
    • Map access rules per source
  2. 02Phase 2

    Knowledge base

    Ingest, split, embed and index your documents.

    • Connect the source systems
    • Build the processing pipeline
    • Choose and configure the vector store
  3. 03Phase 3

    Retrieval

    Wire the search into the assistant your teams use.

    • Implement the query pipeline
    • Tune ranking on the test set
    • Add permission filtering
  4. 04Phase 4

    Production

    Ship, watch, and keep the base current.

    • Integrate with the assistant and the systems
    • Set up the refresh schedule
    • Instrument answer quality
CH.04 · What changes

What changes

Answers you can check, from the knowledge you already have.

Fewer invented answers

The model writes from your passages, so an answer with no source in the base is refused rather than guessed.

Less time searching

The person asking gets the passage and the answer together, without opening four systems.

Answers people trust

A citation on every response means the reader can verify it in one click, and does.

Side by sideModel aloneWith RAG
Where answers come fromTraining data, fixed at a dateYour documents, current at the last index
Company knowledgeNoneEverything you index, by permission
SourcesNoneDocument and section on every answer
When it does not knowGuessesSays so
ControlNone over what it knowsYou decide what is indexed and who sees it
CH.05 · Questions

Questions

Anything with text: PDFs, wikis, tickets, emails, database records, transcripts. Each source keeps its owner and access level, so the assistant only retrieves what the asker may see.

Sources are re-indexed on a schedule you set, from hourly to weekly, and on demand. An answer always states the date of the passage it used.

Yes, with permission filtering at retrieval time: a passage is only considered if the person asking may read the source. The model can run locally when the documents may not leave your perimeter.

Against a question set agreed in discovery, with the expected answers and sources. We report recall, citation accuracy and refusal rate, and review a weekly sample of live answers with your team.

More in Generative AI

The map of the practice

Back to Generative AI
Start

Bring us the problem nobody has cracked yet.

We are a small team of senior specialists. We pick the right model and the right layer, and we build the least machinery that does the job. You get a call with an engineer, not a sales deck.