Working With Your Own Documents and Data
The model does not know your policies, contracts, notes or database. Learn the three ways to give it that knowledge — paste it, upload it, or retrieve it — and when each one stops working.
The idea
Out of the box the model knows the public internet up to its cutoff. To make it useful on your material you have to put that material in front of it. Level one is pasting: fine for a page or two. Level two is uploading files into a workspace — ChatGPT Projects, Claude Projects, Gemini Gems, and Google’s NotebookLM all let you attach documents that every conversation can draw on. The model reads the relevant parts each time you ask.
Level three is retrieval, and it is what powers every "chat with our knowledge base" product. When your documents are too big to fit in the context window — thousands of pages, a whole wiki — a system first finds the handful of passages most relevant to the question (using embeddings, a way of measuring meaning-similarity between texts), then gives only those to the model along with the question. This is retrieval-augmented generation, RAG. The quality of the answer depends far more on whether the right passage was found than on the model.
The rule of thumb for everyone: if the answer to "did the model see the right text?" is no, no amount of prompting fixes it. For non-developers that means check which files are attached and ask the tool to quote its sources. For developers it means most RAG work is chunking, search quality and evaluation — not the LLM call.
Vocabulary
- Grounding
- Making the model answer from supplied sources rather than memory. Ask it to cite them.
- Embedding
- A list of numbers representing the meaning of a text, so similar texts can be found by nearest-neighbour search.
- RAG
- Retrieval-augmented generation: find relevant passages first, then have the model answer using them.
- Chunk
- A piece of a document (a paragraph, a section) stored and retrieved as a unit.
- Vector database
- A store optimised for embedding search (pgvector, Pinecone, Vertex AI Vector Search, and many more).
No-code walkthrough · non-developers
Build a personal expert on a 200-page document with NotebookLM or a Project
For: HR, legal, finance, teachers, researchers — anyone who lives in long documents
- 1Pick a document set you consult often: an employee handbook, a course syllabus and readings, a supplier contract, a product manual.
- 2Google NotebookLM (free): create a notebook, add the PDFs / Docs / URLs as sources. It only answers from those sources and shows inline citations — click one to see the exact passage. Try: "What is the notice period for a manager leaving, and where does it say that?"
- 3Alternative: ChatGPT Project or Claude Project — create one, upload the files, write instructions such as "Answer only from the attached documents. If the answer is not there, say so." Gemini: create a Gem with the files attached.
- 4Ask ten real questions you have actually needed answers to. For each, open the cited passage and confirm. Note any answer with no citation — that is the model guessing.
- 5Ask for a study guide, an FAQ, or a checklist derived from the sources — this is where these tools shine for onboarding and training material.
Using only the attached documents, answer the question below. Quote the exact sentence(s) you relied on and name the document and section. If the documents do not answer it, say "Not covered in the sources" — do not guess. Question: [your question]
Check your work: The test is citations. A grounded tool shows you where the answer came from; if you cannot click through to the passage, treat the answer as a draft.
Developer walkthrough · Gemini · Claude · OpenAI
A minimal RAG pipeline is under 60 lines: chunk documents, embed them, store vectors, embed the question, take the top matches, and prompt the model with them. Below uses each provider’s own embedding model and an in-memory store; swap in pgvector or a managed vector DB for production. Each provider also offers a managed version (Gemini File Search, Claude Files + citations, OpenAI file_search) that does chunking and retrieval for you.
import numpy as np
import voyageai # Anthropic's recommended embeddings partner
from anthropic import Anthropic
vo = voyageai.Client() # export VOYAGE_API_KEY=...
client = Anthropic()
def embed(texts: list[str]) -> np.ndarray:
return np.array(vo.embed(texts, model="voyage-3.5").embeddings)
# 1. Chunk (by paragraph) and embed once
chunks = [c.strip() for c in open("handbook.md").read().split("\n\n") if c.strip()]
index = embed(chunks)
def answer(question: str) -> str:
# 2. Retrieve top-3 chunks by cosine similarity
q = embed([question])[0]
sims = index @ q / (np.linalg.norm(index, axis=1) * np.linalg.norm(q))
top = [chunks[i] for i in sims.argsort()[-3:][::-1]]
context = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(top))
# 3. Generate, grounded on the retrieved chunks
resp = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=600,
system="Answer only from the sources. Cite as [n]. If not covered, say so.",
messages=[{"role": "user",
"content": f"Sources:\n{context}\n\nQuestion: {question}"}],
)
return resp.content[0].text
print(answer("What is the notice period for managers?"))
# Managed alternative: pass documents as content blocks with
# "citations": {"enabled": True} and Claude returns exact source spans.- 1Write 20 question/answer pairs from the document by hand. Measure how often the correct chunk is in the top-3 (retrieval recall) — separately from whether the final answer is right.
- 2Change chunk size (paragraph → 500 tokens with overlap → section) and re-measure. Retrieval usually moves more than the model does.
- 3Add a keyword (BM25) search alongside embeddings and merge results. Hybrid search is the production default.
Practice
Everyone
Build a NotebookLM or Project on a document set from your job. Ask it the ten questions colleagues most often ask you. Share it with one colleague and see whether it answers their questions correctly.
Developers
Extend the pipeline to a folder of PDFs with page-level citations. Add a "not covered" path and an eval of 30 questions with retrieval recall and answer accuracy reported separately.
Pitfalls at this stage
- •Uploading and assuming. Always ask for the source passage; the tool may have skimmed or truncated a long file.
- •For developers: evaluating end-to-end only. When the answer is wrong, you need to know whether retrieval or generation failed.
- •Stale documents. A knowledge base is only as good as its last update — build refresh into the process.
Free resources
- Google NotebookLM ↗Free, source-grounded, the best no-code way to feel RAG.
- DeepLearning.AI: Building and Evaluating Advanced RAG ↗Short course with evaluation focus.
- Anthropic: Contextual Retrieval ↗A simple technique that materially improves retrieval.
- OpenAI Cookbook: RAG examples ↗Runnable notebooks for retrieval, embeddings, evals.