📎

RAG — Giving the Model Reference Materials Without Retraining

Retrieval + Generation — attach external knowledge to prompts without changing model weights

Fine-tuning vs RAG — Fundamental Difference

Fine-tuning: changes model weights (changes what's in its head)
RAG:         doesn't change weights (just gives reference materials)

Analogy: Fine-tuning = teaching a doctor a new disease. RAG = handing the doctor a patient chart.

How It Works

User asks question → retrieve relevant docs from DB → attach to prompt → model generates answer using docs.

Model weights change: zero.

Simplest RAG Code

qa = pipeline("question-answering")
context = "Annual leave: 15 days. Half-day: available."
result = qa(question="How many days off?", context=context)
# → {'answer': '15 days', 'score': 0.98}

That's RAG. No training. Just passed reference material along with the question.

Fine-tuning vs LoRA vs RAG

Fine-tuning LoRA RAG
Weight changes All Partial None
Training needed Yes Yes No
Data update Retrain Retrain Just swap docs

Key Concepts

1

User asks a question

2

Retrieve relevant docs from DB — vector similarity based

3

Attach retrieved docs to prompt — "answer based on these references"

4

model(prompt + references) → generate answer (Generation)

5

Model weights changed: zero — data update = just swap documents

Use Cases

Internal docs QA — build QA bot from company manuals and policies Fresh info — update docs without retraining for up-to-date answers Hallucination prevention — reduce false generation with source-backed answers