CourseRAG · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer) · part 5 of 82
Part 5 · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer)

Topic 5: The Course Tech Stack

10 min read·21 Sept 2026

Everything in this course uses free or free-tier tools, so every learner can run every lab. Here's the full stack, and why each piece is in it.

5.1 The Stack at a Glance

Flowchart LR
LayerCourse defaultAlso coveredWhy the default
LLMGroq (llama-3.3-70b-versatile)Gemini (gemini-3.8-flash), Ollama (llama3.1), Hugging Face local modelsFree tier and very fast responses
Embedding modelHugging Face BAAI/bge-small-en-v1.5all-MiniLM-L6-v2, multilingual-e5-small, Ollama nomic-embed-text, Gemini gemini-embedding-001Free, local, strong retrieval quality for its size
RerankerHugging Face cross-encoder/ms-marco-MiniLM-L-6-v2BAAI/bge-reranker-base, BAAI/bge-reranker-v2-m3Small, fast, runs on CPU
Vector DBQdrant (in-memory for labs)FAISS, Chroma, pgvector, PineconeRuns locally with no setup, and the same code scales to a server
FrameworksPlain Python (to learn how things work)LangChain, LlamaIndexUnderstand the internals first, then use frameworks to go faster

Teaching approach: each concept is first built in plain Python so learners see what's happening, then shown in LangChain or LlamaIndex for real projects.

5.2 Project Setup

Folder structure (used for the whole course)

text
rag-course/
├── .env                  # API keys (never commit this)
├── requirements.txt
├── common/
│   ├── init.py
│   ├── llm.py            # one ask() function for every LLM provider
│   └── embeddings.py     # one Embedder class for every embedding provider
└── module01/
    └── lab.py

Install

python
python -m venv .venv && source .venv/bin/activate     # Windows: .venv\Scripts\activate\

pip install python-dotenv numpy 
    groq "google-genai>=2.3.0" langchain-ollama langchain-core 
    sentence-transformers qdrant-client

Extra packages for specific examples are listed next to each example.

**Local models (for Ollama options)**, after installing Ollama from ollama.com:

ollama pull llama3.1\
ollama pull nomic-embed-text

`.env`** file**

GROQ_API_KEY=your-groq-key          # console.groq.com\
GEMINI_API_KEY=your-gemini-key      # aistudio.google.com\
LLM_PROVIDER=groq                   # groq | gemini | ollama

5.3 LLMs

ProviderModelCostBest for
Groqllama-3.3-70b-versatileFree tier (rate-limited)Default for labs; very fast
Geminigemini-3.8-flashFree tier availableStrong quality; long context
Ollamallama3.1Free, runs on your computerOffline work and private data
Hugging Facee.g. Qwen/Qwen2.5-1.5B-InstructFree, runs on your computerLearning how models run locally

Example: Groq

python
import os\
from groq import Groq\

client = Groq(api_key=os.environ.get("GROQ_API_KEY"))\

chat_completion = client.chat.completions.create(\
    messages=[{"role": "user", "content": "Explain the importance of fast language models"}],\
    model="llama-3.3-70b-versatile",\
)
print(chat_completion.choices[0].message.content)

**Example: Gemini**

from google import genai\

client = genai.Client()   # reads GEMINI_API_KEY\

interaction = client.interactions.create(\
    model="gemini-3.8-flash",\
    input="Explain how AI works in a few words",\
)
print(interaction.output_text)

**Example: Ollama via LangChain**

from langchain_ollama import ChatOllama\

llm = ChatOllama(model="llama3.1", temperature=0)\
print(llm.invoke("Explain how AI works in a few words").content)

**Example: Hugging Face model running locally**

# pip install transformers torch
from transformers import pipelin

generator = pipeline("text-generation", model="Qwen/Qwen2.5-1.5B-Instruct")\
messages = [{"role": "user", "content": "Explain RAG in one sentence."}]\
output = generator(messages, max_new_tokens=100)\
print(output[0]["generated_text"][-1]["content"])

Small local models are great for learning but less reliable at following strict rules (citation formats, JSON-only output). Always validate their output in code.

The course helper: common/llm.py

Every lab calls one function, ask(). To switch providers, change LLM_PROVIDER in .env or pass provider=. The lesson code never changes.

# common/llm.py

python
import json
import os
import re
from functools import lru_cache

from dotenv import load_dotenv

load_dotenv()

DEFAULT_PROVIDER = os.getenv("LLM_PROVIDER", "groq")
MODELS = {
    "groq": "llama-3.3-70b-versatile",
    "gemini": "gemini-3.8-flash",
    "ollama": "llama3.1",
}


@lru_cache
def _groq():
    from groq import Groq
    return Groq(api_key=os.environ.get("GROQ_API_KEY"))


@lru_cache
def _gemini():
    from google import genai
    return genai.Client()


@lru_cache
def _ollama(json_mode: bool, temperature: float):
    from langchain_ollama import ChatOllama
    extra = {"format": "json"} if json_mode else {}
    return ChatOllama(model=MODELS["ollama"], temperature=temperature, **extra)


def ask(prompt: str,
        system: str = "You are a helpful assistant.",
        provider: str | None = None,
        json_mode: bool = False,
        temperature: float = 0.0) -> str:
    """Send one prompt to the chosen provider and return the reply text."""
    provider = provider or DEFAULT_PROVIDER

    if provider == "groq":
        # Groq's JSON mode requires the word "JSON" somewhere in the messages.
        extra = {"response_format": {"type": "json_object"}} if json_mode else {}
        resp = _groq().chat.completions.create(
            model=MODELS["groq"],
            messages=[{"role": "system", "content": system},
                      {"role": "user", "content": prompt}],
            temperature=temperature,
            **extra,
        )
        return resp.choices[0].message.content

    if provider == "gemini":
        interaction = _gemini().interactions.create(
            model=MODELS["gemini"],
            system_instruction=system,
            input=prompt,
            generation_config={"temperature": temperature},
        )
        return interaction.output_text

    if provider == "ollama":
        from langchain_core.messages import HumanMessage, SystemMessage
        reply = _ollama(json_mode, temperature).invoke(
            [SystemMessage(system), HumanMessage(prompt)])
        return reply.content

    raise ValueError(f"Unknown provider: {provider}")


def parse_json(text: str, default=None):
    """Pull the first JSON object out of a model reply. Returns default if none."""
    match = re.search(r"\{.*\}", text or "", re.S)
    try:
        return json.loads(match.group(0)) if match else default
    except json.JSONDecodeError:
        return default

Try it

python

from common.llm import ask

for provider in ["groq", "gemini", "ollama"]:
    print(provider, "→", ask("What is RAG, in one sentence?", provider=provider))

5.4 Embedding Models

An embedding model turns text into a list of numbers (a vector). Texts with similar meaning get similar vectors. That's what makes semantic search possible.

ModelSourceDimensionsRunsBest for
sentence-transformers/all-MiniLM-L6-v2Hugging Face384LocalFast baseline, quick experiments
BAAI/bge-small-en-v1.5Hugging Face384LocalCourse default; strong English retrieval
intfloat/multilingual-e5-smallHugging Face384LocalMany languages (Hindi, Kannada, etc.)
nomic-embed-textOllama768LocalLonger inputs, fully offline
gemini-embedding-001Gemini API3072 (default)CloudHigh quality, 100+ languages

Prefixes matter: some models expect a label in front of the text.

• bge-small-en-v1.5: add Represent this sentence for searching relevant passages:  before queries.

• multilingual-e5-small: add query:  before queries and passage:  before documents.

• nomic-embed-text: add search_query:  and search_document: .

The helper below handles these automatically.

Example: Hugging Face, direct

python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("BAAI/bge-small-en-v1.5")
docs = ["Electronics can be returned within 15 days.", "Express shipping takes 1-2 days."]
query = "Represent this sentence for searching relevant passages: how long do I have to return a laptop?"

doc_vecs = model.encode(docs, normalize_embeddings=True)
query_vec = model.encode(query, normalize_embeddings=True)
print(doc_vecs @ query_vec)    # higher = more similar; the returns doc should win

Example: Hugging Face via LangChain

python
# pip install langchain-huggingface\
from langchain_huggingface import HuggingFaceEmbeddings\
\
emb = HuggingFaceEmbeddings(\
    model_name="sentence-transformers/all-MiniLM-L6-v2",\
    encode_kwargs={"normalize_embeddings": True},\
)\
print(len(emb.embed_query("How long do I have to return a laptop?")))   # 384


Example: Ollama via LangChain

python
from langchain_ollama import OllamaEmbeddings

emb = OllamaEmbeddings(model="nomic-embed-text")
print(len(emb.embed_query("search_query: how long do I have to return a laptop?")))   # 768

Example: Gemini

python
from google import genai
from google.genai import types

client = genai.Client()
result = client.models.embed_content(
    model="gemini-embedding-001",
    contents=["Electronics can be returned within 15 days.",
              "Express shipping takes 1-2 days."],
    config=types.EmbedContentConfig(task_type="RETRIEVAL_DOCUMENT"),
)
print(len(result.embeddings), len(result.embeddings[0].values))

Gemini also offers gemini-embedding-2, which embeds images, audio, and video into the same space. We'll use it in the multimodal module.

The course helper: common/embeddings.py

python
# common/embeddings.py
import numpy as np
from dotenv import load_dotenv

load_dotenv()

DEFAULT_MODELS = {
    "hf": "BAAI/bge-small-en-v1.5",
    "ollama": "nomic-embed-text",
    "gemini": "gemini-embedding-001",
}
QUERY_PREFIX = {
    "BAAI/bge-small-en-v1.5": "Represent this sentence for searching relevant passages: ",
    "intfloat/multilingual-e5-small": "query: ",
    "nomic-embed-text": "search_query: ",
}
DOC_PREFIX = {
    "intfloat/multilingual-e5-small": "passage: ",
    "nomic-embed-text": "search_document: ",
}


def _normalize(vectors) -> np.ndarray:
    m = np.asarray(vectors, dtype="float32")
    return m / np.linalg.norm(m, axis=-1, keepdims=True)


class Embedder:
    """One interface for every embedding backend: 'hf', 'ollama', or 'gemini'."""

    def __init__(self, backend: str = "hf", model: str | None = None):
        self.backend = backend
        self.model = model or DEFAULT_MODELS[backend]
        if backend == "hf":
            from sentence_transformers import SentenceTransformer
            self._st = SentenceTransformer(self.model)
        elif backend == "ollama":
            from langchain_ollama import OllamaEmbeddings
            self._lc = OllamaEmbeddings(model=self.model)
        elif backend == "gemini":
            from google import genai
            self._client = genai.Client()
        else:
            raise ValueError(f"Unknown backend: {backend}")

    def _embed(self, texts: list[str], task: str):
        if self.backend == "hf":
            return self._st.encode(texts)
        if self.backend == "ollama":
            return self._lc.embed_documents(texts)
        from google.genai import types
        vectors = []
        for i in range(0, len(texts), 100):          # send in batches of 100
            result = self._client.models.embed_content(
                model=self.model,
                contents=texts[i:i + 100],
                config=types.EmbedContentConfig(task_type=task),
            )
            vectors += [e.values for e in result.embeddings]
        return vectors

    def embed_documents(self, texts: list[str]) -> np.ndarray:
        prefix = DOC_PREFIX.get(self.model, "")
        return _normalize(self._embed([prefix + t for t in texts], "RETRIEVAL_DOCUMENT"))

    def embed_query(self, text: str) -> np.ndarray:
        prefix = QUERY_PREFIX.get(self.model, "")
        return _normalize(self._embed([prefix + text], "RETRIEVAL_QUERY"))[0]

All vectors are normalized (length 1), so a dot product equals cosine similarity. We'll explain why in the similarity module.

5.5 Rerankers (Hugging Face Cross-Encoders)

A reranker reads the question and each candidate together and scores how well they match. It's slower than vector search but more accurate, so we use it only on the top 20-50 candidates.

ModelSizeBest for
cross-encoder/ms-marco-MiniLM-L-6-v2SmallCourse default; fast on CPU
BAAI/bge-reranker-baseMediumHigher accuracy
BAAI/bge-reranker-v2-m3LargerMultilingual content

Example

python
from sentence_transformers import CrossEncoder

reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
question = "Can I return opened headphones?"
candidates = [
    "Express shipping takes 1-2 business days.",
    "Opened electronics can be returned within 7 days for store credit only.",
    "Unopened electronics can be returned within 15 days.",
]
scores = reranker.predict([(question, c) for c in candidates])
for score, text in sorted(zip(scores, candidates), reverse=True):
    print(f"{score:6.2f}  {text}")

5.6 Vector Databases

A vector database stores vectors and quickly finds the ones closest to a query vector.

Vector DBTypeBest forUsed in this course for
FAISSLibrary (in your Python process)Learning, fast local experimentsUnderstanding index types
ChromaLightweight DBPrototypes, small appsQuick demos, LangChain examples
QdrantFull DB (in-memory, local, server, or cloud)Production with filtersCourse default
pgvectorPostgreSQL extensionTeams already using PostgresMixing SQL filters and vectors
PineconeFully managed cloud serviceZero-ops productionManaged deployment

The examples below share this setup:

python
from common.embeddings import Embedder

docs = [
    {"id": "returns-1", "text": "Unopened electronics can be returned within 15 days."},
    {"id": "returns-2", "text": "Opened electronics can be returned within 7 days for store credit only."},
    {"id": "shipping-1", "text": "Express shipping takes 1-2 business days."},
]
embedder = Embedder("hf")                                   # 384 dimensions
doc_vecs = embedder.embed_documents([d["text"] for d in docs])
query = "Can I send back headphones I already opened?"
query_vec = embedder.embed_query(query)

Example: FAISS

python
# pip install faiss-cpu
import faiss

index = faiss.IndexFlatIP(doc_vecs.shape[1])      # inner product = cosine for normalized vectors
index.add(doc_vecs)
scores, positions = index.search(query_vec.reshape(1, -1), 2)
for score, pos in zip(scores[0], positions[0]):
    print(f"{score:.2f}  {docs[pos]['id']}")

Example: Chroma

python
# pip install chromadb
import chromadb

client = chromadb.EphemeralClient()                # in-memory; use PersistentClient(path=...) to save
collection = client.get_or_create_collection("shopsphere")
collection.add(
    ids=[d["id"] for d in docs],
    documents=[d["text"] for d in docs],
    embeddings=doc_vecs.tolist(),
)
result = collection.query(query_embeddings=[query_vec.tolist()], n_results=2)
print(result["ids"][0], result["distances"][0])     # lower distance = more similar

Chroma's default metric is distance-based, and with normalized vectors it ranks results the same way as cosine similarity.

Example: Qdrant (course default)

python
from qdrant_client import QdrantClient, models

client = QdrantClient(":memory:")                  # same code works with a real server URL
client.create_collection(
    collection_name="shopsphere",
    vectors_config=models.VectorParams(size=doc_vecs.shape[1], distance=models.Distance.COSINE),
)
client.upsert(
    collection_name="shopsphere",
    points=[models.PointStruct(id=i, vector=v.tolist(), payload=d)
            for i, (d, v) in enumerate(zip(docs, doc_vecs))],
)
hits = client.query_points("shopsphere", query=query_vec.tolist(), limit=2).points
for h in hits:
    print(f"{h.score:.2f}  {h.payload['id']}")

Example: pgvector (PostgreSQL)

python
docker run -d --name pgvector -e POSTGRES_PASSWORD=pass -p 5432:5432 pgvector/pgvector:pg16
pip install "psycopg[binary]"

import psycopg

with psycopg.connect("postgresql://postgres:pass@localhost:5432/postgres", autocommit=True) as conn:
    conn.execute("CREATE EXTENSION IF NOT EXISTS vector")
    conn.execute("DROP TABLE IF EXISTS chunks")
    conn.execute("CREATE TABLE chunks (id text PRIMARY KEY, text text, embedding vector(384))")
    for d, v in zip(docs, doc_vecs):
        conn.execute("INSERT INTO chunks VALUES (%s, %s, %s::vector)",
                     (d["id"], d["text"], str(v.tolist())))
    rows = conn.execute(
        "SELECT id, 1 - (embedding <=> %s::vector) AS similarity "
        "FROM chunks ORDER BY embedding <=> %s::vector LIMIT 2",
        (str(query_vec.tolist()), str(query_vec.tolist())),
    ).fetchall()
    print(rows)                                    # <=> is cosine distance

Example: Pinecone (managed cloud):

python
# pip install pinecone      (needs PINECONE_API_KEY; free starter plan available)
import os
from pinecone import Pinecone, ServerlessSpec

pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])
if not pc.has_index("shopsphere"):
    pc.create_index(name="shopsphere", dimension=384, metric="cosine",
                    spec=ServerlessSpec(cloud="aws", region="us-east-1"))
index = pc.Index("shopsphere")
index.upsert(vectors=[{"id": d["id"], "values": v.tolist(), "metadata": {"text": d["text"]}}
                      for d, v in zip(docs, doc_vecs)])
result = index.query(vector=query_vec.tolist(), top_k=2, include_metadata=True)
for match in result.matches:
    print(f"{match.score:.2f}  {match.id}")

Newly upserted vectors can take a few seconds to become searchable in Pinecone.

Dimension must match: a 384-dimension index only accepts 384-dimension vectors. Changing the embedding model means creating a new index and re-embedding everything.

5.7 Frameworks

OptionStrengthWhen we use it
Plain PythonYou see every stepLearning each concept first
LangChainHuge integration library; composable chains; agents via LangGraphReal projects, tool calling, agents
LlamaIndexStrong data loading and indexingDocument-heavy pipelines

Parsing, evaluation, and agent libraries (such as document parsers and RAG evaluation tools) are introduced in the modules where they're needed.

Example: a full RAG chain in LangChain

python
# pip install langchain-huggingface langchain-chroma langchain-groq
from langchain_chroma import Chroma
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_groq import ChatGroq
from langchain_huggingface import HuggingFaceEmbeddings

texts = ["Unopened electronics can be returned within 15 days.",
         "Opened electronics can be returned within 7 days for store credit only.",
         "Express shipping takes 1-2 business days."]
ids = ["returns-1", "returns-2", "shipping-1"]

emb = HuggingFaceEmbeddings(model_name="BAAI/bge-small-en-v1.5",
                            encode_kwargs={"normalize_embeddings": True})
store = Chroma.from_texts(texts, embedding=emb, metadatas=[{"id": i} for i in ids])
retriever = store.as_retriever(search_kwargs={"k": 2})

def format_docs(found):
    return "\n".join(f"[{d.metadata['id']}] {d.page_content}" for d in found)

prompt = ChatPromptTemplate.from_messages([
    ("system", "Answer ONLY from the context. Cite chunk IDs in square brackets. "
               "If the answer isn't in the context, say you don't know."),
    ("human", "Context:\n{context}\n\nQuestion: {question}"),
])

llm = ChatGroq(model="llama-3.3-70b-versatile", temperature=0)
# Swap in another provider with one line:
# from langchain_ollama import ChatOllama;  llm = ChatOllama(model="llama3.1", temperature=0)
# from langchain_google_genai import ChatGoogleGenerativeAI;  llm = ChatGoogleGenerativeAI(model="gemini-3.8-flash")

chain = ({"context": retriever | format_docs, "question": RunnablePassthrough()}
         | prompt | llm | StrOutputParser())
print(chain.invoke("Can I return headphones I already opened?"))c

Example: the same RAG in LlamaIndex (fully local)

python
# pip install llama-index-core llama-index-embeddings-huggingface llama-index-llms-ollama
from llama_index.core import Document, Settings, VectorStoreIndex
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.ollama import Ollama

Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")
Settings.llm = Ollama(model="llama3.1", request_timeout=120.0)

documents = [
    Document(text="Unopened electronics can be returned within 15 days.", metadata={"id": "returns-1"}),
    Document(text="Opened electronics can be returned within 7 days for store credit only.", metadata={"id": "returns-2"}),
    Document(text="Express shipping takes 1-2 business days.", metadata={"id": "shipping-1"}),
]
index = VectorStoreIndex.from_documents(documents)
response = index.as_query_engine(similarity_top_k=2).query("Can I return headphones I already opened?")

print(response)
for node in response.source_nodes:                 # the evidence behind the answer
    print(f"{node.score:.2f}  {node.metadata['id']}")