Topic 5: The Course Tech Stack
Everything in this course uses free or free-tier tools, so every learner can run every lab. Here's the full stack, and why each piece is in it.
5.1 The Stack at a Glance
| Layer | Course default | Also covered | Why the default |
| LLM | Groq (llama-3.3-70b-versatile) | Gemini (gemini-3.8-flash), Ollama (llama3.1), Hugging Face local models | Free tier and very fast responses |
| Embedding model | Hugging Face BAAI/bge-small-en-v1.5 | all-MiniLM-L6-v2, multilingual-e5-small, Ollama nomic-embed-text, Gemini gemini-embedding-001 | Free, local, strong retrieval quality for its size |
| Reranker | Hugging Face cross-encoder/ms-marco-MiniLM-L-6-v2 | BAAI/bge-reranker-base, BAAI/bge-reranker-v2-m3 | Small, fast, runs on CPU |
| Vector DB | Qdrant (in-memory for labs) | FAISS, Chroma, pgvector, Pinecone | Runs locally with no setup, and the same code scales to a server |
| Frameworks | Plain Python (to learn how things work) | LangChain, LlamaIndex | Understand the internals first, then use frameworks to go faster |
Teaching approach: each concept is first built in plain Python so learners see what's happening, then shown in LangChain or LlamaIndex for real projects.
5.2 Project Setup
Folder structure (used for the whole course)
rag-course/
├── .env # API keys (never commit this)
├── requirements.txt
├── common/
│ ├── init.py
│ ├── llm.py # one ask() function for every LLM provider
│ └── embeddings.py # one Embedder class for every embedding provider
└── module01/
└── lab.py
Install
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate\
pip install python-dotenv numpy
groq "google-genai>=2.3.0" langchain-ollama langchain-core
sentence-transformers qdrant-client
Extra packages for specific examples are listed next to each example.
**Local models (for Ollama options)**, after installing Ollama from ollama.com:
ollama pull llama3.1\
ollama pull nomic-embed-text
`.env`** file**
GROQ_API_KEY=your-groq-key # console.groq.com\
GEMINI_API_KEY=your-gemini-key # aistudio.google.com\
LLM_PROVIDER=groq # groq | gemini | ollama
5.3 LLMs
| Provider | Model | Cost | Best for |
| Groq | llama-3.3-70b-versatile | Free tier (rate-limited) | Default for labs; very fast |
| Gemini | gemini-3.8-flash | Free tier available | Strong quality; long context |
| Ollama | llama3.1 | Free, runs on your computer | Offline work and private data |
| Hugging Face | e.g. Qwen/Qwen2.5-1.5B-Instruct | Free, runs on your computer | Learning how models run locally |
Example: Groq
import os\
from groq import Groq\
client = Groq(api_key=os.environ.get("GROQ_API_KEY"))\
chat_completion = client.chat.completions.create(\
messages=[{"role": "user", "content": "Explain the importance of fast language models"}],\
model="llama-3.3-70b-versatile",\
)
print(chat_completion.choices[0].message.content)
**Example: Gemini**
from google import genai\
client = genai.Client() # reads GEMINI_API_KEY\
interaction = client.interactions.create(\
model="gemini-3.8-flash",\
input="Explain how AI works in a few words",\
)
print(interaction.output_text)
**Example: Ollama via LangChain**
from langchain_ollama import ChatOllama\
llm = ChatOllama(model="llama3.1", temperature=0)\
print(llm.invoke("Explain how AI works in a few words").content)
**Example: Hugging Face model running locally**
# pip install transformers torch
from transformers import pipelin
generator = pipeline("text-generation", model="Qwen/Qwen2.5-1.5B-Instruct")\
messages = [{"role": "user", "content": "Explain RAG in one sentence."}]\
output = generator(messages, max_new_tokens=100)\
print(output[0]["generated_text"][-1]["content"])
Small local models are great for learning but less reliable at following strict rules (citation formats, JSON-only output). Always validate their output in code.
The course helper: common/llm.py
Every lab calls one function, ask(). To switch providers, change LLM_PROVIDER in .env or pass provider=. The lesson code never changes.
# common/llm.py
import json
import os
import re
from functools import lru_cache
from dotenv import load_dotenv
load_dotenv()
DEFAULT_PROVIDER = os.getenv("LLM_PROVIDER", "groq")
MODELS = {
"groq": "llama-3.3-70b-versatile",
"gemini": "gemini-3.8-flash",
"ollama": "llama3.1",
}
@lru_cache
def _groq():
from groq import Groq
return Groq(api_key=os.environ.get("GROQ_API_KEY"))
@lru_cache
def _gemini():
from google import genai
return genai.Client()
@lru_cache
def _ollama(json_mode: bool, temperature: float):
from langchain_ollama import ChatOllama
extra = {"format": "json"} if json_mode else {}
return ChatOllama(model=MODELS["ollama"], temperature=temperature, **extra)
def ask(prompt: str,
system: str = "You are a helpful assistant.",
provider: str | None = None,
json_mode: bool = False,
temperature: float = 0.0) -> str:
"""Send one prompt to the chosen provider and return the reply text."""
provider = provider or DEFAULT_PROVIDER
if provider == "groq":
# Groq's JSON mode requires the word "JSON" somewhere in the messages.
extra = {"response_format": {"type": "json_object"}} if json_mode else {}
resp = _groq().chat.completions.create(
model=MODELS["groq"],
messages=[{"role": "system", "content": system},
{"role": "user", "content": prompt}],
temperature=temperature,
**extra,
)
return resp.choices[0].message.content
if provider == "gemini":
interaction = _gemini().interactions.create(
model=MODELS["gemini"],
system_instruction=system,
input=prompt,
generation_config={"temperature": temperature},
)
return interaction.output_text
if provider == "ollama":
from langchain_core.messages import HumanMessage, SystemMessage
reply = _ollama(json_mode, temperature).invoke(
[SystemMessage(system), HumanMessage(prompt)])
return reply.content
raise ValueError(f"Unknown provider: {provider}")
def parse_json(text: str, default=None):
"""Pull the first JSON object out of a model reply. Returns default if none."""
match = re.search(r"\{.*\}", text or "", re.S)
try:
return json.loads(match.group(0)) if match else default
except json.JSONDecodeError:
return default
Try it
from common.llm import ask
for provider in ["groq", "gemini", "ollama"]:
print(provider, "→", ask("What is RAG, in one sentence?", provider=provider))
5.4 Embedding Models
An embedding model turns text into a list of numbers (a vector). Texts with similar meaning get similar vectors. That's what makes semantic search possible.
| Model | Source | Dimensions | Runs | Best for |
| sentence-transformers/all-MiniLM-L6-v2 | Hugging Face | 384 | Local | Fast baseline, quick experiments |
| BAAI/bge-small-en-v1.5 | Hugging Face | 384 | Local | Course default; strong English retrieval |
| intfloat/multilingual-e5-small | Hugging Face | 384 | Local | Many languages (Hindi, Kannada, etc.) |
| nomic-embed-text | Ollama | 768 | Local | Longer inputs, fully offline |
| gemini-embedding-001 | Gemini API | 3072 (default) | Cloud | High quality, 100+ languages |
Prefixes matter: some models expect a label in front of the text.
• bge-small-en-v1.5: add Represent this sentence for searching relevant passages: before queries.
• multilingual-e5-small: add query: before queries and passage: before documents.
• nomic-embed-text: add search_query: and search_document: .
The helper below handles these automatically.
Example: Hugging Face, direct
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-en-v1.5")
docs = ["Electronics can be returned within 15 days.", "Express shipping takes 1-2 days."]
query = "Represent this sentence for searching relevant passages: how long do I have to return a laptop?"
doc_vecs = model.encode(docs, normalize_embeddings=True)
query_vec = model.encode(query, normalize_embeddings=True)
print(doc_vecs @ query_vec) # higher = more similar; the returns doc should winExample: Hugging Face via LangChain
# pip install langchain-huggingface\
from langchain_huggingface import HuggingFaceEmbeddings\
\
emb = HuggingFaceEmbeddings(\
model_name="sentence-transformers/all-MiniLM-L6-v2",\
encode_kwargs={"normalize_embeddings": True},\
)\
print(len(emb.embed_query("How long do I have to return a laptop?"))) # 384
Example: Ollama via LangChain
from langchain_ollama import OllamaEmbeddings
emb = OllamaEmbeddings(model="nomic-embed-text")
print(len(emb.embed_query("search_query: how long do I have to return a laptop?"))) # 768Example: Gemini
from google import genai
from google.genai import types
client = genai.Client()
result = client.models.embed_content(
model="gemini-embedding-001",
contents=["Electronics can be returned within 15 days.",
"Express shipping takes 1-2 days."],
config=types.EmbedContentConfig(task_type="RETRIEVAL_DOCUMENT"),
)
print(len(result.embeddings), len(result.embeddings[0].values))Gemini also offers gemini-embedding-2, which embeds images, audio, and video into the same space. We'll use it in the multimodal module.
The course helper: common/embeddings.py
# common/embeddings.py
import numpy as np
from dotenv import load_dotenv
load_dotenv()
DEFAULT_MODELS = {
"hf": "BAAI/bge-small-en-v1.5",
"ollama": "nomic-embed-text",
"gemini": "gemini-embedding-001",
}
QUERY_PREFIX = {
"BAAI/bge-small-en-v1.5": "Represent this sentence for searching relevant passages: ",
"intfloat/multilingual-e5-small": "query: ",
"nomic-embed-text": "search_query: ",
}
DOC_PREFIX = {
"intfloat/multilingual-e5-small": "passage: ",
"nomic-embed-text": "search_document: ",
}
def _normalize(vectors) -> np.ndarray:
m = np.asarray(vectors, dtype="float32")
return m / np.linalg.norm(m, axis=-1, keepdims=True)
class Embedder:
"""One interface for every embedding backend: 'hf', 'ollama', or 'gemini'."""
def __init__(self, backend: str = "hf", model: str | None = None):
self.backend = backend
self.model = model or DEFAULT_MODELS[backend]
if backend == "hf":
from sentence_transformers import SentenceTransformer
self._st = SentenceTransformer(self.model)
elif backend == "ollama":
from langchain_ollama import OllamaEmbeddings
self._lc = OllamaEmbeddings(model=self.model)
elif backend == "gemini":
from google import genai
self._client = genai.Client()
else:
raise ValueError(f"Unknown backend: {backend}")
def _embed(self, texts: list[str], task: str):
if self.backend == "hf":
return self._st.encode(texts)
if self.backend == "ollama":
return self._lc.embed_documents(texts)
from google.genai import types
vectors = []
for i in range(0, len(texts), 100): # send in batches of 100
result = self._client.models.embed_content(
model=self.model,
contents=texts[i:i + 100],
config=types.EmbedContentConfig(task_type=task),
)
vectors += [e.values for e in result.embeddings]
return vectors
def embed_documents(self, texts: list[str]) -> np.ndarray:
prefix = DOC_PREFIX.get(self.model, "")
return _normalize(self._embed([prefix + t for t in texts], "RETRIEVAL_DOCUMENT"))
def embed_query(self, text: str) -> np.ndarray:
prefix = QUERY_PREFIX.get(self.model, "")
return _normalize(self._embed([prefix + text], "RETRIEVAL_QUERY"))[0]All vectors are normalized (length 1), so a dot product equals cosine similarity. We'll explain why in the similarity module.
5.5 Rerankers (Hugging Face Cross-Encoders)
A reranker reads the question and each candidate together and scores how well they match. It's slower than vector search but more accurate, so we use it only on the top 20-50 candidates.
| Model | Size | Best for |
| cross-encoder/ms-marco-MiniLM-L-6-v2 | Small | Course default; fast on CPU |
| BAAI/bge-reranker-base | Medium | Higher accuracy |
| BAAI/bge-reranker-v2-m3 | Larger | Multilingual content |
Example
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
question = "Can I return opened headphones?"
candidates = [
"Express shipping takes 1-2 business days.",
"Opened electronics can be returned within 7 days for store credit only.",
"Unopened electronics can be returned within 15 days.",
]
scores = reranker.predict([(question, c) for c in candidates])
for score, text in sorted(zip(scores, candidates), reverse=True):
print(f"{score:6.2f} {text}")5.6 Vector Databases
A vector database stores vectors and quickly finds the ones closest to a query vector.
| Vector DB | Type | Best for | Used in this course for |
| FAISS | Library (in your Python process) | Learning, fast local experiments | Understanding index types |
| Chroma | Lightweight DB | Prototypes, small apps | Quick demos, LangChain examples |
| Qdrant | Full DB (in-memory, local, server, or cloud) | Production with filters | Course default |
| pgvector | PostgreSQL extension | Teams already using Postgres | Mixing SQL filters and vectors |
| Pinecone | Fully managed cloud service | Zero-ops production | Managed deployment |
The examples below share this setup:
from common.embeddings import Embedder
docs = [
{"id": "returns-1", "text": "Unopened electronics can be returned within 15 days."},
{"id": "returns-2", "text": "Opened electronics can be returned within 7 days for store credit only."},
{"id": "shipping-1", "text": "Express shipping takes 1-2 business days."},
]
embedder = Embedder("hf") # 384 dimensions
doc_vecs = embedder.embed_documents([d["text"] for d in docs])
query = "Can I send back headphones I already opened?"
query_vec = embedder.embed_query(query)Example: FAISS
# pip install faiss-cpu
import faiss
index = faiss.IndexFlatIP(doc_vecs.shape[1]) # inner product = cosine for normalized vectors
index.add(doc_vecs)
scores, positions = index.search(query_vec.reshape(1, -1), 2)
for score, pos in zip(scores[0], positions[0]):
print(f"{score:.2f} {docs[pos]['id']}")Example: Chroma
# pip install chromadb
import chromadb
client = chromadb.EphemeralClient() # in-memory; use PersistentClient(path=...) to save
collection = client.get_or_create_collection("shopsphere")
collection.add(
ids=[d["id"] for d in docs],
documents=[d["text"] for d in docs],
embeddings=doc_vecs.tolist(),
)
result = collection.query(query_embeddings=[query_vec.tolist()], n_results=2)
print(result["ids"][0], result["distances"][0]) # lower distance = more similarChroma's default metric is distance-based, and with normalized vectors it ranks results the same way as cosine similarity.
Example: Qdrant (course default)
from qdrant_client import QdrantClient, models
client = QdrantClient(":memory:") # same code works with a real server URL
client.create_collection(
collection_name="shopsphere",
vectors_config=models.VectorParams(size=doc_vecs.shape[1], distance=models.Distance.COSINE),
)
client.upsert(
collection_name="shopsphere",
points=[models.PointStruct(id=i, vector=v.tolist(), payload=d)
for i, (d, v) in enumerate(zip(docs, doc_vecs))],
)
hits = client.query_points("shopsphere", query=query_vec.tolist(), limit=2).points
for h in hits:
print(f"{h.score:.2f} {h.payload['id']}")Example: pgvector (PostgreSQL)
docker run -d --name pgvector -e POSTGRES_PASSWORD=pass -p 5432:5432 pgvector/pgvector:pg16
pip install "psycopg[binary]"
import psycopg
with psycopg.connect("postgresql://postgres:pass@localhost:5432/postgres", autocommit=True) as conn:
conn.execute("CREATE EXTENSION IF NOT EXISTS vector")
conn.execute("DROP TABLE IF EXISTS chunks")
conn.execute("CREATE TABLE chunks (id text PRIMARY KEY, text text, embedding vector(384))")
for d, v in zip(docs, doc_vecs):
conn.execute("INSERT INTO chunks VALUES (%s, %s, %s::vector)",
(d["id"], d["text"], str(v.tolist())))
rows = conn.execute(
"SELECT id, 1 - (embedding <=> %s::vector) AS similarity "
"FROM chunks ORDER BY embedding <=> %s::vector LIMIT 2",
(str(query_vec.tolist()), str(query_vec.tolist())),
).fetchall()
print(rows) # <=> is cosine distanceExample: Pinecone (managed cloud):
# pip install pinecone (needs PINECONE_API_KEY; free starter plan available)
import os
from pinecone import Pinecone, ServerlessSpec
pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])
if not pc.has_index("shopsphere"):
pc.create_index(name="shopsphere", dimension=384, metric="cosine",
spec=ServerlessSpec(cloud="aws", region="us-east-1"))
index = pc.Index("shopsphere")
index.upsert(vectors=[{"id": d["id"], "values": v.tolist(), "metadata": {"text": d["text"]}}
for d, v in zip(docs, doc_vecs)])
result = index.query(vector=query_vec.tolist(), top_k=2, include_metadata=True)
for match in result.matches:
print(f"{match.score:.2f} {match.id}")Newly upserted vectors can take a few seconds to become searchable in Pinecone.
Dimension must match: a 384-dimension index only accepts 384-dimension vectors. Changing the embedding model means creating a new index and re-embedding everything.
5.7 Frameworks
| Option | Strength | When we use it |
| Plain Python | You see every step | Learning each concept first |
| LangChain | Huge integration library; composable chains; agents via LangGraph | Real projects, tool calling, agents |
| LlamaIndex | Strong data loading and indexing | Document-heavy pipelines |
Parsing, evaluation, and agent libraries (such as document parsers and RAG evaluation tools) are introduced in the modules where they're needed.
Example: a full RAG chain in LangChain
# pip install langchain-huggingface langchain-chroma langchain-groq
from langchain_chroma import Chroma
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_groq import ChatGroq
from langchain_huggingface import HuggingFaceEmbeddings
texts = ["Unopened electronics can be returned within 15 days.",
"Opened electronics can be returned within 7 days for store credit only.",
"Express shipping takes 1-2 business days."]
ids = ["returns-1", "returns-2", "shipping-1"]
emb = HuggingFaceEmbeddings(model_name="BAAI/bge-small-en-v1.5",
encode_kwargs={"normalize_embeddings": True})
store = Chroma.from_texts(texts, embedding=emb, metadatas=[{"id": i} for i in ids])
retriever = store.as_retriever(search_kwargs={"k": 2})
def format_docs(found):
return "\n".join(f"[{d.metadata['id']}] {d.page_content}" for d in found)
prompt = ChatPromptTemplate.from_messages([
("system", "Answer ONLY from the context. Cite chunk IDs in square brackets. "
"If the answer isn't in the context, say you don't know."),
("human", "Context:\n{context}\n\nQuestion: {question}"),
])
llm = ChatGroq(model="llama-3.3-70b-versatile", temperature=0)
# Swap in another provider with one line:
# from langchain_ollama import ChatOllama; llm = ChatOllama(model="llama3.1", temperature=0)
# from langchain_google_genai import ChatGoogleGenerativeAI; llm = ChatGoogleGenerativeAI(model="gemini-3.8-flash")
chain = ({"context": retriever | format_docs, "question": RunnablePassthrough()}
| prompt | llm | StrOutputParser())
print(chain.invoke("Can I return headphones I already opened?"))cExample: the same RAG in LlamaIndex (fully local)
# pip install llama-index-core llama-index-embeddings-huggingface llama-index-llms-ollama
from llama_index.core import Document, Settings, VectorStoreIndex
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.llms.ollama import Ollama
Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")
Settings.llm = Ollama(model="llama3.1", request_timeout=120.0)
documents = [
Document(text="Unopened electronics can be returned within 15 days.", metadata={"id": "returns-1"}),
Document(text="Opened electronics can be returned within 7 days for store credit only.", metadata={"id": "returns-2"}),
Document(text="Express shipping takes 1-2 business days.", metadata={"id": "shipping-1"}),
]
index = VectorStoreIndex.from_documents(documents)
response = index.as_query_engine(similarity_top_k=2).query("Can I return headphones I already opened?")
print(response)
for node in response.source_nodes: # the evidence behind the answer
print(f"{node.score:.2f} {node.metadata['id']}")