Skip to content

Glossary · 24 terms

GenAI

Prompts, RAG, agents and how models run.

Beginner · Students and graduates

Prompt
The input text a model responds to. Clear goals, context and examples in it usually matter more than clever wording.
Token
Models split text into tokens, often parts of words. Limits and pricing are counted in tokens, not characters.
LLM
A large language model: a neural network trained on text that generates language by predicting likely next tokens.
Model
The learned parameters and architecture that produce predictions. Different models trade off quality, speed and cost.
Chatbot
An app that holds a back-and-forth conversation, sending the chat history to a model on each turn.
Dataset
The data used to train or evaluate a model. Its quality and coverage shape what the model can and can’t do.
Bias
Systematic skew in outputs, often reflecting imbalances in training data. It needs testing and mitigation, not just good intentions.
Context
The context window holds the prompt, history and any added documents. Anything outside it, the model simply doesn’t know.

Practitioner · Working engineers

Embedding
A vector representation where similar meanings sit close together. It powers semantic search and clustering.
RAG
Retrieval-augmented generation: search your own data, put the best matches in the context, and have the model answer from them.
Agent
A system where the model decides which tool to use, sees the result and continues until the task is done.
Fine-tune
Adjusting a pre-trained model’s weights with task-specific data to change its style or behaviour.
Vector DB
Stores vectors and finds the nearest ones to a query quickly, usually with approximate nearest-neighbour indexes.
Grounding
Supplying facts in the context and asking for answers based on them, often with citations, to reduce made-up content.
Tool use
The model returns a structured request to call a tool you defined; your code runs it and sends back the result.
Chunking
Breaking text into sections sized for embedding and retrieval. Chunk size and overlap strongly affect RAG quality.

Lead · Seniors and managers

Attention
Self-attention lets each token look at every other token and decide how much each matters. It’s the core of transformers. Docs
MCP
The Model Context Protocol standardises how applications expose tools, resources and prompts to AI models. Docs
Evals
Sets of inputs with expected outcomes or graders, run on every change so you know whether quality went up or down.
LoRA
Low-Rank Adaptation freezes the base weights and learns small low-rank updates, making fine-tuning far cheaper. Docs
Quantize
Quantization reduces numeric precision, such as 16-bit to 4-bit, cutting memory and speeding up inference with some quality loss.
Reranker
After a fast retrieval step, a reranker scores each query and document pair more precisely and keeps the best few.
Guardrail
Validation, filters and policies applied before and after model calls, such as schema checks or content rules.
Inference
The prediction phase, as opposed to training. Latency, throughput and cost per token are inference concerns.