Glossary · 24 terms
GenAI
Prompts, RAG, agents and how models run.
Beginner · Students and graduates
- Prompt
- The input text a model responds to. Clear goals, context and examples in it usually matter more than clever wording.
- Token
- Models split text into tokens, often parts of words. Limits and pricing are counted in tokens, not characters.
- LLM
- A large language model: a neural network trained on text that generates language by predicting likely next tokens.
- Model
- The learned parameters and architecture that produce predictions. Different models trade off quality, speed and cost.
- Chatbot
- An app that holds a back-and-forth conversation, sending the chat history to a model on each turn.
- Dataset
- The data used to train or evaluate a model. Its quality and coverage shape what the model can and can’t do.
- Bias
- Systematic skew in outputs, often reflecting imbalances in training data. It needs testing and mitigation, not just good intentions.
- Context
- The context window holds the prompt, history and any added documents. Anything outside it, the model simply doesn’t know.
Practitioner · Working engineers
- Embedding
- A vector representation where similar meanings sit close together. It powers semantic search and clustering.
- RAG
- Retrieval-augmented generation: search your own data, put the best matches in the context, and have the model answer from them.
- Agent
- A system where the model decides which tool to use, sees the result and continues until the task is done.
- Fine-tune
- Adjusting a pre-trained model’s weights with task-specific data to change its style or behaviour.
- Vector DB
- Stores vectors and finds the nearest ones to a query quickly, usually with approximate nearest-neighbour indexes.
- Grounding
- Supplying facts in the context and asking for answers based on them, often with citations, to reduce made-up content.
- Tool use
- The model returns a structured request to call a tool you defined; your code runs it and sends back the result.
- Chunking
- Breaking text into sections sized for embedding and retrieval. Chunk size and overlap strongly affect RAG quality.
Lead · Seniors and managers
- Attention
- Self-attention lets each token look at every other token and decide how much each matters. It’s the core of transformers. Docs
- MCP
- The Model Context Protocol standardises how applications expose tools, resources and prompts to AI models. Docs
- Evals
- Sets of inputs with expected outcomes or graders, run on every change so you know whether quality went up or down.
- LoRA
- Low-Rank Adaptation freezes the base weights and learns small low-rank updates, making fine-tuning far cheaper. Docs
- Quantize
- Quantization reduces numeric precision, such as 16-bit to 4-bit, cutting memory and speeding up inference with some quality loss.
- Reranker
- After a fast retrieval step, a reranker scores each query and document pair more precisely and keeps the best few.
- Guardrail
- Validation, filters and policies applied before and after model calls, such as schema checks or content rules.
- Inference
- The prediction phase, as opposed to training. Latency, throughput and cost per token are inference concerns.