https://github.com/mlabonne/llm-course Instruct model Chat model Base model Raw text Instructions Preferences Supervised fine-tuning Preference alignment Autocomplete prompts Follow instructions Optimized for humans Post-Training Pre- training
Covers a wide range of topics Complexity Non-trivial tasks forcing reasoning Find more information in the LLM Datasets repo on GitHub: https://github.com/mlabonne/llm-datasets
Remove the spaces from the following sentence: Fine- tuning is simple. Output Fine-tuningissimple. System (optional) You are a helpful assistant with a great sense of humor. Instruction Tell me a joke about octopuses. Chosen answer Why don't octopuses play cards in casinos? Because they can't count past eight. Rejected answer How many tickles does it take to make an octopus laugh? Ten tickles. Instruction data Preference data
w/ IFEval Keyword exclusion SFT example: Instruction following Example Write a detailed review of the movie "The Social Network". Your entire response should be in English and all lower case (no capital letters whatsoever)
answers Preference example: Ultrafeedback Query LLM 1 Query LLM n … Ganqu Cui et al. "UltraFeedback: Boosting Language Models with Scaled AI Feedback." arXiv preprint arXiv:2310.01377, October 2023.
provide explanation. Think like you are answering to a five- year-old.<|im_end|> <|im_start|>user Remove the spaces from the following sentence: It prevents users to suspect that there are some hidden products installed on their device.<|im_end|> <|im_start|>assistant Itpreventsuserstosuspectthattherearesomehiddenproduc tsinstalledontheirsdevice.<|im_end|> Storage format: Alpaca (Other examples: ShareGPT, OpenAI) Chat template: ChatML (Other examples: Llama 3, Mistral Instruct) System (optional) You are a helpful assistant, who always provide explanation. Think like you are answering to a five- year-old. Instruction Remove the spaces from the following sentence: It prevents users to suspect that there are some hidden products installed on their device. Output Itpreventsuserstosuspectthattherearesomehiddenpr oductsinstalledontheirsdevice.
Quantization-Aware Low-Rank Adaptation of Large Language Models." arXiv preprint arXiv:2309.14717 (2023). LoRA 16-bit precision QLoRA 4-bit precision Full Fine-Tuning 16-bit precision Maximizes quality Very high VRAM usage Fastest training High VRAM usage Low VRAM usage Degrades performance
the parameters are updated during training 1e-6 to 1e-3 2️⃣ Batch size Number of samples processed before updating parameters 8 or 16 (effective) Max length Longest input (in tokens) the model can process 1024 to 4096 Epochs Number of passes through the entire training dataset 3 to 5 Optimizer Algorithm to update the parameters to minimize the loss function AdamW Attention Implementation of the attention mechanism FlashAttention-2
ground- truth answers (e.g., accuracy). Example: MMLU, Open LLM Leaderboard Advantages Consistent & reproducible Cost-effective at scale Clear dimensions (e.g., math) Limitations ❌Not how models are used ❌Hard to evaluate complex tasks ❌Risk of data contamination Clémentine Fourrier and The Hugging Face Community, "LLM Evaluation Guidebook.", 2024. Open LLM Leaderboard by Hugging Face
on specific properties (accuracy, relevance, toxicity, etc.). Example: LLM-as-a-judge, Reward Models, small classifiers Advantages ✅This is how models are used ✅Can handle complex tasks ✅Provide direct feedback Limitations ❌Hidden biases (e.g., length, tone) ❌Quality validation needed ❌Costly at scale Clémentine Fourrier and The Hugging Face Community, "LLM Evaluation Guidebook.", 2024. EQ-Bench by Samuel J. Peach