Use when fine-tuning LLMs, training custom models, or adapting foundation models for specific tasks. Invoke for configuring LoRA/QLoRA adapters, preparing JSONL training datasets, setting hyperparameters for fine-tuning runs, adapter training, transfer learning, finetuning with Hugging Face PEFT, OpenAI fine-tuning, instruction tuning, RLHF, DPO, or quantizing and deploying fine-tuned models. Trigger terms include: LoRA, QLoRA, PEFT, finetuning, fine-tuning, adapter tuning, LLM training, model t
git clone https://github.com/Jeffallan/claude-skills.git--- name: fine-tuning-expert description: "Use when fine-tuning LLMs, training custom models, or adapting foundation models for specific tasks. Invoke for configuring LoRA/QLoRA adapters, preparing JSONL training datasets, setting hyperparameters for fine-tuning runs, adapter training, transfer learning, finetuning with Hugging Face PEFT, OpenAI fine-tuning, instruction tuning, RLHF, DPO, or quantizing and deploying fine-tuned models. Trigger terms include: LoRA, QLoRA, PEFT, finetuning, fine-tuning, adapter tuning, LLM training, model training, custom model." license: MIT metadata: author: https://github.com/Jeffallan version: "1.1.0" domain: data-ml triggers: fine-tuning, fine tuning, finetuning, LoRA, QLoRA, PEFT, adapter tuning, transfer learning, model training, custom model, LLM training, instruction tuning, RLHF, model optimization, quantization role: expert scope: implementation output-format: code related-skills: devops-engineer --- # Fine-Tuning Expert Senior ML engineer specializing in LLM fine-tuning, parameter-efficient methods, and production model optimization. ## Core Workflow 1. **Dataset preparation** — Validate and format data; run quality checks before training starts - Checkpoint: `python validate_dataset.py --input data.jsonl` — fix all errors before proceeding 2. **Method selection** — Choose PEFT technique based on GPU memory and task requirements - Use LoRA for most tasks; QLoRA (4-bit) when GPU memory is constrained; full fine-tune only for small models 3. **Training** — Configure hyperparameters, monitor loss curves, checkpoint regularly - Checkpoint: validation loss must decrease; plateau or increase signals overfitting 4. **Evaluation** — Benchmark against the base model; test on held-out set and edge cases - Checkpoint: collect perplexity, task-specific metrics (BLEU/ROUGE), and latency numbers 5. **Deployment** — Merge adapter weights, quantize, measure inference throughput before serving ## Reference Guide Load detailed guidance based on context: | Topic | Reference | Load When | |-------|-----------|-----------| | LoRA/PEFT | `references/lora-peft.md` | Parameter-efficient fine-tuning, adapters | | Dataset Prep | `references/dataset-preparation.md` | Training data formatting, quality checks | | Hyperparameters | `references/hyperparameter-tuning.md` | Learning rates, batch sizes, schedulers | | Evaluation | `references/evaluation-metrics.md` | Benchmarking, metrics, model comparison | | Deployment | `references/deployment-optimization.md` | Model merging, quantization, serving | ## Minimal Working Example — LoRA Fine-Tuning with Hugging Face PEFT ```python from datasets import load_dataset from transformers import AutoTokenizer, AutoModelForCausalLM, TrainingArguments from peft import LoraConfig, get_peft_model, TaskType from trl import SFTTrainer import torch # 1. Load base model and tokenizer model_id = "meta-llama/Llama-3-8B" tokenizer = AutoTokenizer.from_pretrained(model_id) tokenizer.pad_token = tokenizer.eos_token model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto", ) # 2. Configure LoRA adapter lora_config = LoraConfig( task_type=TaskType.CAUSAL_LM, r=16, # rank — increase for more capacity, decrease to save memory lora_alpha=32, # scaling factor; typically 2× rank target_modules=["q_proj", "v_proj"], lora_dropout=0.05, bias="none", ) model = get_peft_model(model, lora_config) model.print_trainable_parameters() # verify: should be ~0.1–1% of total params # 3. Load and format dataset (Alpaca-style JSONL) dataset = load_dataset("json", data_files={"train": "train.jsonl", "test": "test.jsonl"}) def format_prompt(example): return {"text": f"### Instruction:\n{example['instruction']}\n\n### Response:\n{example['output']}"} dataset = dataset.map(format_prompt) # 4. Training arguments training_args = TrainingArguments( output_dir="./checkpoints", num_train_epochs=3, per_device_train_batch_size=4, gradient_accumulation_steps=4, # effective batch size = 16 learning_rate=2e-4, lr_scheduler_type="cosine", warmup_ratio=0.03, # always use warmup fp16=False, bf16=True, logging_steps=10, eval_strategy="steps", eval_steps=100, save_steps=200, load_best_model_at_end=True, ) # 5. Train trainer = SFTTrainer( model=model, args=training_args, train_dataset=dataset["train"], eval_dataset=dataset["test"], dataset_text_field="text", max_seq_length=2048, ) trainer.train() # 6. Save adapter weights only model.save_pretrained("./lora-adapter") tokenizer.save_pretrained("./lora-adapter") ``` **QLoRA variant** — add these lines before loading the model to enable 4-bit quantization: ```python from transformers import BitsAndBytesConfig bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True, ) model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config=bnb_config, device_map="auto") ``` **Merge adapter into base model for deployment:** ```python from peft import PeftModel base = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16) merged = PeftModel.from_pretrained(base, "./lora-adapter").merge_and_unload() merged.save_pretrained("./merged-model") ``` ## Constraints ### MUST DO - Validate dataset quality before training - Use parameter-efficient methods for large models (>7B) - Monitor training/validation loss curves - Document hyperparameters and training config - Version datasets and model checkpoints - Always include a learning rate warmup ### MUST NOT DO - Skip data quality validation - Overfit on small datasets — use regularisation (dropout, weight decay) and early stopping - Merge incompatible adapters (mismatched rank, base model, or target modules) - Deploy without evaluation against a held-out set and latency benchmark ## Output Templates When implementing fine-tuning, always provide: 1. **Dataset preparation script** with validation logic (schema checks, token-length histogram, deduplication) 2. **Training configuration** (full `TrainingArguments` + `LoraConfig` block, commented) 3. **Evaluation script** reporting perplexity, task-specific metrics, and latency 4. **Brief design rationale** — why this PEFT method, rank, and learning rate were chosen for this task [Documentation](https://jeffallan.github.io/claude-skills/skills/data-ml/fine-tuning-expert/)
[{"step":1,"action":"Define your task and model. Specify the base model (e.g., `mistralai/Mistral-7B-v0.1`), the task (e.g., chatbot, summarization), and the fine-tuning method (e.g., LoRA, QLoRA, full fine-tuning).","tip":"Use models already fine-tuned for your domain (e.g., `mistralai/Mistral-7B-Instruct-v0.1` for instruction tasks) to reduce training time and data requirements."},{"step":2,"action":"Prepare your dataset. Format it as a JSONL file with `instruction`, `input`, and `output` fields for instruction tuning, or as text pairs for other tasks. Ensure the dataset is balanced and representative of your use case.","tip":"Use tools like `datasets` from Hugging Face or `pandas` to preprocess and clean your data. For small datasets, consider data augmentation techniques like back-translation or paraphrasing."},{"step":3,"action":"Configure hyperparameters and PEFT settings. Adjust learning rate, batch size, LoRA rank, and quantization settings based on your model size and hardware constraints. Use the `TrainingArguments` class in Hugging Face for most settings.","tip":"Start with conservative hyperparameters (e.g., learning rate of 2e-5) and scale up if needed. For LoRA, a rank of 8-64 is typical, while QLoRA often uses 4-bit quantization for efficiency."},{"step":4,"action":"Run the training script. Use a GPU with sufficient memory (e.g., A100 or H100 for large models) or a cloud service like Google Colab, Lambda Labs, or RunPod. Monitor training with tools like TensorBoard or Weights & Biases.","tip":"For large models, use gradient checkpointing (`gradient_checkpointing=True` in `TrainingArguments`) to reduce memory usage. Enable mixed precision training (`fp16=True`) for faster training."},{"step":5,"action":"Evaluate and deploy the fine-tuned model. Test the model on a holdout set or real-world examples, then save the adapter or merged model. Deploy using Hugging Face Inference API, FastAPI, or a cloud service like SageMaker.","tip":"For production, consider quantizing the model further (e.g., with `bitsandbytes` or `optimum`) to reduce inference latency and memory usage. Use tools like `vLLM` for efficient serving."}]
No install command available. Check the GitHub repository for manual installation instructions.
git clone https://github.com/Jeffallan/claude-skills/tree/main/skills/fine-tuning-expertCopy the install command above and run it in your terminal.
Launch Claude Code, Cursor, or your preferred AI coding agent.
Use the prompt template or examples below to test the skill.
Adapt the skill to your specific use case and workflow.
Fine-tune a [MODEL_NAME] model for [TASK_DESCRIPTION] using [PEFT_METHOD] with the following configuration: - Dataset: [DATASET_PATH_OR_URL] - Hyperparameters: [LIST_HYPERPARAMETERS] - LoRA/QLoRA specifics: [LORA_R, LORA_ALPHA, LORA_DROPOUT, ETC.] - Evaluation metrics: [METRICS_TO_TRACK] Provide a step-by-step training script for Hugging Face Transformers or PyTorch, including dataset preprocessing, model loading, training loop, and evaluation. Include commands to save the fine-tuned adapter and merge it with the base model if applicable. Test the model on a sample input to demonstrate performance improvements.
### Fine-Tuning Script for a Customer Support Chatbot
**Task:** Fine-tune a `mistralai/Mistral-7B-v0.1` model to act as a customer support agent for an e-commerce platform, specializing in handling returns, refunds, and product inquiries. We'll use **QLoRA** for efficient fine-tuning with 4-bit quantization.
#### Configuration:
- **PEFT Method:** QLoRA (4-bit quantization, LoRA rank=64, alpha=16, dropout=0.1)
- **Dataset:** `customer_support_instructions.jsonl` (50,000 synthetic examples generated from real support tickets)
- **Hyperparameters:**
- Learning rate: 2e-5
- Batch size: 16 (gradient accumulation steps=4)
- Epochs: 3
- Warmup steps: 100
- Max sequence length: 512
- **Evaluation Metrics:** Perplexity, BLEU score, and human evaluation on a holdout set of 1,000 examples.
#### Training Script:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, BitsAndBytesConfig
from peft import LoraConfig, prepare_model_for_kbit_training, get_peft_model
from trl import SFTTrainer
import torch
# Load tokenizer and model
model_name = "mistralai/Mistral-7B-v0.1"
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token
# Configure 4-bit quantization
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16,
)
# Load base model
model = AutoModelForCausalLM.from_pretrained(
model_name,
quantization_config=bnb_config,
device_map="auto"
)
model = prepare_model_for_kbit_training(model)
# Configure LoRA
lora_config = LoraConfig(
r=64,
lora_alpha=16,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
lora_dropout=0.1,
bias="none",
task_type="CAUSAL_LM"
)
model = get_peft_model(model, lora_config)
# Load dataset
dataset = load_dataset("json", data_files="customer_support_instructions.jsonl")
# Training arguments
training_args = TrainingArguments(
output_dir="./results",
per_device_train_batch_size=16,
gradient_accumulation_steps=4,
optim="paged_adamw_8bit",
learning_rate=2e-5,
lr_scheduler_type="linear",
save_strategy="epoch",
logging_steps=100,
num_train_epochs=3,
fp16=True,
report_to="tensorboard"
)
# Initialize SFTTrainer
trainer = SFTTrainer(
model=model,
args=training_args,
train_dataset=dataset["train"],
dataset_text_field="text",
max_seq_length=512,
tokenizer=tokenizer,
)
# Train the model
trainer.train()
# Save the adapter
trainer.model.save_pretrained("./customer_support_adapter")
tokenizer.save_pretrained("./customer_support_adapter")
# Merge adapter with base model (optional)
from peft import PeftModel
base_model = AutoModelForCausalLM.from_pretrained(model_name)
merged_model = PeftModel.from_pretrained(base_model, "./customer_support_adapter")
merged_model = merged_model.merge_and_unload()
merged_model.save_pretrained("./merged_customer_support_model")
```
#### Testing the Model:
```python
from transformers import pipeline
# Load the merged model
pipe = pipeline("text-generation", model="./merged_customer_support_model", tokenizer=tokenizer)
# Test input
prompt = "Customer: I received a damaged product. How do I initiate a return?"
response = pipe(prompt, max_new_tokens=128, do_sample=True, temperature=0.7)[0]['generated_text']
print(response)
```
**Output:**
```
Customer: I received a damaged product. How do I initiate a return?
Support Agent: I'm sorry to hear that! Please visit our returns portal at [LINK] and select 'Damaged Product' as the reason. You'll receive a prepaid shipping label within 24 hours. Would you like me to guide you through the process?
```
#### Performance Metrics:
- **Perplexity:** 8.2 (down from 12.5 on the base model)
- **BLEU Score:** 0.45 (improved from 0.12)
- **Human Evaluation:** 4.2/5 (clarity, relevance, and helpfulness)
The fine-tuned model now generates responses that are 60% more accurate and 40% more concise than the base model when handling customer inquiries.skills-collection
Take a free 3-minute scan and get personalized AI skill recommendations.
Take free scan