🔧

Inference vs Fine-tuning — Using Someone Else's Model vs Making Your Own

pipeline() is ordering delivery, fine-tuning is taking a recipe and adjusting seasoning to your taste

Common Misconception: Fine-tuning = Using a Trained Model?

No. That's inference.

Inference:     Use someone else's model as-is
Fine-tuning:   Retrain someone else's model with your data → becomes YOUR model

Inference Fine-tuning
Analogy Ordering delivery Taking a recipe, adjusting seasoning
Coding npm install and use Fork, add your code, rebuild
Model Their model Your model (based on theirs)
Data Not needed Required (training material)

What We Did Before: Inference

classifier = pipeline("sentiment-analysis")
classifier("I love this!")  # Using someone else's completed model

What Fine-tuning Does: Making Your Model

You need 3 things:
1. Base model — already well-trained (DistilBERT)
2. Dataset — training material ("this sentence is positive, this is negative")
3. Training code — HuggingFace Trainer handles most of it

Where Do Datasets Come From?

HuggingFace Hub has datasets too, not just models.

huggingface.co/models    → 1M+ trained models
huggingface.co/datasets  → 200K+ training datasets  ← here

dataset = load_dataset("imdb")  # 25,000 movie reviews auto-downloaded

This isn't test data — it's training material. Teaching the model "this sentence is positive, this is negative."

Why "Fine" Tuning?

Not training from scratch. Slightly adjusting an already well-trained model.

Pre-training: blank → months on entire internet → base model (thousands of GPUs, millions $)
Fine-tuning:  base model → hours on your few thousand examples → your model (1 laptop GPU)

Full Fine-tuning Code

from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForSequenceClassification, Trainer, TrainingArguments

dataset = load_dataset("imdb")
tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased")
dataset = dataset.map(lambda x: tokenizer(x["text"], truncation=True, padding="max_length"), batched=True)
model = AutoModelForSequenceClassification.from_pretrained("distilbert-base-uncased", num_labels=2)
trainer = Trainer(model=model, train_dataset=dataset["train"].select(range(1000)),
                  args=TrainingArguments(output_dir="./my-model", num_train_epochs=3))
trainer.train()  # ← Your model is created here

my_pipe = pipeline("sentiment-analysis", model="./my-model")
print(my_pipe("Fantastic movie!"))

Why Sentiment Analysis Is Easiest for Fine-tuning

  • Data: 2-column CSV (text + label)

  • Output: single number (positive/negative)

  • Evaluation: accuracy (right or wrong)

  • Training: minutes, not hours

Key Concepts

1

Inference = use their model: immediate results with pipeline("sentiment-analysis")

2

Fine-tuning = make your model: base model + your data → retrain → your model

3

Hub has datasets too: load_dataset("imdb") downloads 25,000 movie reviews

4

Trainer.train() is the core — your model is created in this one line

5

Use your trained model for inference with pipeline(model="./my-model")

Use Cases

Sentiment fine-tuning intro — build positive/negative classifier with IMDB data Korean sentiment — NSMC (Naver movie reviews) + klue/bert-base for Korean model Custom data — any 2-column CSV (text, label) can become a custom classifier