Hugging Face
Hugging Face is the central platform and ecosystem for open-source machine learning, hosting models, datasets, Spaces, and foundational libraries like Transformers, Accelerate, and PEFT.
When to Use
- Pretrained Model Discovery & Inference: Utilizing thousands of open-source vision, NLP, and audio models via
transformers. - Parameter-Efficient Fine-Tuning (PEFT): Fine-tuning LLMs with LoRA / QLoRA using
peftandbitsandbytes. - Dataset Management & Streaming: Accessing multi-gigabyte datasets without memory exhaustion via
datasets. - Production Model Deployment: Deploying containerized endpoints with Hugging Face Text Generation Inference (TGI).
Quick Start
from transformers import pipeline
# One-line inference with modern pre-trained models
classifier = pipeline("sentiment-analysis", model="distilbert-base-uncased-finetuned-sst-2-english")
result = classifier("This framework makes NLP development seamless and fast!")
print(result) # [{'label': 'POSITIVE', 'score': 0.9998}]
Core Concepts
Inference Pipeline with Transformers & PyTorch
Loading state-of-the-art transformer models with automatic tokenization:
from transformers import pipeline
import torch
# Create accelerated text-classification pipeline
classifier = pipeline(
task="text-classification",
model="distilbert/distilbert-base-uncased-finetuned-sst-2-english",
device=0 if torch.cuda.is_available() else -1
)
results = classifier([
"This new AI development platform accelerates productivity significantly.",
"The legacy system crashed repeatedly under moderate load."
])
for res in results:
print(f"Label: {res['label']}, Score: {res['score']:.4f}")
4-Bit Model Quantization with BitsAndBytes
Loading large LLMs on consumer hardware using 4-bit NormalFloat (NF4):
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
model_id = "meta-llama/Llama-3.2-3B-Instruct"
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=bnb_config,
device_map="auto"
)
inputs = tokenizer("Explain quantum computing in one sentence:", return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Streaming Massive Datasets with Datasets Library
Iterating over terabyte-scale datasets without downloading all files upfront:
from datasets import load_dataset
# Stream dataset lazily without full disk download
dataset = load_dataset('allenai/c4', 'en', split='train', streaming=True)
# Iterate through samples
for i, sample in enumerate(dataset.take(3)):
print(f"Sample {i+1}: {sample['text'][:100]}...\n")
Common Patterns
Quantized Model Loading with BitsAndBytes (4-bit / 8-bit)
Problem: 7B-70B parameter models exceed single-GPU VRAM limits in FP16 precision.
Solution:
Load with 4-bit NF4 quantization using BitsAndBytesConfig:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16
)
model_id = "meta-llama/Llama-3.2-3B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=bnb_config,
device_map="auto"
)
Best Practices
Do:
- Use
BitsAndBytesConfig(nf4) or AWQ quantization to run open-weight LLMs with low memory consumption. - Set
device_map="auto"when loading large models to distribute layers automatically across available GPUs and RAM. - Use
datasets.load_dataset(..., streaming=True)for datasets that exceed local storage capacity. - Pass
torch_dtype=torch.bfloat16when running inference on modern Ampere/Hopper GPU architectures.
Don't:
- Load entire models onto CPU before moving to GPU; instantiate directly with
device_map. - Hardcode Hugging Face access tokens in code; authenticate via
huggingface-cli loginorHF_TOKEN. - Run unquantized 70B+ parameter models on consumer hardware; use quantized GGUF, EXL2, or 4-bit NF4.
Troubleshooting
| Error | Cause | Solution |
| :-------------------------------------- | :--------------------------------------------------------------------------- | :--------------------------------------------------------------------------- |
| Gated repo error / 403 Forbidden | Accessing gated model (e.g. Llama/Mistral) without accepting license on Hub. | Accept license on HuggingFace and run huggingface-cli login. |
| CUDA Out of Memory in from_pretrained | Model weights exceed GPU memory in FP32. | Use device_map="auto", torch_dtype=torch.float16, or 4-bit quantization. |
| KeyError in tokenizer output | Mismatched tokenizer and model configurations. | Always load tokenizer and model from the identical model_id. |
微信扫一扫