Learn Large Language Models from first principles — from attention math to production deployment.
┌─────────────────────────────────────────────────────────────────────────────┐
│ LLM DEEPER LEARNING PATH │
├─────────────────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ BASICS │───▶│ HFCORE │───▶│ INFERENCE │ │
│ │ │ │ │ │ & ADVANCED │ │
│ │ • NumPy │ │ • TinyGPT │ │ │ │
│ │ • Attention │ │ • HuggingFace│ │ • MLX / vLLM │ │
│ │ • Positional │ │ • LoRA │ │ • Tool Call │ │
│ │ Encoding │ │ • Fine-tune │ │ • Deployment │ │
│ │ • Multi-Head │ │ • Evaluation │ │ │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
│ │
│ Week 1 Week 2 Week 3+ │
│ Foundations Training Production │
└─────────────────────────────────────────────────────────────────────────────┘
This repository provides a hands-on learning path for understanding LLMs:
| Module | What You'll Learn |
|---|---|
| Basics | NumPy, attention math, positional encoding, multi-head attention |
| HFCore | Build TinyGPT from scratch, HuggingFace, LoRA, fine-tuning, evaluation |
| Inference & Advanced | MLX (Apple Silicon), vLLM (CUDA), tool calling, production deployment |
- Python 3.12+
- uv (recommended) or pip
# Clone the repository
git clone https://github.com/your-username/llm_deeper.git
cd llm_deeper
# Install dependencies with uv
uv sync
# Or with pip
pip install -e .| Platform | Requirements | Notes |
|---|---|---|
| CPU | Any modern CPU | Slow, but works for learning |
| Apple Silicon | M1/M2/M3/M4 Mac | Use MLX for fast inference |
| NVIDIA GPU | 8GB+ VRAM | Use vLLM for production |
uv run python src/basics/03_attention_math.pyuv run python src/hfcore/00_tinygpt.pyuv run python src/hfcore/04_dataset_plus_finetune_lora.pyuv run python src/inferencing_and_advanced/02_mlx_lora.pyllm_deeper/
├── src/
│ ├── basics/ # Foundational concepts
│ │ ├── 00_pandas.py # Data manipulation
│ │ ├── 01_numpy.py # Tensor operations
│ │ ├── 02_math.py # Mathematical foundations
│ │ ├── 03_attention_math.py # Attention step-by-step
│ │ ├── 04_attention_impl.py # Full attention implementation
│ │ ├── 05_attention_positional_encoding.py
│ │ ├── 06_mha.py # Multi-head attention
│ │ └── 07_conclusion.py # Putting it together
│ │
│ ├── hfcore/ # HuggingFace training pipeline
│ │ ├── 00_tinygpt.py # Build GPT from scratch
│ │ ├── 01_transformers.py # HuggingFace basics
│ │ ├── 02_peft_lora.py # LoRA concepts
│ │ ├── 03_dataset_plus_finetune_full.py
│ │ ├── 04_dataset_plus_finetune_lora.py
│ │ ├── 05_finetuned_model_inference.py
│ │ └── 06_evals.py # Evaluation metrics
│ │
│ └── inferencing_and_advanced/ # Production inference
│ ├── 00_base.py
│ ├── 01_peft_to_mlx.py # Convert PEFT → MLX
│ ├── 02_mlx_lora.py # MLX inference
│ ├── 03_server.py # FastAPI server
│ ├── 04_tool_calling_basic.py
│ ├── 05_tool_calling_advanced.py
│ ├── 07_vllm_cuda_plus_tools.py
│ ├── 08_deepseek_example.py # DeepSeek R1 reasoning
│ ├── 09_mistral_example.py # Large model w/ FP8
│ └── 10_diffuser_image.py # Image generation (SD3.5)
│
├── docs/ # Detailed documentation
│ ├── 00_overview.md # Project overview
│ ├── 01_basics.md # Foundations
│ ├── 02_hfcore.md # Training pipeline
│ ├── 03_inferencing.md # Inference
│ ├── 04_tool_calling.md # Function calling
│ └── 05_deployment.md # Production deployment
│
├── output_models/ # Saved models & adapters
├── pyproject.toml
└── README.md
| Document | Description |
|---|---|
| 00_overview.md | Learning path, prerequisites, LLM concepts |
| 01_basics.md | Tensors, attention, positional encoding, MHA |
| 02_hfcore.md | Tokenization, LoRA, fine-tuning, evaluation |
| 03_inferencing.md | Model formats, quantization, MLX, servers |
| 04_tool_calling.md | Function calling, agents, multi-turn |
| 05_deployment.md | vLLM, GPU optimization, cloud deployment |
Attention(Q, K, V) = softmax(QK^T / √d_k) × V
The attention mechanism allows tokens to "look at" other tokens and decide how much to focus on each one.
Instead of fine-tuning all ~600M parameters, LoRA trains only ~3M parameters (0.5%) by adding small adapter matrices:
W' = W + BA (where B is d×r and A is r×d, with r << d)
During generation, transformers compute Key and Value vectors for every token. The KV cache stores these to avoid recomputation:
Without cache: O(n²) per token
With cache: O(n) per token
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM
from trl import SFTTrainer
# Load base model
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-0.6B")
# Apply LoRA
lora_config = LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"])
model = get_peft_model(model, lora_config)
# Train with SFTTrainer
trainer = SFTTrainer(model=model, train_dataset=dataset, ...)
trainer.train()
# Save adapter (only 12MB vs 1.2GB full model)
model.save_pretrained("./medical-lora-adapter")uv run python -m mlx_lm.server --model ./merged-model --port 8000vllm serve Qwen/Qwen3-0.6B --enable-lora --lora-modules medical=./adapter# R1 models use <think>...</think> tokens for chain-of-thought
from vllm import LLM, SamplingParams
llm = LLM(model="deepseek-ai/DeepSeek-R1-Distill-Qwen-14B", enforce_eager=True)
outputs = llm.generate(["Explain quantum computing"], SamplingParams(max_tokens=3000))# 24B model on 40GB GPU using FP8
llm = LLM(
model="mistralai/Devstral-Small-2-24B-Instruct-2512",
quantization="fp8", # 50% memory savings
max_model_len=4096,
)curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "medical", "messages": [{"role": "user", "content": "What are symptoms of diabetes?"}]}'Contributions are welcome! Please feel free to submit a Pull Request.
MIT License - see LICENSE for details.
- HuggingFace for Transformers, PEFT, TRL
- Apple MLX for Apple Silicon support
- vLLM for production inference
- Qwen for excellent open-source models