Ollama, LM Studio & Hugging Face Are Eating Your Disk: How to Manage AI Models on Mac
The Silo Team · 8 min read · September 25, 2026
The AI Model Storage Problem
Running local LLMs on Apple Silicon is genuinely useful. But it comes with a storage cost that sneaks up on you. A typical developer's Mac after 6 months of AI experimentation might contain:
~/.ollama/models/— 15–40 GB of pulled Ollama models~/.cache/huggingface/hub/— 20–60 GB of HuggingFace model shards~/Library/Application Support/LM Studio/models/— another 10–30 GB~/mlx-models/— Apple MLX quantized weights~/ComfyUI/models/— image generation model checkpoints
Total: easily 50–150 GB of model weights, often with significant duplication between hubs.
Cross-Hub Duplication is Common
Ollama's llama3.2:3b and the HuggingFace meta-llama/Llama-3.2-3B-Instruct GGUF Q4_K_M are often byte-for-byte identical. Downloading through different interfaces creates silent duplicates.
Parsing GGUF Headers to Identify Models
A .gguf file begins with a 4-byte magic number (GGUF in ASCII) followed by a structured header containing:
- Model architecture (LLaMA, Mistral, Gemma, Qwen, DeepSeek, Phi, etc.)
- Context length
- Number of layers and heads
- Quantization type (Q4_K_M, Q8_0, F16, etc.)
- Vocabulary size
- Model author and name from metadata
MacPilot parses these headers directly — without loading the full model into memory — to give you meaningful names and sizes rather than opaque blob filenames.
Finding Duplicate Models
MacPilot computes SHA-256 hashes of GGUF files across all detected model directories and surfaces duplicates with their sizes. You can then use APFS clonefile(2) to replace the duplicate with a block-sharing clone, or simply delete it from the hub you use less.
Which Models Are Safe to Delete?
- Safe to delete: Old quantization versions (Q4_0 when you now have Q4_K_M of the same model), experimental pulls you tested once, superseded model versions.
- Keep: Models actively used in workflows, fine-tuned or LoRA-adapted weights (not re-downloadable), models pulled with specific system prompts stored alongside them.
MacPilot's AI Model Registry shows last-accessed timestamps to help you identify unused models.
Keep reading
How to Push Background Tasks to Apple Silicon E-Cores for Better Mac Performance
Apple Silicon has separate Performance (P) and Efficiency (E) cores. Using POSIX QoS APIs, you can deprioritize background processes to E-cores and reclaim P-core headroom for your primary work.
"Your System Has Run Out of Application Memory" on Mac: Fix It Permanently
What the memory pressure alert actually means on Apple Silicon, which processes to terminate first, and how to prevent it from coming back.