| A | B | C | D | E | F | G | H | I | J | K | L | M | N | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1 | Timestamp | Real-world Usability / Task Suitability (Is it actually usable for real tasks?) | Model Name + Size + Quantization (e.g., Llama-3 8B Q4_K_M) | Runtime / Stack (e.g., llama.cpp, MLX, Ollama, LM Studio) | Hardware (Chip + RAM, e.g., M2 Max 64GB) | Throughput (Tokens/sec) | Latency / Response Feel | Practical Context Window Limit | ||||||
2 | 7/25/2026 21:21:16 | gpt4 | gpt5 4b q4 | llama | n/a | 2300 | n/a | n/a | ||||||
3 | 3/20/2026 | Yes - excellent for coding, chat, and general tasks | Llama 3.3 70B Q4_K_M | Ollama / llama.cpp | M4 Max 128GB (40 GPU cores, 546 GB/s) | ~8-10 t/s generation | Noticeable pause before response, but usable for interactive chat | ~8K practical (supports 128K but slows significantly beyond 8K) | ||||||
4 | 3/20/2026 | Yes - fast and capable for most tasks | Qwen 2.5 32B Q4_K_M | Ollama / llama.cpp | M4 Max 64GB (40 GPU cores, 546 GB/s) | ~25-30 t/s generation | Fast and responsive, feels near-instant | ~16K practical (supports 128K but quality degrades beyond 16K) | ||||||
5 | 3/20/2026 | Yes - best small model for coding and chat | Mistral Small 24B Instruct Q4_K_M | Ollama / llama.cpp | M3 Max 36GB (30 GPU cores, 300 GB/s) | ~18 t/s generation | Very responsive, great interactive experience | ~8K practical (supports 32K context) | ||||||
6 | 3/20/2026 | Yes - excellent for quick tasks, summarization, and chat | Phi-3 Mini 3.8B Q4_K_M | Ollama / llama.cpp | M1 MacBook Air 8GB (8 GPU cores, 68 GB/s) | ~14 t/s generation | Very fast, feels instant | ~4K practical (supports 4K-128K depending on version) | ||||||
7 | 3/20/2026 | Yes - strong for coding and reasoning tasks | Qwen 2.5 Coder 32B Q4_K_M | LM Studio / llama.cpp | M4 Max 64GB (40 GPU cores, 546 GB/s) | ~25-30 t/s generation | Fast, smooth streaming output | ~16K practical (supports 128K) | ||||||
8 | 3/20/2026 | Yes - good for general chat and lightweight tasks | Llama 3.2 8B Q4_K_M | Ollama | M2 MacBook Air 16GB (10 GPU cores, 100 GB/s) | ~22 t/s generation | Fast and responsive | ~8K practical (supports 128K) | ||||||
9 | 3/20/2026 | Yes - best balance of quality and speed for M1 Macs | Gemma 2 9B Q4_K_M | MLX / LM Studio | M1 Pro 16GB (16 GPU cores, 200 GB/s) | ~20-25 t/s generation | Smooth streaming, low latency | ~8K practical (supports 8K context) | ||||||
10 | 3/20/2026 | Yes - top-tier coding model, highly recommended | DeepSeek Coder V2 Lite 16B Q4_K_M | Ollama / llama.cpp | M2 Max 32GB (38 GPU cores, 400 GB/s) | ~30-35 t/s generation | Very fast, excellent for interactive coding | ~16K practical (supports 128K) | ||||||
11 | 3/20/2026 | Marginally - usable for simple tasks but slow for complex work | Llama 3.1 70B Q4_K_M | Ollama / llama.cpp | M1 Max 64GB (32 GPU cores, 400 GB/s) | ~5-6 t/s generation | Slow, noticeable waiting between tokens | ~4K practical (context heavily impacts speed) | ||||||
12 | 3/20/2026 | Yes - great speed/quality for everyday use | Mistral 7B Instruct v0.3 Q5_K_M | MLX | M3 MacBook Pro 18GB (10 GPU cores, 150 GB/s) | ~35-40 t/s generation | Instant feeling, excellent for interactive use | ~8K practical (supports 32K context) | ||||||
13 | 3/20/2026 | Yes - fastest large model experience on Apple Silicon | Qwen 2.5 72B Q4_K_M | Ollama / llama.cpp | M4 Max 128GB (40 GPU cores, 546 GB/s) | ~10-12 t/s generation | Moderate latency, comfortable for chat but not instant | ~8K practical (supports 128K but heavily impacts speed) | ||||||
14 | 3/20/2026 | Yes - great for summarization, writing, and RAG tasks | Llama 3.1 8B Instruct Q5_K_M | MLX | M2 Pro 16GB (19 GPU cores, 200 GB/s) | ~35-40 t/s generation | Very fast, feels instant | ~8K practical (supports 128K) | ||||||
15 | 3/20/2026 | No - too slow for interactive use, batch only | Llama 3.1 70B Q4_K_M | Ollama | M2 MacBook Air 24GB (10 GPU cores, 100 GB/s) | ~1-2 t/s (model spills to swap) | Extremely slow, painful wait times | ~2K practical (RAM insufficient, heavy swapping) | ||||||
16 | 3/20/2026 | Yes - excellent for creative writing, roleplay, and general chat | Gemma 2 27B Q4_K_M | Ollama / llama.cpp | M3 Max 36GB (40 GPU cores, 400 GB/s) | ~15-18 t/s generation | Good responsiveness, slight initial delay | ~8K practical (supports 8K context) | ||||||
17 | 3/20/2026 | Yes - smallest viable model for basic tasks on low-RAM Macs | Qwen 2.5 3B Instruct Q4_K_M | Ollama | M1 MacBook Air 8GB (7 GPU cores, 68 GB/s) | ~14-16 t/s generation | Fast responses, limited task capability | ~4K practical (supports 32K) | ||||||
18 | 3/20/2026 | Yes - MLX delivers native Apple Silicon performance | Llama 3.1 8B Instruct 4-bit MLX | MLX (mlx-community) | M3 MacBook Pro 18GB (10 GPU cores, 150 GB/s) | ~40-50 t/s generation | Near-instant, best-in-class latency for small models | ~8K practical (supports 128K) | ||||||
19 | 3/20/2026 | Yes - competitive with GPT-3.5 for many tasks | Mistral Small 24B Instruct Q5_K_M | LM Studio | M4 Pro 48GB (20 GPU cores, 273 GB/s) | ~20-25 t/s generation | Smooth and fast, great for sustained conversations | ~16K practical (supports 32K context) | ||||||
20 | 3/20/2026 | Yes - strong for code generation and reasoning | Qwen 2.5 14B Instruct Q4_K_M | Ollama / llama.cpp | M2 Pro 32GB (19 GPU cores, 200 GB/s) | ~22-28 t/s generation | Fast and interactive | ~8K practical (supports 128K) | ||||||
21 | 3/20/2026 | Yes - surprisingly capable for its size | Phi-4 14B Q4_K_M | Ollama / llama.cpp | M3 Pro 18GB (14 GPU cores, 150 GB/s) | ~20-25 t/s generation | Fast, fluid streaming | ~8K practical (supports 16K context) | ||||||
22 | 3/20/2026 | Yes - best ultra-large model experience on Mac | Qwen 2.5 72B Q4_K_M | MLX | M2 Ultra 192GB (76 GPU cores, 800 GB/s) | ~12-15 t/s generation | Reasonable for interactive use, slight prompt processing delay | ~16K practical (supports 128K) | ||||||
23 | ||||||||||||||
24 | ||||||||||||||
25 | ||||||||||||||
26 | ||||||||||||||
27 | ||||||||||||||
28 | ||||||||||||||
29 | ||||||||||||||
30 | ||||||||||||||
31 | ||||||||||||||
32 | ||||||||||||||
33 | ||||||||||||||
34 | ||||||||||||||
35 | ||||||||||||||
36 | ||||||||||||||
37 | ||||||||||||||
38 | ||||||||||||||
39 | ||||||||||||||
40 | ||||||||||||||
41 | ||||||||||||||
42 | ||||||||||||||
43 | ||||||||||||||
44 | ||||||||||||||
45 | ||||||||||||||
46 | ||||||||||||||
47 | ||||||||||||||
48 | ||||||||||||||
49 | ||||||||||||||
50 | ||||||||||||||
51 | ||||||||||||||
52 | ||||||||||||||
53 | ||||||||||||||
54 | ||||||||||||||
55 | ||||||||||||||
56 | ||||||||||||||
57 | ||||||||||||||
58 | ||||||||||||||
59 | ||||||||||||||
60 | ||||||||||||||
61 | ||||||||||||||
62 | ||||||||||||||
63 | ||||||||||||||
64 | ||||||||||||||
65 | ||||||||||||||
66 | ||||||||||||||
67 | ||||||||||||||
68 | ||||||||||||||
69 | ||||||||||||||
70 | ||||||||||||||
71 | ||||||||||||||
72 | ||||||||||||||
73 | ||||||||||||||
74 | ||||||||||||||
75 | ||||||||||||||
76 | ||||||||||||||
77 | ||||||||||||||
78 | ||||||||||||||
79 | ||||||||||||||
80 | ||||||||||||||
81 | ||||||||||||||
82 | ||||||||||||||
83 | ||||||||||||||
84 | ||||||||||||||
85 | ||||||||||||||
86 | ||||||||||||||
87 | ||||||||||||||
88 | ||||||||||||||
89 | ||||||||||||||
90 | ||||||||||||||
91 | ||||||||||||||
92 | ||||||||||||||
93 | ||||||||||||||
94 | ||||||||||||||
95 | ||||||||||||||
96 | ||||||||||||||
97 | ||||||||||||||
98 | ||||||||||||||
99 | ||||||||||||||
100 |