The Model Index
35 open models · 16 machines · speeds measured by the community

Run it yourself

Pick the machine you own or could buy, and see the strongest open-weight models it can run, how compressed, and how fast people have measured them.

Your machinememory decides what fits

Mac
Mini PC
GPU
Rig

64 GB Mac (M5 Pro: MacBook Pro or Mac mini) · price not confirmed · 307 GB/s memory bandwidth. Usable figure is an estimate (about 75%). The Mac mini M5 Pro starts at $1,699 with base memory and tops out at 64 GB; the 64 GB configuration price was not confirmed. apple.com ↗

48GBmemory a model can useof 64 GB shared memory
16of 35models fit at some size9 at 4-bit or better

What fitsindex score against memory needed · log scale

Fits at 4-bit or betterFits only heavily compressedDoes not fit

Each model sits at the memory its 4-bit version needs. Scores are measured on the full-precision model; 4-bit and above stay close to that, and quality falls off below it. Sparse (mixture-of-experts) models must hold every weight in memory but only use a fraction per word, so they run much faster than their size suggests.

Models for this machine16 fit in 48 GB

ModelIndex scoreSize · activeFits asMeasured speedGet it
Qwen3.8-27B
152
27B · all8-bit33 GB–lmstudio-community · MLX ↗
Qwen3.6 27B
145
27B · all8-bit33 GB–lmstudio-community · MLX ↗
Qwen3.6-35B-A3B
139
35B · 3B8-bit42.2 GB–lmstudio-community · MLX ↗
Gemma 4
138
30.7B · all8-bit37.2 GB–lmstudio-community · MLX ↗
Gemma 4 26B A4B
136
25.2B · 3.8B8-bit30.9 GB–lmstudio-community · MLX ↗
Qwen3.5-9B
130
9B · all8-bit12.3 GB–lmstudio-community · MLX ↗
gpt-oss-20b
128
21B · 3.6B8-bit26.1 GB–mlx-community · MLX ↗
Nemotron 3 Nano
not scored yet
31.6B · 3B8-bit38.3 GB–lmstudio-community · MLX ↗
Muse Glimmer 30B
not scored yet
29.6B · all8-bit36 GB–mlx-community · MLX ↗
Qwen3.8-Flash-Next
~169
125B · 6B2-bit45.9 GB–mlx-community · MLX ↗
Nemotron 3 Super
~132
120B · 12B2-bit44.1 GB–mlx-community · MLX ↗
gpt-oss-120b
131
117B · 5.1B2-bit43.1 GB–mlx-community · MLX ↗
Mistral Small 4
~129
119B · 6.5B2-bit43.8 GB–mlx-community · MLX ↗
Qwen3.5-122B-A10Bnot in the index122B · 10B2-bit44.8 GB–mlx-community · MLX ↗
Step 3.7 Flash
not scored yet
198B · 11B1-bit47.4 GB–mlx-community · MLX ↗
Devstral 2
not scored yet
123B · all2-bit45.2 GB–unsloth · GGUF ↗
MiMo-V2.6-Pro
~178
1T · 42Bneeds 236 GB+–mlx-community · MLX ↗
Kimi K3
172
2.8T · 104Bneeds 645 GB+–mlx-community · MLX ↗
Qwen3.8-2.4T-A95B
~169
2.4T · 95Bneeds 553 GB+–unsloth · GGUF ↗
GLM-5.3
167
744B · 40Bneeds 173 GB+–mlx-community · MLX ↗
DeepSeek-V4-Pro
167
1.6T · 49Bneeds 369 GB+–mlx-community · MLX ↗
MiMo-V2.6-Flash
~167
309B · 15Bneeds 72.9 GB+–mlx-community · MLX ↗
DeepSeek-V4.1-Flash
165
552B · 16Bneeds 129 GB+–mlx-community · MLX ↗
DeepSeek-V4-Flash
165
284B · 13Bneeds 67.2 GB+–mlx-community · MLX ↗
GLM-5.2
158
744B · 40Bneeds 173 GB+–mlx-community · MLX ↗
GLM-5.3-Flash
158
320B · 18Bneeds 75.4 GB+–mlx-community · MLX ↗
Kimi K2.6
156
1T · 32Bneeds 232 GB+–unsloth · GGUF ↗
Inkling-Small
154
276B · 12Bneeds 65.3 GB+–mlx-community · MLX ↗
Inkling
150
975B · 41Bneeds 226 GB+–mlx-community · MLX ↗
Qwen3.5
146
397B · 17Bneeds 93.1 GB+–mlx-community · MLX ↗
MiniMax-M3
145
428B · 23Bneeds 100 GB+–mlx-community · MLX ↗
Nemotron 3 Ultra
145
550B · 55Bneeds 128 GB+–mlx-community · MLX ↗
MiniMax-M2.7
143
229B · 10Bneeds 54.6 GB+–mlx-community · MLX ↗
LongCat 2.0
~140
1.6T · 48Bneeds 369 GB+–original weights ↗
Hy4 preview
not scored yet
770B · 49Bneeds 179 GB+–mlx-community · MLX ↗

Memory is an estimate: parameters × bits per weight, plus 8% and 2 GB for the context. Speeds are only shown where someone published a measurement on this machine; nothing is calculated. Size · active: total parameters, and the share used for each word.

Measured on this machine1 published results

Llama 2 7B (reference test)

Q4_0 · llama.cpp (Metal) · reads 1,621 tok/s. Standard llama.cpp reference test: Llama 2 7B at Q4_0, prompt 512 (pp512) and generation 128 (tg128). MacBook Pro M5 Pro, 64 GB.

Recipeshow people set it up

The people who make this possible21 projects and builders

Ahmad Osman@TheAhmadOsman

r/LocalLLaMA moderator and founder of Osmantic, which builds the open-source ODS self-hosted AI stack. Long-running advocate for running open models on hardware you own; ran two workshops at the AI Engineer World's Fair 2026 and showed frontier open models on a DGX Station on stage. Handle confirmed from his GitHub profile and his own link page.

theahmadosman.world ↗
Georgi Gerganov@ggerganov

Creator of llama.cpp and the ggml library; maintains the community benchmark threads for Apple Silicon, CUDA and DGX Spark. GitHub profile lists Hugging Face as his employer.

github.com ↗
llama.cpp / ggml-orgggml-org

The C/C++ inference engine most local tools build on (Ollama, LM Studio, many GUIs). About 130,000 stars.

github.com ↗
Awni Hannun@awnihannun

Co-created MLX, Apple's array framework for Apple silicon, and mlx-lm; first to publish 1-trillion-parameter models running on two Mac Studios. GitHub bio now says he does research at Anthropic.

github.com ↗
MLX / mlx-lmml-explore

Apple's open-source framework and LLM runner for Apple silicon; the engine behind most MLX quants on Hugging Face.

github.com ↗
Daniel Han and Michael Han (Unsloth)@danielhanchen, @unslothai

Unsloth: dynamic GGUF quantisations that often appear on release day, plus run-it-locally guides for nearly every major open model.

unsloth.ai ↗
Donato Capitella (kyuz0)kyuz0

Maintains the Strix Halo toolboxes and the local-llm-benchmarks site with reproducible speed measurements at different context depths.

strix-halo-toolboxes.com ↗
Alex CheemaAlexCheema

Co-founder and CEO of EXO Labs; exo clusters Macs and other devices (including RDMA over Thunderbolt 5) to run models too large for one machine.

github.com ↗
exo labsexo-explore

Open-source distributed inference project, about 47,000 stars, Apache 2.0.

github.com ↗
Jun Kim@jundotkim

Creator of oMLX, an Apple-silicon inference server with continuous batching and SSD caching, and a community benchmark database of real Mac results.

github.com ↗
Bartowskibartowski

One of the most-used community GGUF quantisation publishers, with over 2,400 model repos.

huggingface.co ↗
mradermacher teammradermacher

Very large-scale GGUF and imatrix quantisation of community models (over 70,000 repos).

huggingface.co ↗
LM Studio community and MLX publisherslmstudio-community, mlx-community

Ready-to-run GGUF and MLX conversions of new releases, usually within days of launch.

huggingface.co ↗
Kawrakowikawrakow

Author of ik_llama.cpp, a llama.cpp fork with state-of-the-art quantisation types and faster CPU and hybrid CPU+GPU inference for large MoE models.

github.com ↗
Johannes GaesslerJohannesGaessler

Core llama.cpp CUDA and multi-GPU contributor.

github.com ↗
vLLM project (Woosuk Kwon)@woosuk_k

vLLM, the standard high-throughput serving engine; its recipes site lists tested configurations for new models on 5090s, GB300s and more.

github.com ↗
SGLang (Lianmin Zheng and LMSYS)@lm_zheng

SGLang serving framework; the LMSYS team also published the DGX Spark review and the KTransformers hybrid-inference work.

github.com ↗
KTransformers (KVCache.AI)kvcache-ai

CPU+GPU hybrid inference kernels for running huge MoE models with modest GPU memory; integrated into SGLang.

github.com ↗
Ollama (Jeffrey Morgan and Michael Chiang)@mchiang0610

The easiest way to run open models locally for most people; about 180,000 GitHub stars.

ollama.com ↗
LM Studiolmstudio-ai

Desktop app for local models with GGUF and MLX engines.

lmstudio.ai ↗
Simon Willison@simonw

Widely read writer who documents hands-on results with local and open models, and relayed early MLX multi-Mac results.

simonwillison.net ↗

Every speed and recipe on this page is their work, linked to where they published it. Engines, quantised files and benchmarks are all free and open source.