> Which open AI models fit on the hardware you own: 35 models against 16 machines by memory, with measured speeds and setup recipes from the people who ran them.

35 open models · 16 machines · speeds measured by the community

# Run it yourself

Pick the machine you own or could buy, and see the strongest open-weight models it can run, how compressed, and how fast people have measured them.

## Your machine · memory decides what fits

Mac

MacBook Air 16 GB

12 GB usable

MacBook Air 32 GB

24 GB usable

Mac 64 GB

48 GB usable

Mac 128 GB

96 GB usable

Mac Studio 256 GB

192 GB usable

Mac Studio 512 GB

384 GB usable

Mini PC

Strix Halo 128 GB

124 GB usable

Strix Halo 192 GB

180 GB usable

DGX Spark

120 GB usable

GPU

RTX 3090

24 GB usable

RTX 4090

24 GB usable

RTX 5090

32 GB usable

RTX PRO 6000

96 GB usable

Rig

4× RTX 3090

94 GB usable

4× RTX PRO 6000

380 GB usable

DGX Station

740 GB usable

64 GB Mac (M5 Pro: MacBook Pro or Mac mini) · price not confirmed · 307 GB/s memory bandwidth. Usable figure is an estimate (about 75%). The Mac mini M5 Pro starts at $1,699 with base memory and tops out at 64 GB; the 64 GB configuration price was not confirmed. [apple.com ↗](https://www.apple.com/macbook-pro/specs/)

48GB

memory a model can use

of 64 GB shared memory

152

strongest model that fits well

Qwen3.8-27B at 8-bit

14months

behind the closed frontier

GPT-5 reached this in Aug 2025

16of 35

models fit at some size

9 at 4-bit or better

## What fits · index score against memory needed · log scale

Fits at 4-bit or better

Fits only heavily compressed

Does not fit

_Chart: Open models by index score and memory needed, against Mac 64 GB_

Each model sits at the memory its 4-bit version needs. Scores are measured on the full-precision model; 4-bit and above stay close to that, and quality falls off below it. Sparse (mixture-of-experts) models must hold every weight in memory but only use a fraction per word, so they run much faster than their size suggests.

## Models for this machine · 16 fit in 48 GB

| MODEL | INDEX SCORE | SIZE · ACTIVE | FITS AS | MEASURED SPEED | GET IT |
| --- | --- | --- | --- | --- | --- |
| Qwen3.8-27B | 152 | 27B · all | 8-BIT33 GB | – | [lmstudio-community · MLX ↗](https://huggingface.co/lmstudio-community/Qwen3.8-27B-MLX-4bit) |
| Qwen3.6 27B | 145 | 27B · all | 8-BIT33 GB | – | [lmstudio-community · MLX ↗](https://huggingface.co/lmstudio-community/Qwen3.6-27B-MLX-4bit) |
| Qwen3.6-35B-A3B | 139 | 35B · 3B | 8-BIT42.2 GB | – | [lmstudio-community · MLX ↗](https://huggingface.co/lmstudio-community/Qwen3.6-35B-A3B-MLX-4bit) |
| Gemma 4 | 138 | 30.7B · all | 8-BIT37.2 GB | – | [lmstudio-community · MLX ↗](https://huggingface.co/lmstudio-community/gemma-4-31B-it-MLX-4bit) |
| Gemma 4 26B A4B | 136 | 25.2B · 3.8B | 8-BIT30.9 GB | – | [lmstudio-community · MLX ↗](https://huggingface.co/lmstudio-community/gemma-4-26B-A4B-it-MLX-4bit) |
| Qwen3.5-9B | 130 | 9B · all | 8-BIT12.3 GB | – | [lmstudio-community · MLX ↗](https://huggingface.co/lmstudio-community/Qwen3.5-9B-MLX-4bit) |
| gpt-oss-20b | 128 | 21B · 3.6B | 8-BIT26.1 GB | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/gpt-oss-20b-MXFP4-Q8) |
| Nemotron 3 Nano | not scored yet | 31.6B · 3B | 8-BIT38.3 GB | – | [lmstudio-community · MLX ↗](https://huggingface.co/lmstudio-community/NVIDIA-Nemotron-3-Nano-30B-A3B-MLX-4bit) |
| Muse Glimmer 30B | not scored yet | 29.6B · all | 8-BIT36 GB | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Muse-Glimmer-30B-4bit) |
| Qwen3.8-Flash-Next | ~169 | 125B · 6B | 2-BIT45.9 GB | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Qwen3.8-Flash-Next-4bit) |
| Nemotron 3 Super | ~132 | 120B · 12B | 2-BIT44.1 GB | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/NVIDIA-Nemotron-3-Super-120B-A12B-4bit) |
| gpt-oss-120b | 131 | 117B · 5.1B | 2-BIT43.1 GB | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/gpt-oss-120b-MXFP4-Q8) |
| Mistral Small 4 | ~129 | 119B · 6.5B | 2-BIT43.8 GB | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Mistral-Small-4-119B-2603-4bit) |
| Qwen3.5-122B-A10B | not in the index | 122B · 10B | 2-BIT44.8 GB | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Qwen3.5-122B-A10B-4bit) |
| Step 3.7 Flash | not scored yet | 198B · 11B | 1-BIT47.4 GB | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Step-3.7-Flash-4bit) |
| Devstral 2 | not scored yet | 123B · all | 2-BIT45.2 GB | – | [unsloth · GGUF ↗](https://huggingface.co/unsloth/Devstral-2-123B-Instruct-2512-GGUF) |
| MiMo-V2.6-Pro | ~178 | 1T · 42B | needs 236 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/MiMo-V2.6-Pro-RL-mxfp4-q8) |
| Kimi K3 | 172 | 2.8T · 104B | needs 645 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Kimi-K3-mlx-reap160-2bit) |
| Qwen3.8-2.4T-A95B | ~169 | 2.4T · 95B | needs 553 GB+ | – | [unsloth · GGUF ↗](https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF) |
| GLM-5.3 | 167 | 744B · 40B | needs 173 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/GLM-5.3-4bit) |
| DeepSeek-V4-Pro | 167 | 1.6T · 49B | needs 369 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/DeepSeek-V4-Pro-4bit) |
| MiMo-V2.6-Flash | ~167 | 309B · 15B | needs 72.9 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/MiMo-V2.6-Flash-RL-mxfp4-q8) |
| DeepSeek-V4.1-Flash | 165 | 552B · 16B | needs 129 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/DeepSeek-V4.1-Flash-MLX-4bit) |
| DeepSeek-V4-Flash | 165 | 284B · 13B | needs 67.2 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/DeepSeek-V4-Flash-0731-2.4bit-mixed) |
| GLM-5.2 | 158 | 744B · 40B | needs 173 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/GLM-5.2-mxfp4) |
| GLM-5.3-Flash | 158 | 320B · 18B | needs 75.4 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/GLM-5.3-Flash-4bit) |
| Kimi K2.6 | 156 | 1T · 32B | needs 232 GB+ | – | [unsloth · GGUF ↗](https://huggingface.co/unsloth/Kimi-K2.6-GGUF) |
| TM Inkling-Small | 154 | 276B · 12B | needs 65.3 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Inkling-Small-mxfp4) |
| TM Inkling | 150 | 975B · 41B | needs 226 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Inkling-mlx-4bit) |
| Qwen3.5 | 146 | 397B · 17B | needs 93.1 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Qwen3.5-397B-A17B-4bit) |
| MiniMax-M3 | 145 | 428B · 23B | needs 100 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/MiniMax-M3-4bit) |
| Nemotron 3 Ultra | 145 | 550B · 55B | needs 128 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Nemotron-3-Ultra-550B-A55B-4bit) |
| MiniMax-M2.7 | 143 | 229B · 10B | needs 54.6 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/MiniMax-M2.7-4bit) |
| LongCat 2.0 | ~140 | 1.6T · 48B | needs 369 GB+ | – | [original weights ↗](https://huggingface.co/meituan-longcat/LongCat-2.0) |
| Hy4 preview | not scored yet | 770B · 49B | needs 179 GB+ | – | [mlx-community · MLX ↗](https://huggingface.co/mlx-community/Hy4-preview-4bit) |

Memory is an estimate: parameters × bits per weight, plus 8% and 2 GB for the context. Speeds are only shown where someone published a measurement on this machine; nothing is calculated. Size · active: total parameters, and the share used for each word.

## Measured on this machine · 1 published results

66.3 tok/s

#### Llama 2 7B (reference test)

Q4_0 · llama.cpp (Metal) · reads 1,621 tok/s. Standard llama.cpp reference test: Llama 2 7B at Q4_0, prompt 512 (pp512) and generation 128 (tg128). MacBook Pro M5 Pro, 64 GB.

[hnfong](https://github.com/hnfong)

Aug 2026

[source ↗](https://github.com/ggml-org/llama.cpp/discussions/4167#discussioncomment-18213344)

## Recipes · how people set it up

Strix Halo llama.cpp toolboxes (Vulkan and ROCm containers)

Donato Capitella (kyuz0) · llama.cpp ↗

Strix Halo vLLM toolboxes

Donato Capitella (kyuz0) · vLLM ↗

vLLM Recipes: Qwen3.8-27B (RTX 5090 single and dual card, GB300)

vLLM project · vLLM ↗

Framework Desktop: DeepSeek V4 Flash 0731 UD-IQ2_XXS (90.9 GB) on 128 GB

Framework · llama.cpp ↗

Framework Desktop: Qwen3.5-122B-A10B Q3_K_S (52.5 GB) on 64 GB

Framework · llama.cpp ↗

Qwen3.8-27B NVFP4 serving recipe: full 256K context on a single RTX 5090

ayayalar · vLLM 0.27.1 ↗

MTP speculative decoding on Strix Halo (Qwen3.6 27B and 35B-A3B)

Donato Capitella (kyuz0) · llama.cpp ↗

Accelerating hybrid CPU and GPU inference in SGLang with KTransformers

KVCache.AI and Approaching AI · SGLang with KTransformers ↗

Performance of llama.cpp on NVIDIA DGX Spark, with setup guide

Georgi Gerganov (ggerganov) · llama.cpp ↗

## The people who make this possible · 21 projects and builders

Ahmad Osman

@TheAhmadOsman

r/LocalLLaMA moderator and founder of Osmantic, which builds the open-source ODS self-hosted AI stack. Long-running advocate for running open models on hardware you own; ran two workshops at the AI Engineer World's Fair 2026 and showed frontier open models on a DGX Station on stage. Handle confirmed from his GitHub profile and his own link page.

theahmadosman.world ↗

Georgi Gerganov

@ggerganov

Creator of llama.cpp and the ggml library; maintains the community benchmark threads for Apple Silicon, CUDA and DGX Spark. GitHub profile lists Hugging Face as his employer.

github.com ↗

llama.cpp / ggml-org

ggml-org

The C/C++ inference engine most local tools build on (Ollama, LM Studio, many GUIs). About 130,000 stars.

github.com ↗

Awni Hannun

@awnihannun

Co-created MLX, Apple's array framework for Apple silicon, and mlx-lm; first to publish 1-trillion-parameter models running on two Mac Studios. GitHub bio now says he does research at Anthropic.

github.com ↗

MLX / mlx-lm

ml-explore

Apple's open-source framework and LLM runner for Apple silicon; the engine behind most MLX quants on Hugging Face.

github.com ↗

Daniel Han and Michael Han (Unsloth)

@danielhanchen, @unslothai

Unsloth: dynamic GGUF quantisations that often appear on release day, plus run-it-locally guides for nearly every major open model.

unsloth.ai ↗

Donato Capitella (kyuz0)

kyuz0

Maintains the Strix Halo toolboxes and the local-llm-benchmarks site with reproducible speed measurements at different context depths.

strix-halo-toolboxes.com ↗

Alex Cheema

AlexCheema

Co-founder and CEO of EXO Labs; exo clusters Macs and other devices (including RDMA over Thunderbolt 5) to run models too large for one machine.

github.com ↗

exo labs

exo-explore

Open-source distributed inference project, about 47,000 stars, Apache 2.0.

github.com ↗

Jun Kim

@jundotkim

Creator of oMLX, an Apple-silicon inference server with continuous batching and SSD caching, and a community benchmark database of real Mac results.

github.com ↗

Bartowski

bartowski

One of the most-used community GGUF quantisation publishers, with over 2,400 model repos.

huggingface.co ↗

mradermacher team

mradermacher

Very large-scale GGUF and imatrix quantisation of community models (over 70,000 repos).

huggingface.co ↗

LM Studio community and MLX publishers

lmstudio-community, mlx-community

Ready-to-run GGUF and MLX conversions of new releases, usually within days of launch.

huggingface.co ↗

Kawrakow

ikawrakow

Author of ik_llama.cpp, a llama.cpp fork with state-of-the-art quantisation types and faster CPU and hybrid CPU+GPU inference for large MoE models.

github.com ↗

Johannes Gaessler

JohannesGaessler

Core llama.cpp CUDA and multi-GPU contributor.

github.com ↗

vLLM project (Woosuk Kwon)

@woosuk_k

vLLM, the standard high-throughput serving engine; its recipes site lists tested configurations for new models on 5090s, GB300s and more.

github.com ↗

SGLang (Lianmin Zheng and LMSYS)

@lm_zheng

SGLang serving framework; the LMSYS team also published the DGX Spark review and the KTransformers hybrid-inference work.

github.com ↗

KTransformers (KVCache.AI)

kvcache-ai

CPU+GPU hybrid inference kernels for running huge MoE models with modest GPU memory; integrated into SGLang.

github.com ↗

Ollama (Jeffrey Morgan and Michael Chiang)

@mchiang0610

The easiest way to run open models locally for most people; about 180,000 GitHub stars.

ollama.com ↗

LM Studio

lmstudio-ai

Desktop app for local models with GGUF and MLX engines.

lmstudio.ai ↗

Simon Willison

@simonw

Widely read writer who documents hands-on results with local and open models, and relayed early MLX multi-Mac results.

simonwillison.net ↗

Every speed and recipe on this page is their work, linked to where they published it. Engines, quantised files and benchmarks are all free and open source.

---
Source: https://themodelindex.org/local/ · The Model Index · data as of 2026-10-01
