Wink Pings

Running LLMs Locally: 90% of People Ignore This Parameter That Matters More Than VRAM

When building a local AI PC, everyone's screaming about grabbing GPUs and piling up VRAM, but almost no one clearly explains what actually core factor impacts user experience. A foreign developer tested 19 commercially available GPUs, and one enthusiast even pulled off a fully GPU-free solution, fitting an entire 405B-parameter large language model into system memory to get it running. This post compiles real test data and pitfalls to avoid, as a reference for anyone planning to build their own local AI machine.

Recently browsing overseas AI communities, I came across two really interesting projects for running local large language models. One completely shatters everyone's obsession with GPUs, the other pierces the marketing myth that "more VRAM is everything".

Let's start with the first one: an enthusiast built a local LLM host that doesn't use a GPU at all:

The entire build is extremely simple: an AMD EPYC 7000 series processor, 8 memory channels fully populated with 512GB of DDR5 ECC memory, using llama.cpp for CPU inference. No GPU, no CUDA required, and no need to scour the secondhand market fighting for a 3090. The goal was to fit an entire hundreds-of-billions-parameter LLM into system memory, no reliance on GPU chunked loading.

In the end, he successfully loaded the entire 405B-parameter DeepSeek V3 into memory and got it running. This approach is the complete opposite of the mainstream: instead of relying on GPU acceleration, the upgrade is purely to system memory capacity. For a similarly sized LLM to fit fully into GPU VRAM, you'd need to build an 8x H100 server, which costs at least $40,000.

Community discussions got pretty heated, and takes are interesting. Some people point out that fitting into memory doesn't mean it's usable, the speed must be glacial. Others note this is just a different tradeoff: sacrifice per-token generation speed to get to run a 400B-scale LLM at home, no cloud rental required, and all data stays local. Right now so many people are still waiting out GPU price hikes and shortages, and this guy just completely bypassed the GPU race, building a fully functional LLM machine with off-the-shelf server motherboard and memory.

Then there's the real-world testing by developer beamnxw, which directly called out the lie pushed by many marketing accounts: he tested 19 GPUs available for purchase on Amazon, and reached the conclusion that for local LLM inference, bandwidth is far more important than VRAM.

![This is an illustration depicting the data transfer process between VRAM and GPU. On the left is a rectangular box representing VRAM (Video Random Access Memory), and on the right is a square chip with multiple connectors representing the GPU (Graphics Processing Unit). The two are connected by an arrowed pipe labeled "bandwidth". Next to the GPU there is a small dot marked "tokens/sec", representing the number of tokens processed per second. At the bottom there is also a line of text: "faster pipe → faster tokens", meaning a faster pipe equals faster token processing. The entire background is light colored, decorated with faint coffee-stain-like patterns.](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHNIgsQfXMAAXzPd%3Fformat%3Djpg%26name%3Dlarge)

Everyone keeps saying "To run LLMs, first get enough VRAM", which isn't wrong—if the model can't fit into VRAM it definitely won't run. But almost no one talks about the fact that after fitting, what really determines whether your experience is laggy is bandwidth.

Every time an LLM generates one token, it has to read the entire model's weights from VRAM, then write the results back after calculation. In most cases, the GPU cores are just idle waiting for data. This problem is called the memory bottleneck, and it's the core reason local LLMs run slow.

In beamnxw's testing, running the same 13B-parameter model, a machine with 96GB unified memory was so slow it was unusable, while a card with only 24GB VRAM but double the bandwidth ran far more smoothly. The problem isn't insufficient core count, it's that the memory bus can't feed enough data to the GPU. He posted a set of real test data that makes this extremely clear:

![This is a chart of performance metrics for the 13B Q4 model. It shows read speed and per-token processing time at different bandwidth levels. From top to bottom: 936GB/s corresponds to ~19ms/token, 504GB/s corresponds to ~36ms/token, 256GB/s corresponds to ~70ms/token. A box on the left notes that the model uses ~0.7 bytes per parameter, and processing one token requires approximately 18GB of memory. On the right it mentions that under the same model and VRAM conditions, bandwidth increased by 3.7x.](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHNIhnb2XoAABDtt%3Fformat%3Djpg%26name%3Dlarge)

Same 13B Q4 model, same full VRAM utilization:

- 936 GB/s bandwidth → 19 milliseconds per token

- 504 GB/s bandwidth → 36 milliseconds per token

- 256 GB/s bandwidth → 70 milliseconds per token

The speed difference is nearly 4x, which translates to a night-and-day experience in real use: one is real-time conversation, the other is waiting forever for a single word.

Next, he sorted all 19 currently available GPUs with 16GB+ VRAM by bandwidth, compiled their prices, for regular users to reference directly:

|GPU| VRAM | Mem Type | Bus | Bandwidth | TDP | Price (USD)|

|---------------------------------|------|-----------|-----------|-----------|-------|-------------|

|H100 PCIe | 80GB | HBM3e | 5120-bit | 2039 GB/s | 700W | ~$25k-$40k|

|A100 PCIe 80GB | 80GB | HBM2e | 5120-bit | 1935 GB/s | 400W | ~$12k-$18k|

|RTX PRO 6000 Blackwell | 96GB | GDDR7 | 512-bit | 1792 GB/s | 300W | ~$12k-$13k|

|RTX 5090 | 32GB | GDDR7 | 512-bit | 1792 GB/s | 575W | ~$2k-$2.6k|

|RTX PRO 5000 (48GB) | 48GB | GDDR7 | 384-bit | 1344 GB/s | 300W | ~$6.8k-$7.4k|

|RTX PRO 5000 (72GB) | 72GB | GDDR7 | 384-bit | ~1344 GB/s| 300W | ~$9.9k|

|RTX 4090 / 3090 Ti | 24GB | GDDR6X | 384-bit | 1008 GB/s | 350W | ~$1.6k-$2.2k|

|RTX 6000 Ada | 48GB | GDDR6 | 384-bit | 960 GB/s | 300W | ~$7.3k|

|RX 7900 XTX | 24GB | GDDR6 | 384-bit | 960 GB/s | 355W | ~$800-$1k|

|RTX 5080 | 16GB | GDDR7 | 256-bit | 960 GB/s | 360W | ~$1.2k-$1.3k|

|RTX 3090 | 24GB | GDDR6X | 384-bit | 936 GB/s | 350W | ~$600-$1.05k|

|L40S / L40 | 48GB | GDDR6 ECC | 384-bit | 864 GB/s | 320W | ~$7k-$9k|

|RTX A6000 (Ampere) | 48GB | GDDR6 | 384-bit | 768 GB/s | 300W | ~$2.5k-$4.5k|

|RTX A5000 (24GB) | 24GB | GDDR6 | 384-bit | 768 GB/s | 250W | ~$1.3k-$2k|

|A40 | 48GB | GDDR6 | 384-bit | 696 GB/s | 250W | ~$3k-$4.5k|

|RTX PRO 4000 Blackwell | 24GB | GDDR7 ECC | 192-bit | 672 GB/s | 140W | ~$2.2k|

|Radeon AI PRO R9700 | 32GB | GDDR6 | 256-bit | 645 GB/s | 300W | ~$1.35k|

|RTX 4080 | 16GB | GDDR6X | 256-bit | 717 GB/s | 320W | ~$800-$900|

He also specifically called out two marketing traps that aren't worth buying:

The first trap is the AMD Ryzen AI Max+ 395 mini PC with 96GB unified memory. It's marketed as "AI ready", but the DDR5 memory bandwidth used by the integrated iGPU is only 256GB/s—less than a third of the RTX 3090 in the table above—and it costs $2,000 to $3,000. Even running a 13B model is too slow for daily use.

![This is a picture showing two GMKtec cases. Both have a modern, tech-forward design with a silver and black color scheme. The rear panel of the left case has multiple ports, including HDMI, USB, and a power button, for connecting external devices. The front panel of the right case has a similar port layout, plus a blue LED light strip for extra visual appeal. Overall, both cases look very professional and suitable for high-performance computing or gaming.](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHNIkpd6WUAAu6qQ%3Fformat%3Djpg%26name%3Dlarge)

The second trap is the NVIDIA DGX Spark, marketed as a personal AI supercomputer with 128GB LPDDR5X. It looks good but costs $4,000 to $4,700. Its actual bandwidth is less than 300GB/s, while a used $800 RTX 3090 has 3.4x the bandwidth, for a quarter of the price. It's only worth considering if you absolutely need to load models that won't fit in 24GB of VRAM; otherwise it's just an expensive paperweight.

![Image](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHNIloArXkAAOLRp%3Fformat%3Djpg%26name%3Dlarge)

Finally, he gave a few high cost-performance recommendations:

### RTX 3090 (used on eBay: $600-$1050)

24GB GDDR6X, 936GB/s bandwidth. Bandwidth is on par with the RTX 5080, it has more VRAM, and it costs a fraction of the price. As long as your model fits in 24GB, it's more than fast enough for daily use. Refurbished units on Amazon are a bit more expensive, but you can often find great deals on the secondhand market.

![Image](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHNIntgPXEAAUwij%3Fformat%3Djpg%26name%3Dlarge)

### RX 7900 XTX ($800-$1000)

AMD accidentally ended up with the best value card here: 960GB/s bandwidth, 24GB VRAM, full 384-bit bus. Bandwidth ties with the $10k-class RTX 6000 Ada, for a fraction of the price. ROCm works perfectly now in 2025, it runs out of the box with llama.cpp on Linux with no tinkering required. Two of these will set you back less than $2,000 total for 48GB of VRAM.

![Image](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHNIqBFGWcAE52Rf%3Fformat%3Djpg%26name%3Dlarge)

### Radeon AI PRO R9700 ($1350)

A workstation card with ECC, 32GB VRAM, 645GB/s bandwidth isn't the highest, but the capacity hits that sweet spot. A good option if you don't want to buy used and need stability, and drivers work perfectly these days.

![Image](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FHNIqYiOXgAA22lN%3Fformat%3Djpg%26name%3Dlarge)

To wrap up, here's a quick purchase guide matching your needs directly:

|Need|Recommended Purchase|

|---------------------------------|------------------------------------------|

|Money no object for maximum performance | RTX PRO 6000 Blackwell (~$13k)|

|Best value pick for regular users | Used RTX 3090 (~$600-$1,050)|

|Most bandwidth per dollar | RX 7900 XTX (~$800-$1k)|

|ECC and workstation stability | Radeon AI PRO R9700 or used RTX A5000|

|Want 48GB+ without paying the datacenter tax | Used RTX A6000 or dual RX 7900 XTX|

|Small form factor compact build | RTX PRO 4000 Blackwell (~$2.2k)|

Putting these two approaches side by side is really interesting: one bypasses GPUs entirely to stack memory for ultra-large models, the other tells you not to just look at VRAM when picking a GPU—look at bandwidth. At their core, both make the same point: don't get swept up by marketing hype, figure out your own needs and your core bottleneck before opening your wallet.

If you're building a local AI machine right now, would you prioritize stacking memory or a better GPU? Have you used any configuration that delivered way better value than you expected? Feel free to share in the comments.

发布时间: 2026-07-15 03:44