Radar van Elk Solutions

Hackaday · Creativity & design

Running Large Language Models on Gaming PCs

New projects are enabling larger language models (LLMs) to run on more modest hardware, such as gaming PCs, rather than requiring extensive GPU setups.

The Strata project allows users to run a 125-billion-parameter LLM, like variants of the Qwen3.8 model, on hardware commonly found in gaming PCs. It utilizes VRAM as a cache for the large model, keeping frequently used components on the GPU while the rest resides in system RAM. A lookup table of approximately 29 GB is stored on the SSD.

The software employs speculative decoding to check multiple candidate tokens in a single pass. An RTX 5070 with 12 GB of VRAM can achieve speeds of 50 to 90 tokens per second, depending on quantization. Setup may involve compatibility issues with compilers and headers, but once configured, it loads in a few minutes.

Strata offers interaction via a web browser and exposes OpenAI- and Anthropic-compatible APIs for integration with existing tools. However, performance may not match larger models for all tasks. For instance, when queried about Hackaday, the model provided some correct information but also exhibited confusion regarding its founders and authors.

The model performed better on tasks like identifying problematic code or outlining compiler porting procedures. While "modest" hardware requirements include at least 32 GB of system RAM, 12 GB of VRAM, and 80 GB of storage, the project demonstrates how mixture-of-experts models and efficient memory management can extend the capabilities of ordinary PC hardware.

AI-samenvatting op basis van de bron.

Hackaday