Quantization

Homelab
Performance Tuning Ollama and Local LLMs Without Guesswork

Make Ollama and other local LLM tools feel faster by matching the model, context length, quantization, GPU offload, and logs to the hardware you actually have.

Read more
AI
Local AI Hardware Sizing: CPU, NPU, GPU, RAM, and VRAM

A local AI hardware sizing guide explaining model memory, quantization, KV cache, CPU-only use, NPUs, GPUs, RAM, VRAM, and upgrade tiers.

Read more