Local AI
Private AI infrastructure, local inference, model serving, GPU planning, RAG, privacy, and operations for systems you control.
From Homelab Chatbot to a Private AI Operations Platform
Extend a local chatbot with document retrieval, bounded tools, monitoring, and backups only after defining permissions, ownership, failure behavior, and recovery; begin with read-only assistance.
Running Local AI Alongside Media Services Without Fighting Your GPU
Run Ollama, Tdarr, Plex, Jellyfin, and Frigate on the same homelab more safely by understanding NVENC, NVDEC, CUDA, VRAM, scheduling, and monitoring.
Performance Tuning Ollama and Local LLMs Without Guesswork
Tune Ollama from a repeatable baseline, checking model fit, context, memory placement, concurrency, and answer quality; change one setting at a time and keep a rollback path.
Model Routing: Different Models for Different Jobs
Compare local models on representative tasks before routing prompts automatically; enforce data restrictions first, make each route visible, and retain an allowed fallback and manual override.
Securing a Self-Hosted AI Server
Keep Ollama's API private, control Open WebUI accounts and tools, protect documents and logs, and test denied access before sharing a self-hosted AI service.
Open WebUI Deep Dive: Users, Models, Documents, Prompts, and Safe Sharing
Configure Open WebUI accounts, model and document access, provider connections, backups, and updates, with controlled signup and restricted tools before sharing an instance over LAN or VPN.
Building the Right PC for Local AI
A plain-English guide to choosing the right CPU, RAM, GPU, VRAM, storage, power supply, and cooling for a local AI homelab PC.
Choosing the Right Local AI Model: Gemma, Qwen, Llama, Mistral, DeepSeek, Nemotron
New to local AI? This beginner-friendly guide explains how to choose between popular open-weight model families like Gemma, Qwen, Llama, Mistral, DeepSeek, and Nemotron for your homelab.








