Local AI
Ollama vs llama.cpp vs LM Studio vs Lemonade: Test the Runtime, Not the LogoNew!!
A local-AI runtime comparison framework for Ollama, llama.cpp, LM Studio, and Lemonade: hardware fit, context, APIs, concurrency, privacy, and evidence.
MCP 2026-07-28 Migration Lab: Stateless Core, Auth, and What Breaks
A practical MCP migration lab for the 2026-07-28 protocol revision: stateless behavior, routing headers, auth checks, cache hints, approvals, tasks, and rollback.
Model Routing: Different Models for Different Jobs
A noob-friendly guide to routing local AI prompts to the right model: Gemma for chat, Qwen Coder for code, Nemotron or reasoning models for planning, small models for summaries, and large models for hard tasks.
Nemotron and Agentic Local AI
A practical guide to NVIDIA Nemotron models and agentic local AI: what a homelab can run, what needs bigger GPUs, and how to keep tools bounded and logged.
Local Coding Assistants: Codex CLI, Continue, Aider, and OpenHands
A noob-friendly guide to choosing and safely using AI coding assistants in a local homelab, including Codex CLI, Continue, Aider, OpenHands, local models, hosted models, Git safety, and task scoping.
Adding RAG: Chat with Documents Locally
A beginner-friendly guide to local document chat, embeddings, vector databases, Open WebUI knowledge bases, privacy, and common RAG mistakes.
Open WebUI Deep Dive: Users, Models, Documents, Prompts, and Safe Sharing
A noob-friendly tour of Open WebUI after installation: user accounts, model switching, documents, prompt shortcuts, admin settings, OpenAI-compatible endpoints, backups, updates, reverse proxy basics, and LAN/VPN security.
Local AI Hardware Sizing: CPU, NPU, GPU, RAM, and VRAM
A local AI hardware sizing guide explaining model memory, quantization, KV cache, CPU-only use, NPUs, GPUs, RAM, VRAM, and upgrade tiers.
Building the Right PC for Local AI
A plain-English guide to choosing the right CPU, RAM, GPU, VRAM, storage, power supply, and cooling for a local AI homelab PC.










