Run Google Gemma Locally with Ollama, Open WebUI, and Codex CLI
The short answer: run a downloaded Gemma model in Ollama on loopback, then connect a single Open WebUI instance and Codex CLI to that local provider. This guide uses native Linux host networking with WebUI itself bound to 127.0.0.1:3000, reached through SSH. Disable optional cloud features and review tools, plugins and remote providers separately. Local model weights alone do not prove that every part of the workflow stays offline.
This guide is for a single Linux homelab host or workstation. It covers model sizing, Open WebUI installation, Ollama, optional NVIDIA acceleration, and Codex CLI. It assumes you can use SSH, Docker, systemd, and a firewall; it does not assume that the machine is safe to expose to the public internet.
The target stack:
- Google Gemma as the local open-weight model family
- Ollama as the model runtime and local API server
- Open WebUI as the ChatGPT-like browser interface
- Codex CLI in OSS mode for local coding-agent workflows
- Optional NVIDIA GPU acceleration
- Optional coexistence with Tdarr or other GPU workloads
Important distinction: Gemma and Gemini are not the same thing. Gemini is Google's hosted API model family. Gemma is Google's open-weight model family that you can download and run locally.
Who This Is For
This guide is for homelab users, developers, and media-server owners who want to run Gemma locally instead of relying entirely on hosted AI services. It assumes basic comfort with Linux, Docker, SSH, and terminal commands.
What You'll Build
The intended acceptance targets are:
- Ollama serving a Gemma model locally on port 11434
- Open WebUI available in a browser on port 3000
- A working local chat interface for Gemma
- A Codex CLI workflow that can target the local Ollama/Gemma provider
- A safer mental model for sharing GPU resources with Tdarr or ffmpeg
Start With the Local Model
Complete the installation and storage prerequisites below before starting services. First prove one local prompt; add the browser and coding-agent workflows separately. There is no tested one-command installer for this stack here.
ollama --version &&
ollama pull gemma4:e2b &&
ollama run gemma4:e2b "Explain what Open WebUI does in one paragraph."
Security warning: keep Ollama and WebUI off public interfaces. The setup below uses SSH access to loopback, not a publicly reachable login page.
Interactive Flow Diagram: Local AI Architecture
Interactive Stack Flow: Browser to Local Model
Select each stage to see what it does and where it runs.
Click a stage in the flow to view operational details.
Hardware Planning
Memory is a capacity constraint, not a speed or quality guarantee. Choose the exact model artifact and quantization, then budget for context, vision inputs, runtime buffers and competing processes. E2B/E4B are effective-parameter labels; their full weight footprint is larger than those labels alone suggest.
Practical Gemma sizing:
| Model tag | Ollama listing, October 2, 2026 | Planning limit |
|---|---|---|
gemma4:e2b | 4.6-7.5 GB | Variant-dependent download size, not free-memory requirement. |
gemma4:e4b | 6.6-9.5 GB | Include full weights and context, not only effective parameters. |
gemma4:12b | 7.7-8.0 GB | Runtime and quantization choice affect fit. |
gemma4:26b | 16-19 GB | Active expert count does not remove resident-weight cost. |
gemma4:31b | 19-20 GB | Leave additional memory for context and other workloads. |
Earlier Tag-Size Snapshot
This retained graphic shows the earlier article's approximate tag-size snapshot, not the current listing. Use the dated table above and the exact downloaded artifact for planning. Neither these bars nor download sizes establish runtime memory fit; context, images, buffers and other GPU processes add overhead.
For Gemma 4 31B, Google lists approximate inference memory at about 17.5GB for 4-bit Q4_0, 34.9GB for 8-bit SFP8, and 69.9GB for BF16. Those figures are approximate GPU/TPU memory to load the model with overhead, not a guarantee that a full 256K-context Ollama session will fit; KV cache, runtime overhead, multimodal inputs, and other GPU processes add more.
For a 31B workstation, evaluate these components against the intended artifact and workload, not a promised fit:
- GPU: compare supported backend, usable memory and measured workload latency. A model-file size below nominal VRAM is not proof that generation fits.
- RAM: account for operating-system use, CPU-offloaded weights and concurrent services. Offload can change both fit and speed.
- CPU: evaluate prompt processing, CPU offload and other host workloads before paying for additional cores.
- Storage: budget for selected model files, previous versions, WebUI data and protected recovery copies. No universal 2TB minimum is established.
- PSU: follow the exact GPU/system requirements, connectors and electrical limits rather than a model-size-based wattage rule.
- OS: Linux preferred for server-style use
If you also run Tdarr, Plex, Jellyfin, Frigate, Stable Diffusion, or any other GPU workload, leave extra headroom. A model that barely fits while idle may fail when ffmpeg is already using VRAM.
Useful monitoring command:
watch -n 1 nvidia-smi
Step 1: Prepare the Linux Host
Use a supported native Linux host and the distribution-specific Docker Engine installation procedure. Resolve conflicting packages and review firewall behavior before installing. Install Git, curl, CA certificates and OpenSSL through the distribution package manager. Do not execute a downloaded installer unseen.
Privilege warning: Docker daemon access is effectively root-equivalent. The commands below assume an authorized, trusted administrator already has access. Do not add untrusted users to the Docker group to bypass an error.
docker version &&
docker compose version
For an NVIDIA deployment, install a supported driver using the vendor/distribution instructions and check it independently:
nvidia-smi
A visible GPU does not prove that Ollama used it; retain the inference and contention tests below.
Step 2: Install Ollama
Install Ollama using its official Linux instructions. Review the installer or documented manual installation steps before executing them, and match the runtime to the intended model. Do not pipe a fresh network response directly into a shell.
ollama --version
systemctl status --no-pager ollama
Service configuration: for a systemd installation, use sudo systemctl edit ollama.service to merge Environment="OLLAMA_HOST=127.0.0.1:11434" and Environment="OLLAMA_NO_CLOUD=1" under [Service]. Preserve existing overrides. Reload systemd, restart Ollama during a suitable window, and confirm the listener and cloud-disabled log before using private data. Setting a variable only in the client shell does not configure an already-running service.
Verify the service:
ollama --version
curl http://127.0.0.1:11434/api/tags
Pull a Gemma model:
ollama pull gemma4:e2b
For a stronger system, 12B is the next sensible step before 26B or 31B:
ollama pull gemma4:12b
For 24GB+ VRAM or high-RAM systems:
ollama pull gemma4:26b
ollama pull gemma4:31b
Test a prompt:
ollama run gemma4:e2b "Explain what Open WebUI does in one paragraph."
If you are on smaller hardware, start with E2B. Do not begin with 31B unless you know your system has the memory for it.
Step 3: Configure One Open WebUI Instance
Use a new, private project directory on native Linux. Do not run this initialization against an existing installation or data volume. The block refuses an existing .env; a failed key-generation command stops before writing it. Existing users must preserve their current key and complete data volume, not generate replacements.
(
set -euo pipefail
set -C
umask 077
mkdir -p "$HOME/open-webui"
cd "$HOME/open-webui"
test ! -e .env
test ! -L .env
secret="$(openssl rand -hex 32)"
test "${#secret}" -eq 64
printf 'WEBUI_SECRET_KEY=%s\nOPEN_WEBUI_IMAGE=\n' "$secret" > .env
)
Edit the new ~/open-webui/.env with a trusted editor and set OPEN_WEBUI_IMAGE to the reviewed release tag or digest, including the ghcr.io/open-webui/open-webui repository. Leave the generated key unchanged. Check the selected release supports the HOST and PORT settings; the current project startup source passes them to its web server. This guide does not choose a perpetual safe image version.
Create ~/open-webui/docker-compose.yml with the following content. Refuse an existing file until you have backed it up and compared its volume, secret and network settings. The explicit volume name preserves the original guide's open-webui volume identity; do not silently attach an unrelated existing volume.
services:
open-webui:
image: ${OPEN_WEBUI_IMAGE:?Set a reviewed image tag or digest}
container_name: open-webui
restart: unless-stopped
network_mode: host
env_file: .env
environment:
HOST: 127.0.0.1
PORT: "3000"
OLLAMA_BASE_URL: http://127.0.0.1:11434
volumes:
- open-webui:/app/backend/data
volumes:
open-webui:
name: open-webui
Network boundary: host networking lets the trusted WebUI process reach Ollama on host loopback. It also gives that container access to other host-network services, so do not treat it as network isolation. No Docker published-port mapping applies here. WebUI must actually listen on 127.0.0.1:3000; if the selected image ignores that setting, stop it rather than expose it. A bridge container cannot reach host loopback merely by naming host.docker.internal.
Step 4: Start, Verify and Recover
cd "$HOME/open-webui" &&
docker compose config --quiet &&
docker compose pull &&
docker compose up -d &&
docker compose ps
From the trusted workstation, open an SSH tunnel. Replace the account and host placeholder; keep this session open while using http://127.0.0.1:3000 locally:
ssh -N -L 127.0.0.1:3000:127.0.0.1:3000 YOUR_ADMIN_USER@YOUR_SERVER
On the server, inspect the actual listeners and logs before creating the first admin account:
ss -ltnp
curl --fail --silent --show-error http://127.0.0.1:11434/api/tags
docker logs --tail=100 open-webui
In WebUI, create the first administrator through the tunnel, review registration/account policy, and check the Ollama connection in Admin Settings. Select the downloaded model and send a harmless prompt. Confirm that saved connection settings have not retained another provider or URL. Before sharing, review the hardening guide; Tools and Functions execute server-side code and are not ordinary prompt text. Disable unused extensions and test unauthorized access from another device.
For an update, stop WebUI, take a consistent protected backup of the complete data volume and project configuration, and verify the recovery image is available. Select the reviewed target image in the existing .env, then rerun the guarded Compose sequence above. Do not regenerate the secret, change the volume name or forcibly delete the old container before checking recovery. A database migration can require restoring the matching pre-update data as well as the old image.
If a listener, migration or privacy check fails, stop WebUI with the command below and investigate. Rollback does not retract information already sent to a provider:
cd "$HOME/open-webui" &&
docker compose stop open-webui
Step 5: Tie In Codex CLI with Local Gemma
Codex CLI can run against local open-source providers. With Ollama installed, Codex can use Gemma through Ollama's local API.
Install Codex CLI using the current official CLI instructions and an approved package/release. Inspect any installer before execution. Check the installed flags without starting an agent:
codex --version
codex --help
Authentication depends on the selected provider and managed policy. Official Codex documentation supports local providers and provider configurations without OpenAI authentication; a ChatGPT sign-in is not a universal prerequisite for local inference. Hosted Codex features and remote integrations have their own requirements. This guide has not tested an unauthenticated Gemma session on your installed version.
Verify:
codex --version
codex --help
Make sure Ollama has a model:
ollama pull gemma4:e2b
Set Ollama as the default OSS provider by adding this as a root-level key near the top of ~/.codex/config.toml, before any [section] headers:
oss_provider = "ollama"
Put that in:
~/.codex/config.toml
Run Codex in OSS mode against Ollama:
codex --oss -c 'oss_provider="ollama"' -m gemma4:e2b
Run a one-shot prompt:
codex exec --oss -c 'oss_provider="ollama"' -m gemma4:e2b "Summarize this repository and identify the main entry points."
Inside a project directory:
cd /path/to/your/repo &&
codex --oss -c 'oss_provider="ollama"' -m gemma4:e2b
The current documented local-provider flag can make the provider choice explicit. Begin with a read-only sandbox in a disposable, credential-free repository; verify model/tool compatibility before allowing edits:
codex --oss --local-provider ollama --sandbox read-only -m gemma4:e2b
Use codex --help to confirm which flags your installed version supports.
Scope: these examples target local Codex CLI inference through Ollama, not a claim that hosted Codex uses this local model. Authentication is provider-dependent; ChatGPT/API-key login is not universally required for local inference. Inspect enabled hooks, MCP integrations, environment access and sandbox settings separately before calling the workflow private.
Command safety warning: A local coding agent can propose and run shell commands. Keep work in git, review diffs, and avoid granting broad filesystem access unless you understand the risk.
Interactive Flow Diagram: Codex + Gemma Coding Flow
Interactive Codex Flow: Local Gemma Provider
Tdarr and GPU Sharing Notes
If the same system also runs Tdarr, your GPU is shared, but the contention is not identical: Tdarr/ffmpeg GPU transcodes usually use fixed-function NVENC/NVDEC engines, which are separate from the CUDA/Tensor compute Ollama uses. They can still conflict through VRAM, memory bandwidth, copy engines, power, thermals, and driver scheduling.
Check usage:
nvidia-smi
If you see ffmpeg processes using VRAM, Ollama has less room to load Gemma.
Practical scheduling:
- Begin with the smallest suitable artifact and measure it with the intended Tdarr workload; the E2B label alone does not establish fit or safety.
- Test coexistence at the actual quantization and context. Record memory and encoder/decoder utilization rather than assuming a 16GB threshold guarantees coexistence.
- Pause Tdarr before loading 26B/31B on 24GB-class cards, and avoid 31B plus active transcodes unless you have substantial VRAM headroom and have tested it.
- Use Tdarr schedules to keep heavy transcodes overnight.
- Consider a second GPU if this machine is both a media server and an AI server.
For a dedicated AI box, keep it separate from Tdarr. For a combined homelab system, use smaller models and accept slower responses.
Security Recommendations
Do not expose Open WebUI directly to the internet without protection.
Better options:
- Use Tailscale, WireGuard, or ZeroTier.
- Put Open WebUI behind a reverse proxy with HTTPS.
- Require strong accounts and passwords.
- Keep the Ollama API port private.
- Do not publish port 11434 to the public internet.
- Back up the Open WebUI volume.
Example firewall stance:
Keep WebUI on 127.0.0.1:3000 and use the SSH tunnel.
Keep Ollama on 127.0.0.1:11434.
Do not expose Docker socket.
Do not run random model-generated shell commands without review.
Troubleshooting
Open WebUI cannot see Ollama:
curl -fsS http://127.0.0.1:11434/api/tags
docker inspect --format '{{.HostConfig.NetworkMode}}' open-webui
docker logs open-webui
For this guide's host-network deployment, the inspect result should be host and WebUI's effective Ollama URL should be http://127.0.0.1:11434. If the host request succeeds but WebUI fails, check saved connection settings and application logs. Do not broaden Ollama to all interfaces as a troubleshooting shortcut.
Try setting the Ollama URL in Open WebUI to:
http://127.0.0.1:11434
Model is too slow:
- Use a smaller model.
- Use a 4-bit quantized model.
- Reduce context length.
- Stop Tdarr/ffmpeg workloads.
- Check whether the GPU is being used.
GPU memory error:
- Run
nvidia-smi. - Stop competing processes.
- Use E2B or E4B instead of 12B/26B/31B.
- Lower context.
Codex local model gives weak edits:
- Try a coding-tuned model through Ollama for coding tasks.
- Keep tasks small and specific.
- Ask Codex to inspect before editing.
- Use git and review diffs carefully.
- For complex coding work, use a stronger hosted model or a larger local GPU.
Before Adding More Models
Keep one model and one browser workflow working before adding Codex tool use or competing GPU jobs. Reuse the existing Compose project and secret. The small-model prompt above checks inference only; the acceptance tests below cover isolation, egress, memory and task quality.
Final Thoughts
The best first version of this stack is simple:
Ollama + Gemma E2B/E4B + Open WebUI
Once that works, add Codex CLI in OSS mode:
Codex CLI + Ollama + Gemma
That gives you two local workflows: a browser-based chat interface for everyday use and a terminal-based coding agent for repo work. Smaller Gemma models are reasonable for learning and tightly scoped tasks. A larger model may improve difficult repo work, but only a task-based evaluation can show whether the extra latency, memory, and power are justified.
For a mixed-use homelab server that also runs Tdarr, start small. For a dedicated local AI workstation, build around VRAM first.
How to Verify This Stack
Evidence status: this guide is documentation-backed, not a completed deployment. The networking, secret-handling and Codex examples were rechecked on October 2, 2026; shell syntax was checked without running the services. TechGeeks did not install this exact stack or measure throughput, VRAM, power, egress or coding-task success. The checks below are acceptance tests to perform, not reported results.
- Record versions and tags. Save the output of
ollama --version,codex --version,docker image inspect, andollama show gemma4:e2b. A mutablelatestormaintag is not enough for a reproducible result. - Prove the execution path. While a prompt is running, check
ollama psand GPU telemetry such asnvidia-smi. Record processor placement, model size, context setting, and peak memory. A downloaded model file does not prove the runtime used the GPU. - Prove the service boundary. Use
ss -ltnpto confirm port 11434 is bound only where intended. From an unauthorized VLAN or guest device, model enumeration and generation requests should fail. - Prove local-only behavior. Enable
OLLAMA_NO_CLOUD=1ordisable_ollama_cloud, restart Ollama, and confirm the log says cloud is disabled. Then temporarily block WAN egress and repeat a local Open WebUI chat and a Codex task. Disable web search, remote providers, pipes, tools, and browser plugins for this test. - Test the real workload. Run at least five cold and five warm trials with a fixed prompt. Record time to first token, generation rate, peak RAM/VRAM, errors, and whether a simultaneous Tdarr job causes an out-of-memory event or unacceptable delay.
- Test Codex in a disposable repository. Give it one bounded edit with a test, inspect every command and diff, and verify that sandbox and approval settings prevent writes outside the workspace. Do not place real credentials in the test repo.
Risk, Recovery, and Legal Boundaries
Before an Open WebUI update, stop the application and make a cold backup of its complete /app/backend/data volume plus the Compose file and protected .env. Restore that backup to a separate volume and sign in before calling it recoverable. Pin a known image tag or digest so rollback means recreating the old container against a compatible copy of the data, not guessing which build main used.
If the stack starts exposing data or consuming resources unexpectedly, the break-glass action is to stop Open WebUI and Ollama, remove the network route or firewall allowance, revoke connected provider keys, and preserve logs before changing configuration. Codex changes should live on a Git branch; rollback is a reviewed revert or branch deletion, not a blind cleanup command. A model rollback can restore behavior, but it cannot retract text already sent to a remote provider or tool.
Gemma is open-weight under Google's terms, not automatically Apache-licensed software. Review the exact model license and prohibited-use policy before redistribution or commercial use. Only index or submit code and documents you own or are authorized to process. Local execution does not remove copyright, privacy, employment, export, or client-confidentiality obligations.
What This Evidence Does Not Prove
- Google's weight-memory table and Ollama's download sizes do not prove that a model fits at your chosen context length; key-value cache, vision input, runtime buffers, and competing GPU work add memory.
- A successful chat does not prove coding-agent reliability, correct tool calls, safe command selection, or acceptable results across another repository.
OLLAMA_NO_CLOUD=1does not control Open WebUI plugins, remote model connections, package downloads, Codex integrations, or third-party browser extensions.- A WAN-blocked test shows behavior during that observation window. It does not audit every future update, dependency, administrator change, or user-installed tool.
- Loopback binding protects against direct remote access; it does not protect the host from a compromised local process or an overprivileged container.
Related TechGeeks Reading
- NVIDIA vs Intel GPUs for Local AI
- Docker Compose for Normal People
- Remote Access Without Opening Router Ports
- AI Workflow Notes: Start Here
References
- Google Gemma docs: https://ai.google.dev/gemma/docs
- Gemma 4 model and memory requirements: https://ai.google.dev/gemma/docs/core
- Run Gemma with Ollama: https://ai.google.dev/gemma/docs/integrations/ollama
- Open WebUI docs: https://docs.openwebui.com/
- Open WebUI project: https://github.com/open-webui/open-webui
- Codex CLI docs: https://developers.openai.com/codex/cli
- Codex advanced configuration and local providers: https://developers.openai.com/codex/config-advanced
- Ollama downloads: https://ollama.com/download
- Ollama Gemma 4 library and current tag sizes: https://ollama.com/library/gemma4
- Ollama cloud controls and network binding: https://docs.ollama.com/faq
- Ollama Launch integration for Codex: https://ollama.com/blog/launch
- Open WebUI hardening guidance: https://docs.openwebui.com/getting-started/advanced-topics/hardening/
- GitHub Security Lab, Open WebUI tool-restriction bypass analysis: https://securitylab.github.com/advisories/GHSL-2026-002_Open_WebUI/


