Running the Web UI with a Local GPU LLM¶
This guide runs generated Wiki pages and Ask against a local repository using a
local GPU LLM, without cloud credentials. The repository must contain at least
one supported source language, or you must pass an
explicit supported --language.
The setup has three services running in separate terminals:
| Service | Script | Where to run |
|---|---|---|
| LLM server (llama-cpp-python) | scripts/start_llm.sh |
GPU node |
| CodeNib backend (FastAPI) | scripts/start_web.sh |
Main machine |
| Vite frontend | cd web && npm run dev |
Main machine |
Prerequisites¶
Main machine¶
- A CodeNib source checkout with development dependencies:
make dev - Node.js
^20.19.0or>=22.12.0, plus npm:make web-deps(once) - Conda env
codenibactive
GPU node¶
- Access to a node with CUDA 12.4+ driver and enough VRAM (7B model needs ~5 GB)
llama-cpp-python[server]with GPU support installed in thecodenibenv:
# Install the pre-built CUDA 12.4 wheel (works with any CUDA 12.4+ driver)
conda activate codenib
pip install "llama-cpp-python[server]==0.3.29" \
--extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu124
Verify GPU is detected:
python -c "import llama_cpp; print('GPU:', llama_cpp.llama_supports_gpu_offload())"
# Should print: GPU: True
- A GGUF model file. If you have Ollama installed, qwen2.5-coder:7b is at: Otherwise download any GGUF from HuggingFace and note the path.
Step 1 — Index a repo¶
For a repository you want to explore, build its BM25 index and register it in one command:
# Clone the repo (skip if already cloned)
git clone https://github.com/<owner>/<repo> ~/projects/<repo>
# Build the index and register the repo
conda activate codenib
cd ~/projects/CodeNib/CodeNib
python scripts/index_repo.py /absolute/path/to/your/repo
The script uses CodeNib's shared language registry to detect every supported
language represented by the repository's source extensions. Override detection
with --language go; comma- or slash-separated values such as
javascript/typescript also work. If detection finds nothing supported, the
script exits instead of silently indexing the repository as Python.
Indexes are written below
$CODENIB_HOME/repositories/<repo>-<id>/indexes (default
~/.codenib/repositories/...). The web registry defaults to
.codenib_qa/qa_registry.json; change that path with --registry. Restart an
already-running backend after registering a repository.
Step 2 — Configure the LLM¶
No tracked config edit is required. scripts/start_web.sh points the backend at
the local OpenAI-compatible endpoint by exporting:
For local-only config changes, copy the template and edit the ignored file:
When present, scripts/start_web.sh automatically uses qa_config.local.yaml.
Override CODENIB_DEMO_CONFIG or CODENIB_DEMO_MODEL before running
start_web.sh if your local server exposes a different model name or config
path.
Step 3 — Start all three services¶
Terminal 1 — LLM server (on GPU node)¶
The script will ask for your GGUF model path and start an OpenAI-compatible server on port 8080.
Terminal 2 — CodeNib backend (on main machine)¶
The script asks for the GPU node hostname (default: vscode-dsmlp-l40s) and
starts the FastAPI backend on port 8000 pointed at your LLM server.
Terminal 3 — Frontend (on main machine)¶
Opens at http://localhost:3000.
Step 4 — Load or generate a page¶
- Open http://localhost:3000
- Click on your repo
- Select a page
- On its first request, wait while the backend generates and caches it
Wiki pages are cached under <data_dir>/wiki_cache (by default
.codenib_qa/wiki_cache) so subsequent loads avoid another model call.
Refresh this wiki only re-fetches the current tree and page; it does not
invalidate that cache or force generation. To regenerate, stop the backend,
delete the wiki_cache directory, and start the backend again.
Troubleshooting¶
| Symptom | Fix |
|---|---|
GPU: False from llama_cpp |
Wrong wheel installed. Reinstall with --extra-index-url as shown above. |
ContextWindowExceededError |
LLM server started without --n_ctx 8192. The start script sets this automatically. |
Connection refused on port 8080 |
LLM server not running, or firewall blocking the GPU node port. Check Terminal 1. |
| Wiki says "Couldn't load this page" | Check backend terminal for WARNING outline generation failed. Usually an LLM connectivity issue. |
repos: 0 at /api/health |
qa_registry.json missing or wrong path. Check .codenib_qa/qa_registry.json. |
| Backend stuck on "Loading repositories…" | Index not built. Run Step 1 again. |
| Wiki still shows degraded text after fixing the model | Stop the backend, remove <data_dir>/wiki_cache, and restart it so both in-memory and on-disk results are cleared. |
Running over SSH¶
With the default or another loopback API configuration, forward port 3000 only:
The browser talks same-origin to the Vite dev server, which proxies
/api/* server-side to the FastAPI backend at CODENIB_API_BASE (default
http://127.0.0.1:8000; see web/vite.config.ts) — port 8000 does not need to
be forwarded.
If the LLM server is on a different node than the backend, only the backend needs to reach port 8080 on the GPU node — the browser never talks to port 8080 directly.