V8 · local SLM serving

Serve a GGUF language model inside your own infrastructure.

CortexLab V8 can connect to an internal llama.cpp server. The model file stays on your machine or server. Cortex retrieval planning evidence and policy layers remain separate from language generation.

local gguf inference

runtime offline

CortexLab connects only to the internal llama.cpp service configured by LOCAL_SLM_URL. No external model API is required.

model output
No output yet.

Local weights

Place a compatible GGUF model in the models directory and select it through LOCAL_SLM_MODEL_FILE.

Private network

The provided Docker network is internal. The SLM service is not exposed publicly by default.

Replaceable runtime

The Next.js adapter talks to a local completion endpoint so the underlying quantized model can be upgraded independently.