V8 · local SLM serving
Serve a GGUF language model inside your own infrastructure.
CortexLab V8 can connect to an internal llama.cpp server. The model file stays on your machine or server. Cortex retrieval planning evidence and policy layers remain separate from language generation.
runtime offline
CortexLab connects only to the internal llama.cpp service configured by LOCAL_SLM_URL. No external model API is required.
No output yet.
Local weights
Place a compatible GGUF model in the models directory and select it through LOCAL_SLM_MODEL_FILE.
Private network
The provided Docker network is internal. The SLM service is not exposed publicly by default.
Replaceable runtime
The Next.js adapter talks to a local completion endpoint so the underlying quantized model can be upgraded independently.