V9 MODEL OPERATIONS

GPU-aware model manager

inspect local hardware model artifacts quantization choices and serving placement without sending model metadata to a cloud AI API.

gpu inventory

serving hardware

probing model manager…

registered models

placement planner

register GGUF or adapter artifacts to see serving plans.

LORA PIPELINE

adapter training

Generate a reproducible local config with npm run cortex:v9:lora:config then run the optional CUDA trainer using npm run v9:train. The trainer uses local model files and local JSONL data.

QUANTIZATION

GGUF serving export

With llama.cpp tools installed run npm run cortex:v9:quantize -- input.gguf output.gguf Q4_K_M cortex-local. The resulting artifact is added to the local model registry.