b7882b2434f2d558acf3fc3906772920089fa3aa
RTX 3090 was at 94.9% VRAM at 262K context. Reduced to 192K (196608), freeing ~2.4GB. VRAM now at 85% with room for active inference.
Description
SyslogAI Inference Harness — 3-GPU router, dashboard, LiteLLM proxy
916 KiB
Languages
Python
77%
HTML
22%
Shell
0.8%
Dockerfile
0.2%