Projects with this topic
-
Self-hosted LLM serving where the model store IS the registry: drop a GGUF in the directory and it is servable, no registration step.
Updated -
Personal ComfyUI fork (upstream: github.com/comfyanonymous/ComfyUI, GPL-3.0) with a fully local Wan 2.2 + Z-Image Turbo image/video setup tuned for 8 GB VRAM on WSL2: setup scripts, a model downloader, and ready-to-load workflows.
Updated -
Automated LLM Benchmarking on GPU - tokens/sec, latency percentiles, VRAM profiling, multi-format support (HuggingFace, GGUF, GPTQ)
Updated -
LLM quantization & benchmarking on GPU - GGUF, GPTQ, AWQ, bitsandbytes | Quantification et benchmark de modeles LLM sur GPU
Updated