hn.today

Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

github.com95 points17 comments
Screenshot of Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

Janus is a single Go binary that runs local .gguf LLMs (via llama.cpp) with Vulkan GPU acceleration on AMD/Intel/NVIDIA or a CPU fallback, exposing an OpenAI-compatible API and a built-in web UI. It routes prompts, runs inference, and can execute toolchains so models drive behavior while Go handles execution and routing. Key features include /v1/chat/completions and /v1/models endpoints (streaming supported), hot-swappable models in the UI, an extensible toolset (file I/O, shell commands, math, docx/PDF export, OCR via Tesseract), optional Ollama proxying, and admin/auth/safe-mode controls. It targets local workflows with no Python/Docker required and supports integration with standard OpenAI clients.

Practical specifics: Windows is the primary target but Linux and macOS are supported; Go 1.22+ is required and a Vulkan-capable GPU is recommended for GPU inference. Typical model files are 2-8 GB and Janus uses ~50 MB for the binary plus llama.dll; initial model load can take 10-60 seconds. Quick start covers cloning, building (prebuilt llama.cpp Vulkan DLLs provided for Windows), placing a .gguf in ./models, editing .env (INFERENCE_BACKEND, JANUS_MODEL_PATH, VRAM and token limits, JANUS_AUTH), and running on http://127.0.0.1:8990. Common pitfalls and fixes (DLLs, port conflicts, wrong paths, missing Tesseract) are documented in the repo.

Read on github.com17 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.