LiquidAI d1-3B is a 3-billion-parameter multimodal decision model packaged for image-text-to-text tasks and optimized for classification, calibration, and conversational decision-making. Built on LiquidAI’s LFM2.5-VL-3B base, it supports 16 languages and is tagged for edge deployment, system-one workflows, and custom-code integrations. Model artifacts are provided in safetensors (BF16 parameter count ~3.12B; total files around 6.25 GB), licensed under lfm1.0, region-marked US, and marked compatible with inference endpoints and quantizations (hasQuantizations true, isQuantized false). Metadata highlights links to the LFM2 technical report (arXiv:2511.23404), 23 likes and 15 downloads, and pipeline classification labels emphasizing decision, calibration, and multimodal conversational use.
Usage instructions target practical deployment: examples show how to run an image-text-to-text pipeline with Transformers (pipeline("image-text-to-text"), AutoProcessor and AutoModelForMultimodalLM) and how to serve via vLLM with an OpenAI-compatible API. Example inputs illustrate chat-style messages combining image URL objects and text prompts and recommend device_map/auto and processor.apply_chat_template before model.generate. The package is positioned for researchers and engineers who need a multilingual, edge-capable multimodal decision model with ready integration into common tooling (Transformers, vLLM) and endpoint workflows.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.