Mistral Large 4 is a public-preview, open-weight, general-purpose multimodal model built on a granular Mixture-of-Experts (MoE) architecture. It exposes 49 billion active parameters backed by 1.05 trillion total parameters and includes a 1.6 billion‑parameter vision encoder, supporting very long contexts (up to 1 million tokens). The documentation highlights a variant labeled mistral-large-4+1, tradeoffs between speed and performance, and example pricing/tiering for inference: per‑million‑token costs for input, cached input, and output are shown across tiers, with sample values in the roughly $0.07-$4.18 per‑M‑token range depending on caching and priority.
The model is integrated into Mistral’s inference product and supports production features relevant to developers and engineers: structured outputs, function calling, document Q&A, prefix prompts, chat completions, batching, agents and conversations, and built‑in tools. Corresponding API paths are documented (chat completions, batching, agents/conversations, etc.), and the offering is presented alongside other Mistral models and platform tooling (Studio, Vibe, code/compute resources). The page gives enough technical and commercial detail to evaluate capability, latency/quality tradeoffs, and cost implications before deeper experimentation or deployment.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.