hn.today

Inspect: An open-source framework for large language model evaluations

inspect.aisi.org.uk5 points0 comments
Screenshot of Inspect: An open-source framework for large language model evaluations

Inspect is an open-source framework for evaluating large language models, built to measure capabilities across coding, agentic behavior, reasoning, factual knowledge, and multimodal understanding. It provides composable components - datasets, solvers, scorers, agents, and tools - so evaluations are modular and reusable, and ships with over 200 ready-to-run benchmarks. Key tooling includes a web-based Inspect View and a VS Code extension for authoring and monitoring runs, support for more than 20 model providers plus local inference backends (Hugging Face, vLLM, SGLang), caching, concurrency, batch mode, and a Python API and CLI for launching evals. Installation is via pip and standard environment API keys; documentation tailored for LLM assistance is available to help integrate large models into development workflows.

Evaluations are defined as Tasks that combine a dataset, a solver (simple generate() calls or full agents), and a scorer (text comparison, model-graded QA, custom scorers). Example snippets show a SimpleQA benchmark using hf_dataset, generate(), and model_graded_qa(), and a sandboxed Capture-the-Flag agent using react() with bash() and todo_write() tools and a Docker sandbox. The framework supports flexible tool calling and multi-agent setups, sandboxing for running untrusted code across Docker, Kubernetes, Modal, and others via an extension API, and analysis features - logs, dataframes, scanners, and visualizations - to inspect transcripts and results.

Read on inspect.aisi.org.uk0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.