Griffin is a new class of Human Interaction Model (HIM) that performs real-time, face-to-face conversation by perceiving audio and video continuously and generating synchronized speech, facial expressions, gestures, and whole-scene video pixels on the fly. It replaces cascaded pipelines (speech-to-text → LM → renderer) with a unified video-to-video, full‑duplex system: a continuous conversational modeling engine decides when and how to respond every sub-second, while streaming speech and video generators convert those decisions into audiovisual behavior. In live studies Griffin achieved a 48% pass rate on a real-time video Turing test (nearly half of participants thought they spoke with a real person), ranked first on an independent NVIDIA benchmark, and outperformed the next best system by 37%. A limited Griffin‑Lite preview is available to early testers, with broader model releases planned.
Capabilities emphasize natural, timed interaction: back‑channeling, interrupting or waiting appropriately, emotion and behavior modeling, temporal awareness of silences, and perception-driven responses (e.g., coaching during tasks, reacting to objects or gestures). It generates full scenes - face, body, shadows and background - from a reference image in real time. The technical core is concurrent perception, decision-making, and generation so the model can alter its output mid‑utterance. Safety, limitations, and responsible deployment practices are acknowledged, and further evaluation and expansion are forthcoming.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.