GPT-6 Astra is presented as the strongest vision model tested by Roboflow, excelling across object detection, counting, visual reasoning, segmentation, re-identification, and prompting use cases. On Roboflow Vision Evals Astra achieved 82.1% mAP@50 for detection at low reasoning effort - 5.4 points above Qwen3.8 Max and 13.7 above GPT-5.6 Sol - and hit near-perfect scores on specific tasks (99.8% mAP on distinguishing LEGO brick types). Counting reached 80.2% (low effort) and 81.1% (high), while visual reasoning scored 87.2% and 91.2% respectively. Astra is now Roboflow Auto Annotate’s default; it often produces boxes more accurate than manual labels except for highly specialized, rarely seen classes.
Astra’s multimodal prompting is a key strength: text-class prompts, box prompts (positive/negative examples), and cross-image examples dramatically improve precision (dice detection rose from 21.3% to 90.1% with extra examples; football to 100%). Segmentation outputs polygons as JSON and outperforms SAM 3 on language-directed labeling (finding and naming small text regions), though SAM 3 provides finer boundary masks and is faster/cheaper; a recommended pipeline is Astra for selection then SAM 3 for mask refinement. For video, Astra can re-identify players, read jersey numbers, and supply cross-frame identities, but per-frame costs (~$0.02-$0.08) make full-frame processing expensive, motivating hybrid pipelines that use local detectors/trackers with intermittent Astra calls.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.