A Jev-driven SRE diagnosis pipeline programmatically collects Kubernetes cluster state, groups observations by component, and summarizes failure signals for Jev to evaluate. Instead of an LLM agent issuing commands, the collector supplies multiple-choice options and numbered evidence items; Jev selects a likely origin, classifies the object type (for example an admission webhook), and picks the evidence that best supports its verdict. The pipeline then gathers focused detail on that candidate, repeats selection until the hypothesis is supported, and assembles a final diagnosis report. The implementation explores a one-at-a-time investigation strategy, keeps Jev decisions constrained to supplied options, and is available on GitHub.
Across 21 SREGym-Lite faults run five times each (105 attempts), the pipeline produced 80 passing diagnoses (76.2%) with a median diagnosis time of 14.6 seconds. Sixteen faults passed in all five runs, five failed every run, and runs were unusually consistent per fault. Example success: a mutating webhook that rewrote Pod memory limits (template 256Mi → Pod 16Mi) was identified as the origin, Jev selected the webhook evidence, and the submitted diagnosis met rubric criteria. Evaluation used a nine-question rubric scored by gpt-6-astra with a 0.70 pass threshold. Runtime stats: 252 Jev calls total (2.4 per attempt), 3.48M input tokens, median summed Jev latency 0.53s, estimated Jev inference cost ~$0.15.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.