Jevstiller is a proxy that trains a small local model to reproduce a remote classifier called Jev so answers can be returned in ~15 ms instead of ~300 ms. Its core contract is explicit and measurable: pick an agreement target A (for example 98%), and the system guarantees that at least A of requests receive the label Jev would have given. Formally, with c the local model’s coverage and e its disagreement rate on those answered, overall agreement is A = 1 − c·e, and the router must maximize coverage while keeping c·e below the budget β = 1 − A. Simple calibration by choosing the loosest confidence threshold that passes on held-out data fails in finite samples and broke the 2% budget many times in benchmarks; point estimates understate uncertainty and pick thresholds biased by sampling noise.
To make the guarantee hold with high confidence Jevstiller combines four practices: define the loss exactly as the indicator of local disagreement and use an exact Clopper-Pearson upper bound; test candidate thresholds in fixed-sequence order on a pre-defined grid; reserve calibration rows from any tuning; and fit conservatively (threshold at 85% of budget) with a shadow check on fresh traffic. It also supports treating low Jev confidence as an “unsure” flag so local answers don’t hide Jev’s uncertainty. Production uses a constant audit slice (e.g., 2%) to detect drift and force retrain/rollback; experiments show fast recovery after silent label shifts. Costs are reduced coverage (typically 4-8 points versus optimistic point estimates) and assumptions that calibration traffic is representative; agreement is not the same as ground-truth accuracy and per-class guarantees are future work.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.