OpenAI Decisions API beats Jev on real-world checks
- OpenAI Decisions API75.2%±3.3
- Mercury Decide70.5%±3.6
- deck-31B66.8%±3.3
- Kev-27B64.8%±3.8
- pplx-decider v1.1 27B64.7%±3.6
OpenAI Decisions API beats Jev on real-world checks
Frontier LLMs are up to 25 points more accurate than Jev, but cost over 100x more
Gemma 4 31B leaves the least PII in on the LangWatch policy test
Frontier LLMs beat Jev on 9 of 11 tasks
Open models caught up with Jev
Jev beats every tiny open model
Each bar is a score with its 95% interval. Orange: ahead alone. Green: tied for the lead. Gray: the rest. Hatched with *: trained on the test data, so not ranked. Average: the unweighted mean of a model's task scores; tasks use different metrics, so read it as a summary, not a ranking.