What changed
TypeSafe AI opened early access to Jev on September 15, 2026, and developers are still coming off a waitlist 1. You send model set to jev-latest and a questions map to POST https://api.typesafe.ai/v1/systemone 2. Choice and Score answers carry a confidence between 0 and 1, and a yes/no answer comes back as noul, from 0 (no) to 1 (yes) 2. Input costs $0.042 per MTok and output tokens are free 1.
The row Jev does not have
TypeSafe’s eval site averages accuracy, cost and time over four workflows, and runs each comparison model twice, as an explicit workflow and as one prompt 3:
| Model | Prompt | Workflow | Workflow cost | Workflow time |
|---|---|---|---|---|
| haiku 4.5 | 18.1% | 53.6% | $0.0195 | 12.5 s |
| sonnet 5 | 60.4% | 67.8% | $0.1174 | 78.1 s |
| opus 5 | 64.8% | 73.1% | $0.1761 | 37.8 s |
| sol | 63.4% | 74.1% | $0.0836 | 23.3 s |
| Jev | no row | 67.8% | $0.0004 | 0.4 s |
Anthony Maio reads this as every comparison model getting more accurate, faster and cheaper inside the workflow 4. I checked all eight pairs, and it holds on all three axes 3. Jev appears only as a workflow row 3. I think that is its real adoption cost: a model that only takes typed questions forces the decomposition that already took haiku 4.5 from 18.1% to 53.6% 3. Jev ties the sonnet 5 workflow row at 67.8% and trails sol and opus 5 3. What it wins is $0.0004 and 0.4 s against $0.1174 and 78.1 s for sonnet 5 3.
Impact on your team
If you route, triage or score with one long prompt, split it into typed questions now, on the model you already pay for. In my reading that is where TypeSafe’s table puts the gain, and it needs no waitlist 3. Keep the labels the split forces you to write, then test Jev against them when access arrives. Do not gate escalations on confidence yet. The launch post says all answers come with calibrated probabilities and confidence scores 1, the docs derive confidence from probabilities 2, and yet TypeSafe has not said which statistic, and its calibration methodology is undisclosed 4. Typing fixes the shape of an answer and leaves the judgment unconstrained, as Maio puts it 4, so plot confidence against your own labels first.