What is Jev
Jev is TypeSafe AI's first System One model. It doesn't write text for people to read. It makes decisions inside software: you give it state and a typed question, it gives back one of the answers you allowed, with calibrated probabilities your code can branch on.
Claims on this page are TypeSafe's own, linked below. Jev Lab is unofficial: tests by Prateek Hitli, not affiliated with TypeSafe AI.
System One, Not a Chatbot
LLMHuman preference: write-ups and chat replies.
JevCalibrated decisions: answers with honest probabilities (RLCD, Reinforcement Learning for Calibrated Decisions).
LLMA string you parse and validate, and hope it isn't a hallucinated label.
JevA typed value from the options you defined. TypeSafe: “the model never makes type errors.”
LLMNone you can trust; ask it and it writes a number.
JevCalibrated probabilities on every answer, so code can act, or escalate when unsure.
LLM3 to 329 seconds end to end on these tasks (TypeSafe's figures).
Jev70 to 500 ms end to end, which TypeSafe puts at 40–200× faster.
LLM$0.20 to $10 per million input tokens; output ~5× more.
Jev$0.042 per million input tokens. Output tokens are free.
Three Questions It Can Answer
Every call asks one or more typed questions about the same state. They're answered in parallel, so adding questions barely adds time, and you can mix all three in one request.
Pick one of your options (up to 255). Comes back with a confidence and the probability of every option.
"team": {
"type": "choice",
"instructions": "Which team handles this?",
"criteria": {
"billing": "Charges, refunds",
"technical": "Bugs, outages",
"other": "Anything else"
}
}
→ {"choice": "billing",
"confidence": 0.93,
"probabilities": {...}}Rate against an ordered rubric of 2 to 10 levels, lowest first. Returns the expected level, a confidence and each level's probability.
"severity": {
"type": "score",
"instructions": "How severe is it?",
"criteria": [
"Cosmetic",
"Some users hit",
"Outage"
]
}
→ {"score": 1.8,
"confidence": 0.81,
"probabilities": {...}}How likely a statement is to be true, 0 to 1. No confidence field: the probability is the answer, and your code picks the threshold.
"urgent": {
"type": "noul",
"instructions": "It is urgent.",
"criteria": {
"true": "Outage or deadline",
"false": "No time pressure"
}
}
→ {"noul": 0.97}A Smart If-Statement
TypeSafe's pitch is automation, not conversation: workflows with smart if-statements, map-reduce over big data, real-time apps, and score / judge / verify / guardrail steps.
Ask atomic questions, one judgment each. Combine them in code, not in a prompt. Counting, facts and hard rules stay in code too.
Then use the confidence: act when it's high, escalate when it isn't. That's the pattern every one of the 24 experiments here tests.
import httpx
r = httpx.post(
"https://openrouter.ai/api/alpha/decisions",
headers={"Authorization": f"Bearer {OPENROUTER_API_KEY}"},
json={
"model": "typesafe/jev-1.13",
"state": {"ticket": "I was charged twice for order A-104."},
"questions": {"team": {"type": "choice", ...}}, # as many as you like
},
)
answer = r.json()["answers"]["team"]
if answer["confidence"] >= 0.8:
route(answer["choice"]) # sure enough: act
else:
send_to_human() # not sure: escalateWhere It's Weak
TypeSafe's jaggedness report ↗Unusually, TypeSafe publishes its model's weak spots. Design around these: parse numbers and dates in code, keep the state short, write criteria that don't overlap.
Answers the question you wrote, not the one you meant.
Struggles with counting and numeric precision.
Can't reliably judge how close two hex values or RGB triples are.
Score levels are a rubric, not a ruler: don't average or interpolate them.
Reads dates as text, not as ordered quantities. Parse them in code first.
Double negatives and multi-hop reasoning cost accuracy.
Unrelated detail in the state acts as a distractor.
Text written to steer the model can influence the answer.
Overlapping options or rubric levels give confident answers that mean little.
1 − noul is not the probability of the opposite statement.
How it holds up on real work: the 24 experiments, measured →
Where to Find Jev
Official channels only. Look-alike sites and community orgs using the name aren't TypeSafe's and aren't listed.