made byTypeSafe AI
modeltypesafe/jev-1.13
launched15 Sep 2026 · early access
price$0.042 / M input · output free

What is Jev

Jev is TypeSafe AI's first System One model. It doesn't write text for people to read. It makes decisions inside software: you give it state and a typed question, it gives back one of the answers you allowed, with calibrated probabilities your code can branch on.

Claims on this page are TypeSafe's own, linked below. Jev Lab is unofficial: tests by Prateek Hitli, not affiliated with TypeSafe AI.

System One, Not a Chatbot

Trained for

LLMHuman preference: write-ups and chat replies.

JevCalibrated decisions: answers with honest probabilities (RLCD, Reinforcement Learning for Calibrated Decisions).

Returns

LLMA string you parse and validate, and hope it isn't a hallucinated label.

JevA typed value from the options you defined. TypeSafe: “the model never makes type errors.”

Confidence

LLMNone you can trust; ask it and it writes a number.

JevCalibrated probabilities on every answer, so code can act, or escalate when unsure.

Speed

LLM3 to 329 seconds end to end on these tasks (TypeSafe's figures).

Jev70 to 500 ms end to end, which TypeSafe puts at 40–200× faster.

Price

LLM$0.20 to $10 per million input tokens; output ~5× more.

Jev$0.042 per million input tokens. Output tokens are free.

Three Questions It Can Answer

Every call asks one or more typed questions about the same state. They're answered in parallel, so adding questions barely adds time, and you can mix all three in one request.

Choice

Pick one of your options (up to 255). Comes back with a confidence and the probability of every option.

"team": {
  "type": "choice",
  "instructions": "Which team handles this?",
  "criteria": {
    "billing": "Charges, refunds",
    "technical": "Bugs, outages",
    "other": "Anything else"
  }
}
→ {"choice": "billing",
   "confidence": 0.93,
   "probabilities": {...}}
Score

Rate against an ordered rubric of 2 to 10 levels, lowest first. Returns the expected level, a confidence and each level's probability.

"severity": {
  "type": "score",
  "instructions": "How severe is it?",
  "criteria": [
    "Cosmetic",
    "Some users hit",
    "Outage"
  ]
}
→ {"score": 1.8,
   "confidence": 0.81,
   "probabilities": {...}}
Noul

How likely a statement is to be true, 0 to 1. No confidence field: the probability is the answer, and your code picks the threshold.

"urgent": {
  "type": "noul",
  "instructions": "It is urgent.",
  "criteria": {
    "true": "Outage or deadline",
    "false": "No time pressure"
  }
}
→ {"noul": 0.97}

A Smart If-Statement

TypeSafe's pitch is automation, not conversation: workflows with smart if-statements, map-reduce over big data, real-time apps, and score / judge / verify / guardrail steps.

Ask atomic questions, one judgment each. Combine them in code, not in a prompt. Counting, facts and hard rules stay in code too.

Then use the confidence: act when it's high, escalate when it isn't. That's the pattern every one of the 24 experiments here tests.

the exact call this lab makes, via OpenRouter
import httpx

r = httpx.post(
    "https://openrouter.ai/api/alpha/decisions",
    headers={"Authorization": f"Bearer {OPENROUTER_API_KEY}"},
    json={
        "model": "typesafe/jev-1.13",
        "state": {"ticket": "I was charged twice for order A-104."},
        "questions": {"team": {"type": "choice", ...}},   # as many as you like
    },
)
answer = r.json()["answers"]["team"]
if answer["confidence"] >= 0.8:
    route(answer["choice"])        # sure enough: act
else:
    send_to_human()                # not sure: escalate

Unusually, TypeSafe publishes its model's weak spots. Design around these: parse numbers and dates in code, keep the state short, write criteria that don't overlap.

Literal reading

Answers the question you wrote, not the one you meant.

Math and numbers

Struggles with counting and numeric precision.

Numeric representations

Can't reliably judge how close two hex values or RGB triples are.

Math using score

Score levels are a rubric, not a ruler: don't average or interpolate them.

Dates and times

Reads dates as text, not as ordered quantities. Parse them in code first.

Indirection

Double negatives and multi-hop reasoning cost accuracy.

Irrelevant state

Unrelated detail in the state acts as a distractor.

Adversarial content

Text written to steer the model can influence the answer.

Contradictory criteria

Overlapping options or rubric levels give confident answers that mean little.

P(noul) + P(not noul) ≠ 1

1 − noul is not the probability of the opposite statement.

How it holds up on real work: the 24 experiments, measured →

Where to Find Jev

Official channels only. Look-alike sites and community orgs using the name aren't TypeSafe's and aren't listed.