03
Severity Scoring
jev primitiveScore
labelslowmediumhigh
dataset18 labeled rows
tracingone call, not traced
Scores how severe a ticket is: low, medium, or high. Jev answers one Score question over three ordered levels and returns a probability-weighted score (0–2) with a confidence. The label is the most likely level, and the score feeds an "escalate if score ≥ t" sweep. The baselines pick a level in free text, or from an enum. Measured: accuracy, latency, cost, and the escalation sweep.
With Jev
◐ ▮▮▮ ▦Without Jev
LLM · CODEExamples
Calls Jev and every baseline for real. Nothing is saved.
Runs every labeled row, one at a time: 18 rows × every variant, all real, paid calls. Rows run sequentially so latency measures the model, not a traffic jam. The run is saved as it goes and lands in History.