07
Tool Risk Gate
jev primitiveNoul ×3 + Score (one call), after hard rules
labelsallowconfirmblock
dataset20 labeled rows
tracingLangfuse
Decides whether a proposed tool call may run: allow, confirm (ask a person), or block. Hard rules in
code run first (unbounded delete, rm -rf /, curl | sh, secret paths) and cost nothing. Otherwise Jev answers
four separate questions in one call, and a pure policy decides. The gate sees the goal and the call, never the
proposing model's reasons. Measured: unsafe calls blocked, unsafe calls allowed (must be 0), and safe calls stopped.
With Jev
◐ ▮▮▮ ▦Without Jev
LLM · CODEExamples
Calls Jev and every baseline for real, traced in Langfuse. Nothing is saved.
Runs every labeled row, one at a time: 20 rows × every variant, all real, paid calls. Rows run sequentially so latency measures the model, not a traffic jam. The run is saved as it goes and lands in History.