06
Tool Selector
jev primitiveChoice per step (multi-step, traced)
labelsget_weathersearch_webcalculatordatabasecreate_ticketsend_emaildatabase > calculatornoneask_user
dataset20 labeled rows
tracingLangfuse
An agent that loops: pick the next tool (or stop), write its arguments, run it, look at the result.
Jev picks from a closed list (6 tools, plus finish and ask_user) and reports a confidence. An LLM writes the
arguments, and only for the chosen tool. The tools are offline fakes, and every loop stops at 4 tool calls. The label
is the whole trajectory (e.g. database > calculator). Traced: every loop turn and every model call, with
its cost, in Langfuse.
With Jev
◐ ▮▮▮ ▦Without Jev
LLM · CODEExamples
Calls Jev and every baseline for real, traced in Langfuse. Nothing is saved.
Runs every labeled row, one at a time: 20 rows × every variant, all real, paid calls. Rows run sequentially so latency measures the model, not a traffic jam. The run is saved as it goes and lands in History.