Published

Inside MindsHub

A quicker way to check whether an agent finished

In one of our test conversations, a user asks an agent to check last month’s AWS spend. The lookup fails with a rate-limit error. The agent replies anyway:

“Last month’s AWS bill came out at $214,462.28.”

There is no evidence for that number anywhere in the conversation. The reply looks finished. The task isn’t.

This is the kind of failure we need to catch in Anton, the open-source agent harness behind MindsHub Agents. Anton reads files, runs code, searches the web, and writes answers. Sometimes it stops after doing only part of the work. Sometimes a tool fails, and it responds as if it had succeeded.

We use a second, cheaper LLM to check whether Anton has finished. We wanted to find out whether a decision model could do that job faster and at a lower cost. So we tested TypeSafe’s Jev against our existing check on 990 generated conversations.

Defining finished work

Every Anton task that uses a tool ends with a completion check. The checker reads a compact view of the conversation and chooses what should happen next:

  • COMPLETE: the work is done, so Anton ends the task.

  • WAITING: Anton needs something from the user, so it ends the current run and waits. We also use this label for refusals of unsafe requests.

  • INCOMPLETE: Anton stopped early but can continue. We send it back to work with a note about what is missing.

  • STUCK: a blocker prevents further progress. Anton explains the blocker and stops.

Those distinctions matter. A successful query that finds nothing can be complete. A refusal ends the interaction without satisfying the original request. An agent that gives up after one failed lookup may still have another way to get the answer.

Jev looked like a reasonable fit for this part of the system. It returns structured decisions and probabilities. We could give it our four labels and the rules for choosing between them.

Making the test harder to pass

We needed conversations with known expected labels and no customer data. We wrote a generator with 17 scenario recipes, varying the requests, tools, errors, and wording across 16 data sources. It covered finished work, partial answers, empty query results, refusals, and failures like the unsupported AWS bill.

The team reviewed the set before we relied on it. We also had a second model review 37 cases from an earlier version. It used the same rules but couldn’t see our labels. It agreed with every label. It also flagged 15 cases as too easy.

Some replies repeated the checker’s own rules. Others practically announced that the work was unfinished. Many cases differed mainly in their numbers: once we masked those numbers, 1,000 examples became about 340 distinct conversations.

We rewrote the generator. It now rejects replies that echo the rules and repeated request-and-reply pairs within a recipe. Tool results identify their sources. Access requests appear only for private sources, and missing-library errors are more realistic.

The revised set contained 990 distinct examples. They still came from 17 recipes, so this was a test of known situations, with more difficult and unfinished tasks than everyday traffic. It wasn’t a substitute for evaluating real conversations.

A faster check with some telling mistakes

We ran every case once through each model, for 1,980 judgments. The existing checker used the MindsHub Air model (mindshub_air), through Anton’s production code path, including its forced tool call and output budgets. Jev received a four-way Choice using the same production rules, plus a Noul question for the check’s extra yes/no field.

Both received the same transcript through the production MindsHub gateway from the same client. Prompts were about 1,700 tokens, compared with roughly 2,700 in production. Timings included the gateway and network. We scored the completion labels; we didn’t evaluate the LLM’s written explanations or the extra yes/no field.

Measure Current LLM check Jev
Accepted completion label 976 / 990 (98.6%) 982 / 990 (99.2%)
Unsupported numbers labeled COMPLETE 0 / 160 0 / 160
Wrong continue-or-stop action 5 / 830 0 / 830
Median response time 2.56 s 0.78 s
95th-percentile response time 3.85 s 0.95 s
Cost per 1,000 checks about $0.40 about $0.07

Jev’s median response was about three times faster, at roughly one-sixth of the cost per check. These are single-call costs, before fallbacks, calculated using the benchmark’s list prices: $0.20 per million input tokens and $1.20 per million output tokens for mindshub_air; $0.042 per million input tokens with free output for Jev.

Both models rejected all 160 unsupported-number cases. For those cases, our rules accepted either INCOMPLETE or STUCK. They tested whether the checker rejected an unsupported answer, without fixing whether continuing or stopping was right.

The remaining 830 cases specified a single continue-or-stop action. Jev got that action right every time in this run. Its eight label errors all swapped one stopping label for another: six requests for user input became STUCK, and two refusals became COMPLETE. Those errors could still affect what the user sees.

The existing LLM made nine similar swaps between stopping labels. Its other five errors changed the action: three finished answers were sent back to work, one premature stop was accepted, and one hard blocker prompted another attempt.

That was encouraging, but it left a practical problem. When Anton needs to continue, it needs a note explaining what is missing. When it is stuck, it needs an explanation for the user. Jev doesn’t write those explanations.

A shortcut for clearly finished work

We could still use Jev to recognize clearly complete tasks and leave the other cases with the existing checker.

To explore that, we looked at the probability Jev assigned to COMPLETE. Its separate confidence field describes the shape of the probability distribution; it isn’t a measured success rate on our task.

At a COMPLETE probability of 0.80 or higher, Jev selected 306 cases. All 306 were correctly classified as complete. That covered 306 of the 310 finished tasks, about 99%. The two refusals it had labeled COMPLETE didn’t reach this threshold.

That suggests a useful shortcut. It doesn’t establish a production threshold. The uncertain cases included refusals and successful queries that found nothing, both of which need to appear in the next evaluation.

Our proposed rollout would work like this:

  1. Start in shadow mode. After our data-processing review, run Jev alongside the current check and record its decisions. Keep the existing check in control while we review real traffic.

  2. Choose the threshold on human-labeled conversations. Confirm it on a separate set that wasn’t used to choose it.

  3. Enable the COMPLETE shortcut only after validation. Let qualifying completions end the task. Send everything else, including Jev errors and timeouts, to the existing LLM check.

We would also need to preserve the check’s extra yes/no behavior, which this benchmark didn’t evaluate. Giving WAITING its own shortcut would require separate validation.

The savings depend on the fallback

The tempting number is the difference between $0.40 and $0.07 per 1,000 checks. But a system that calls Jev first pays for Jev on every case, and for the LLM whenever it falls back.

Our 306 qualifying completions were only 31% of the test set. With the remaining cases going to the LLM, the estimated cost would be about $0.35 per 1,000 checks. That’s a saving of roughly 13%.

Production has a different mix. Over two weeks in September, 73% of customer completion checks came back COMPLETE on the builds most customers were running. A newer release with stricter rules showed 60%, though on far fewer checks.

If the threshold caught about 99% of real finished tasks as it did in the generated set, roughly 59% to 72% of checks could skip the LLM. That would put the estimated cost saving at about 40% to 55%. Those are projections for completion checks, not measured production savings or reductions in the cost of the whole agent run.

Latency needs a separate measurement. A shortcut case avoids the LLM call; a fallback case waits for both calls in sequence. Running them in parallel avoids the extra fallback wait but pays for both every time.

And Anton’s answer is already visible when this check runs. A faster check can help the agent finish or resume sooner. It won’t make the first answer appear sooner.

What other agent builders can use

Start with the decision your code needs to make and the action each label triggers. Include cases where a confident answer has no evidence, where a successful query returns nothing, and where the agent needs the user to respond. Review the test data for clues that make the answer too easy.

Then look for a narrow decision you can validate independently. For Anton, that is recognizing clearly finished work while keeping the existing checker available for everything else. Measure how often the shortcut applies, what happens when it fails, and the cost and latency of the combined system.

Jev gave us a reason to test that design on real traffic. The next step is to find out whether it can recognize finished work just as reliably in long, messy conversations, where the rules are harder to apply and an unsupported answer may look perfectly reasonable.