Grading agent findings with Jev and openevals

5 October 2026

Planet volumes dl7 Cav4 v0 E unsplash

I have been building evals for an autonomous agent that produces structured findings and a written report. The deterministic checks were easy enough: does the output parse, are the required fields present, does the prose pass our house style gate? What none of them can tell me is whether the evidence really supports the finding it is attached to. Is the severity honest? Does the report read like a person wrote it? Those are judgements, and for a while the obvious answer was to ask a general-purpose LLM to make them.

What drew me to Jev was how closely its design fitted those questions. Each one asks for a decision against criteria I can define, with an answer my code needs to act on. Jev is TypeSafe's first System One model, their term for a model built for that kind of bounded judgement. A general-purpose LLM generates text, which can include an answer constrained to a schema; Jev works within a defined answer space and returns typed values and probabilities directly. That gave me a way to assess the agent's findings while leaving the investigating and writing to the agent itself.

This post explains how I use those typed questions, how they fit the openevals evaluator interface and what the tests can actually tell me.

What I wanted from a judge

Structured outputs can make a chat model's answer easy to consume. Bind a schema and you can get back {"genuine": true}. The schema does not tell you how much uncertainty sits behind that answer, though. Asking for a confidence number adds another output to validate against examples where you know the answer.

I also wanted to ask several small questions about every finding: is it genuine, does the evidence support it, is it a duplicate, is the severity justified? The cost of judging matters when you repeat those questions across a suite. It influences how much you check and how often you can afford to check it.

Jev appealed to me because typed decisions and probability distributions are part of its interface, alongside the prospect of making those repeated judgements much more cheaply. Lower judging costs would let me check more findings and run the suite more often. The interface removes a whole category of faff too, although it still leaves me responsible for deciding whether its judgements are useful on my data.

The questions Jev answers

You give Jev a state, which is the material to assess, and a set of typed questions. There are three question types:

  • A Choice selects an option from criteria you define. It returns the chosen key, probabilities for the options and a confidence value describing how concentrated that distribution is.

  • A Noul answers a yes/no question with the model's probability that the answer is yes. You can supply explicit descriptions of what counts as true and false.

  • A Score assesses the state against ordered descriptions. It returns a probability-weighted mean of their level indices, so the result can fall between levels. In my evaluator, I normalise that result to 0–1.

A Noul suits the question of whether a finding is genuine. A value near one is a strong yes; near zero is a strong no. A value near the middle means the model assigns similar probability to both answers. It is not a measure of how serious the issue is. That needs a separate question with its own criteria.

For detection, I use one Choice per known objective: which finding, if any, reports this issue? I include an option for no match. A finding only counts as a match when it is the top choice and the distribution is sufficiently concentrated. An uncertain match counts as a miss in my metric, which makes it a conservative measure of detected objectives. It also means a low score can reflect uncertainty in the judge as well as something the agent missed.

I sum those accepted matches in code and divide by the number of objectives. There is no benefit in asking the judge to repeat arithmetic I can already do from its answers.

Where predefined answers become a limit

The answer space needs to be defined before each call. I can build the Choice options dynamically from the findings in a run, but Jev cannot invent an additional option or write an explanation of a problem I did not ask about. That makes it a poor fit for a request such as “review this report and tell me what I have missed”. I need to turn that into specific checks, or use a person or a generative model to explore it first.

I do not need to anticipate every possible input, but I do need to give unexpected cases somewhere to go. For a Choice, that can mean an explicit other or none of the above option, as the TypeSafe documentation recommends, with a route to further review. Low confidence can also trigger review, but it is not a substitute for a missing option: a concentrated distribution only describes the choices I supplied. In this evaluator, Jev can help match findings to known objectives. Discovering an issue outside those objectives still depends on the agent and on how I review its output.

Wiring it into openevals

The openevals evaluator interface accepts a callable with keyword arguments such as outputs and reference_outputs. It can return one result dictionary or a list. My evaluators return a list with one entry per metric, using key, score, comment and metadata, and I keep the numeric scores in the range 0–1.

The pattern I settled on is a factory that binds a judge and returns that callable. Here is a reduced example for one metric, including the connection to the Jev client. Each finding is represented as a dictionary containing the claim and the evidence available to assess it.

from typing import Any, Mapping, Protocol

from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient


GENUINE = Noul(
    instructions=(
        "Does the supplied evidence establish that `finding` is "
        "a genuine issue worth reporting?"
    ),
    criteria=NoulCriteria(
        true="The evidence establishes a real, reportable issue",
        false=(
            "The finding is a false positive, an informational note "
            "or a claim not established by the supplied evidence"
        ),
    ),
)


class Judge(Protocol):
    def ask(
        self, state: Any, questions: Mapping[str, Any]
    ) -> Mapping[str, Any]: ...


class JevJudge:
    def __init__(self, client: TypeSafeClient, model: str):
        self.client = client
        self.model = model

    def ask(self, state, questions):
        response = self.client.system_one(
            model=self.model,
            state=state,
            questions=dict(questions),
        )
        return response.answers


def create_genuineness_evaluator(judge: Judge, *, threshold: float):
    def evaluator(*, outputs: list[dict[str, Any]], **_):
        if not outputs:
            raise ValueError("Genuineness rate needs at least one finding")

        accepted = 0
        for finding in outputs:
            answers = judge.ask(
                state={"finding": finding},
                questions={"genuine": GENUINE},
            )
            accepted += answers["genuine"].noul >= threshold

        return [{
            "key": "finding_genuine_rate",
            "score": accepted / len(outputs),
            "comment": (
                f"{accepted} of {len(outputs)} findings accepted "
                f"at noul >= {threshold}"
            ),
            "metadata": {"threshold": threshold},
        }]

    return evaluator


# Illustrative input and threshold, not a calibration result.
with TypeSafeClient() as client:
    judge = JevJudge(client, model="jev-1.13.0")
    evaluator = create_genuineness_evaluator(judge, threshold=0.7)
    result = evaluator(outputs=[{
        "claim": "The export endpoint returns an error",
        "evidence": "Captured response: GET /export returned HTTP 500",
    }])
    print(result)

This example refuses to calculate a rate over an empty set. A run with no findings needs a separate check against the expected objectives; an empty report should not quietly earn a perfect quality score. The example also lets API failures surface as errors, so they cannot be mistaken for findings that failed the rubric.

In the fuller evaluator, I build a filtered state containing the finding, its cited evidence and relevant peers. Several questions share that state in one request. Jev supports evaluating those questions together, and its input-token pricing makes that an attractive way to organise the work. Batching is useful with other judges too; the important thing is to measure the cost of the whole suite with representative inputs.

The Judge interface lets me inject a stub with canned answers. Those offline tests check threshold boundaries, aggregation and result formatting without a network call. They tell me whether the harness handles answers correctly. Assessing the quality of the answers requires the real judge and labelled examples.

Deciding whether to trust the answers

Every question traces back to a line in our grading rubric. That gives me something concrete to inspect when the judge and I disagree: the evidence it saw, the wording of the question and the rule I expected it to apply.

My initial checks use a known-good finding and deliberately corrupted variants. Removing the supporting evidence should lower the evidence-support result. Inflating severity should lower the result for whether that severity is justified, without necessarily changing whether the underlying issue exists. These are useful sanity checks on individual questions.

They are not enough to establish calibration. That needs a representative set of labelled cases, including ambiguous findings, plausible false positives and genuine issues expressed in different ways. To claim that a probability of 0.8 has that meaning on this task, I would need to check whether comparable predictions are correct at roughly that rate. A few clean examples cannot show that.

There are also two different thresholds to choose. The per-finding threshold decides which answers count as accepted. A baseline floor decides whether the aggregate result is acceptable for a run. I keep those decisions separate: choosing a sensible acceptance threshold does not establish how many weak findings a report should be allowed to contain. Thresholds should be selected on labelled examples and checked on held-out cases before being used to gate changes.

I pin the judge model to jev-1.13.0 and record the version alongside the results. Pinning is useful for any judge that supports it. It removes one source of change when comparing runs, although it does not by itself establish that repeated answers will be identical. Changes to the model, rubric or supplied context need their own validation.

What the saved runs can tell me

I separate generating an agent run from grading it. The end-to-end run produces a captured workspace, including the findings and evidence, and the judge can work against that fixture repeatedly.

That separation makes ad hoc review easy too. I can ask a coding agent to review the captured findings and evidence, or open the workspace and work through them myself, without having to run the original agent again. If a score looks odd, the material behind it is there to inspect.

It also helps when changing the evaluator. I can hold the agent output fixed while checking a different question or threshold. The offline stub tests can run on every change, while checks that call Jev need credentials and incur judging costs.

There is an important limit here. Regrading a captured run cannot reveal a regression introduced by a new agent prompt, tool description or model. To evaluate that change, I need fresh outputs from the changed agent on comparable tasks, graded with a fixed evaluator. Since my end-to-end runs happen less often, the saved-fixture checks do not amount to an agent-quality gate on every change.

What I would carry forward

The useful result so far is an evaluator whose questions are explicit, whose answers are straightforward to consume and whose aggregation I can test offline. That is enough to make Jev worth using in this harness. A broader claim about matching frontier-model accuracy would need a comparison on labelled findings, with costs and failure cases reported alongside the scores.

If you are building something similar, start with a judgement you can describe clearly and a collection of examples you can label yourself. Check where the judge disagrees, including the awkward cases near your acceptance threshold. Then connect it to fresh agent runs before relying on it to approve changes. I am interested in taking the same approach into customer applications, but that validation work needs to follow it into each new task.

Matthew Wilson

Senior Architect & AWS Community Builder