Fouad Wahabi
Software Engineering Lead
Alex Barksdale
Senior Software Engineer
Miguel Tulla Lizardi
Software Engineer
TypeSafe AI released Jev in September 2026 to do one thing: make decisions. Give it a state (a string or a JSON object) plus a set of typed questions, and it returns typed answers with probabilities. It never explains itself, and that constraint is the whole idea.
Evaluation pipelines have spent the last two years asking text generators for yes/no verdicts, wrapping the reply in a JSON schema, and paying generation prices for what amounts to a single bit. Jev is built for that bit, which makes it a good fit for the two evaluation surfaces in Datadog Agent Observability: online evals, which score production spans as they arrive, and experiments, which score a dataset offline. The same rubric can drive both.
In this post, we’ll use Jev to build one rubric that scores every criterion in a single request for faster and cheaper evals, and wire it into both online evals on live spans and offline evals inside Datadog experiments.
What Jev returns
Each question for Jev is identified by a key, and its answer is returned under that key. Jev supports three question types, and each one returns a different kind of answer:
| Question Type | Returned Signal | Example evaluation |
|---|---|---|
| Noul | Probability that a yes/no proposition is true | Are the reply’s policy claims supported by the retrieved excerpts? |
| Choice | Selected category, probabilities for all categories, and confidence | What is the reply’s main failure mode? |
| Score | A probability-weighted average of rubric levels, plus the level distribution and confidence | How severe is the potential customer impact? |
A Noul probability expresses uncertainty about a proposition; it does not measure how much of the reply is correct. Choice probabilities and Score confidence summarize the spread of their probabilities, though keep in mind that a confident answer can still be wrong.
Questions in the same request are evaluated independently against a shared state. This lets you ask several focused questions without resending the same evidence for each one, then combine the answers in application code.
Putting Jev to the test
We’ll use a support agent for a fictional airline, Vega Air, as an example for using Jev. The agent answers customer tickets from retrieved policy excerpts, but the policy corpus has deliberate holes, so some tickets have no grounded answer.
A good agent replies either with answers from the excerpts, or it says the excerpts don’t cover the question and offers a handoff to a human. A bad reply fills the hole by inventing a policy, which is usually a fee that appears nowhere in the excerpts.
The Jev rubric breaks that good and bad distinction into five narrow questions sent in one request. Two of them are below and are a Noul and a Choice question type. Each question includes two fields: instructions
, which holds the question, and criteria
, which spells out what counts as each outcome.
instructions
can be a plain string or a JSON object. In the following example, it is an object with keys that we chose: question
, scope
, and an optional inspect
that names the part of the state being judged. failure_mode
skips inspect
because its question already names reply
. You can find the other three questions, and a walkthrough of each field, in our Jev rubric notebook.
The unclear
option in failure_mode
is deliberate. A Choice question always returns the option with the highest probability, so Jev never abstains. If an evaluation needs a way to say “cannot judge this one,” that outcome has to exist in the criteria. TypeSafe’s self-consistency cookbook uses the same pattern for moderation.
Here’s a real response. The ticket asked about cancellation compensation when the retrieved policy only covers delays, so the agent declined to answer and offered a handoff:
The interesting part sits inside failure_mode
. Jev picked none
, but none
at 0.46 and partial_answer
at 0.42 are nearly tied, and confidence came back at 0.34. Flattening that to the string none
throws the interesting part away. A near-tie between two categories is a signal in its own right, and a natural trigger for routing the trace to a human reviewer.
The composite verdict stays in application code:
That composite verdict could have been a sixth Jev question. Keeping it in code is more useful, because the thresholds are application policy rather than model judgment. Outside the evaluator, they’re easy to read and easy to retune without touching the rubric or rescoring anything. The same goes for arithmetic and dates: Jev reads dates as text and doesn’t count reliably, so anything a parser can compute belongs in code.
The call itself is one request against a pinned model:
TypeSafeClient
reads TYPESAFE_API_KEY
from the environment, so nothing else needs configuring. The state is three named fields rather than the whole trace. Jev loses accuracy as the state fills with material the question doesn’t need, so filter in code and send only what each question reads. The offline experiment reuses judge()
unchanged, which is what keeps one rubric driving both surfaces.
Scoring live spans with online evals
For online evals, the appeal of Jev is that you pay for the verdict and nothing else. Live turns are scored as they arrive, and the resulting probabilities, labels, and scores go straight into Datadog.
The traced application never imports Jev. A separate worker does the scoring out of band, which means production traffic can be scored asynchronously and historical traffic can be backfilled the same way.
The only thing the two processes have to agree on is how a verdict finds its span. Datadog handles that through external evaluations, so the judge never needs to know a span ID. The application tags each span with a domain key (e.g., turn_id
) and the scorer joins on that tag.
In this example, the scorer reads a JSON Lines file, standing in for whatever queue, table, or store already holds the turns you want judged:
Passing the turn’s own timestamp_ms
rather than the current time keeps reruns idempotent. Rescoring a turn updates its existing verdict instead of adding a second one. Both halves end-to-end can be found in the online evals notebook in our GitHub repo.
Mapping Jev answers to Datadog metrics
LLMObs.submit_evaluation
accepts four metric types: score
, categorical
, boolean
, and json
. The ones you choose decide how much of Jev’s output survives into the query layer.
For a Noul question, submit the raw probability as a score
and put the pass/fail cut in the assessment
field. Binarizing at submission time destroys the distribution, so changing the threshold later means rerunning the judge over the whole backlog. Keeping the probability turns a threshold change into a query change. Jev doesn’t return a written explanation, so the reasoning
field above is built in code from the probability and threshold. Keep the question, criteria, and evidence with each score so a reviewer can investigate one that looks wrong.
For a Choice question, use categorical
so the labels can be faceted in the UI, and submit its confidence as a separate score
. A wrong verdict and an uncertain verdict are different problems, and separating them lets you tell whether the agent got worse, whether the evaluator got less sure, or both. Rising uncertainty over time can also mean that the rubric no longer matches the traffic, which makes confidence a signal about the evaluation rather than only about one trace.
Tag every metric with judge_model
. Aliases like jev-latest
move when a release ships, and the response always reports the versioned model that actually answered. Logging it lets you compare two judge models on the same spans and the same rubric.
Using Jev inside a Datadog experiment for offline evals
A Datadog Agent Observability experiment takes a dataset, a task, and a list of evaluators. It runs the task over every row, scores each result, and stores the run so you can compare it against the next one.
The interface expects one evaluator object per metric. Ported naively, that’s one Jev request per evaluator per row. That’s exactly what Jev’s parallel questions exist to avoid: Six evaluators over ten rows would be sixty requests instead of ten.
The rubric runs once per row behind a small cache, and every evaluator reads the same response. One detail matters in that cache: guard the dictionary, not the request. experiment.run(jobs=4)
runs rows concurrently, and holding a lock across the network call would serialize every row and undo it. Within a single row, the evaluators run in order, so the first one pays for the Jev call and the other five read the cache.
The evaluators themselves are thin:
Wiring it up looks like any other experiment:
The offline surface can also ask something that the online one cannot. Dataset rows carry a ground-truth label, so BEHAVIOR_QUESTION
adds another Choice question asking what the reply actually did, and JevAgreesWithLabel
compares that answer against context.expected_output
. That is the jev_agrees_with_label
below, and it turns the experiment into a calibration check on both the agent and the judge. Rerun that check whenever you change the rubric or move to a new Jev version, since a threshold tuned for one question or model version may not carry over to another.
The cache class, all six evaluators, and the dataset setup are in our third Jev rubric for experiments.
Get started with Jev for your Datadog evals
You can access Jev directly or through AI gateways like OpenRouter or Vercel AI Gateway. TypeSafe provides Python and JavaScript SDKs, as well as a TypeSafe agent skill that gives coding agents the full API context. Like any judge, Jev has known limitations, so measure its agreement with human reviewers and its repeatability on your own traffic before you rely on it.
On the Datadog side, everything runs on the released ddtrace>=v4.5.0
package. LLMObs.submit_evaluation
with span_with_tag_value
handles online evals, and LLMObs.experiment
with BaseEvaluator
subclasses handles offline evals. There’s no preview builds and no private endpoints. Clone the repo, add your keys, and run them in order:
Our GitHub repo contains three notebooks that are a complete, runnable version of everything in this post:
1-jev-rubric.ipynb builds the rubric one question at a time and reads the answers. This notebook needs a TypeSafe key to get started.
2-online-evals.ipynb traces the agent, then scores the spans out of band.
3-experiments.ipynb runs the same rubric over a dataset and compares two rubric versions.
Check out our Agent Observability documentation to learn more about monitoring and evaluating your LLM applications. And read our blog on using evaluation frameworks with Agent Observability.
If you’re new to Datadog, get started with a free 14-day trial.
Facts Only
* TypeSafe AI released Jev in September 2026.
* Jev is designed to return typed answers with probabilities based on a provided state and typed questions.
* Jev supports three question types: Noul, Choice, and Score.
* Noul returns the probability that a yes/no proposition is true.
* Choice returns a selected category, probabilities for all categories, and confidence.
* Score returns a probability-weighted average of rubric levels, level distribution, and confidence.
* Jev is integrated into Datadog Agent Observability for online evaluations of production spans and offline evaluations in experiments.
* Datadog LLMObs accepts four metric types: score, categorical, boolean, and json.
* TypeSafe provides Python and JavaScript SDKs for Jev.
* Datadog requires ddtrace >= v4.5.0 for these evaluations.
Executive Summary
TypeSafe AI has introduced Jev, a specialized decision-making tool designed to replace expensive text-generation processes with high-efficiency, typed evaluations. By returning probabilities rather than natural language explanations, Jev reduces the cost and latency associated with scoring LLM outputs. The system is specifically tailored for integration with Datadog Agent Observability, allowing developers to apply a single evaluation rubric across both live production traffic and offline datasets.
The technical implementation focuses on separating model judgment from application policy. While Jev provides raw probabilities and category confidence, the final "verdict" or threshold for success is handled in the application code. This allows for the retuning of success criteria without the need to rescore historical data. The workflow involves scoring production spans asynchronously through external evaluations and running offline experiments to calibrate the judge against ground-truth labels. While Jev increases efficiency, its accuracy depends on the quality of the state provided and the precision of the rubric's instructions.
Full Take
The strongest version of this narrative is that the "LLM-as-a-judge" paradigm is evolving from clumsy, conversational critiques toward a streamlined, probabilistic API. By stripping away the "explanation" phase of AI evaluation, the process becomes a measurable engineering metric rather than a subjective conversation, enabling scalable observability.
However, this is a classic vendor advertorial. The central persuasion vector is the promise of efficiency and "one rubric for both surfaces," using the vendor's own product as the sole evidence for the validity of this approach. The narrative frames the current state of evaluation pipelines as wasteful ("paying generation prices for what amounts to a single bit") to position Jev as the inevitable solution. This creates a closed loop where the tool provides the data, and the tool's own metrics define the success of the system.
Patterns detected: ARC-0043 Authority Game
The underlying paradigm is the "quantification of quality." It assumes that complex human linguistic nuances (like "truth" or "severity") can be accurately flattened into a probability distribution. The second-order consequence is a potential atrophy of human critical review; when a "confidence score" is available, there is a strong temptation to trust the number over the actual text.
If this were an influence campaign, the playbook would involve creating a "problem" (the cost of LLM evals) and offering a proprietary "standard" (Jev) to solve it, thereby locking the user into a specific observability ecosystem. The content matches this structural pattern, as it is authored by the vendor to drive adoption of their SDK and API.
Bridge Questions:
1. If a model is "confident" but wrong, how does a probabilistic score protect the user more than a textual explanation would?
2. What happens to the objectivity of an evaluation when the judge's thresholds are shifted in code after the data is collected?
Counterstrike Scan: The content matches the structural pattern of a vendor-driven narrative designed to establish a proprietary standard for a new technical problem.
Sentinel — Human
This article reads as a detailed, technically grounded explanation of a specific methodology for LLM evaluation, strongly suggesting authorship by a subject matter expert integrating proprietary framework details.
