A customer asking about opening hours may only need a predefined reply. A question about an existing order needs account data, and a more involved request may need an LLM to compose a response. Sending every message to the same model can add cost without improving the answer.
Jev is a decision model from TypeSafe AI. It evaluates the information you provide and returns typed answers and probabilities that your application can use to choose what happens next.
I tried Jev in a WhatsApp chatbot to see whether it could help me decide which messages needed an LLM. In this article, I’ll explain its question types, walk through the integration, and discuss the limitations to consider before using it in your own application.
Jev accepts a state and a set of typed questions. The state contains the information you want it to evaluate, such as a customer’s message or recent conversation history. Its answers can identify a category, assign a score, or estimate the probability that a statement is true.
TypeSafe calls Jev a System One model, a reference to the fast, intuitive thinking popularized by Daniel Kahneman. In practice, it gives developers a way to ask narrowly defined questions and use the answers in application logic.
Your code controls the workflow: it supplies the context, evaluates the returned probabilities, and decides whether to act, request more information, or use another model. Jev doesn’t write the customer-facing response or generate images. Its role in this chatbot is classification and routing.
A request contains a state, which can be a string, JSON object, or array, and the questions Jev should answer about it. Each question specifies the kind of answer you need.
This request, adapted from the TypeSafe quick start, asks Jev to rate a customer’s frustration:
{
"model": "jev-latest",
"state": "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
"questions": {
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
}
}
}
The corresponding example response assigns probabilities to the three levels and returns a score:
{
"model": "jev-1.13.0",
"answers": {
"frustration": {
"type": "score",
"score": 1.0,
"confidence": 1.0,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language"
},
"probabilities": {
"0": 0.0,
"1": 1.0,
"2": 0.0
}
}
}
}
Here, score is 1.0, the position of “Frustrated but civil” on the scale. The probability distribution puts all its weight on that level. The score itself is not a probability.
Jev supports three question types:
| Question type | Use it to | Example |
|---|---|---|
| Noul | Estimate the probability that a statement is true | Does this message need a reply? |
| Choice | Select from a set of categories | Which team should handle this ticket? |
| Score | Rate an input against ordered levels | How frustrated does this customer appear? |
A Noul question returns a number between 0 and 1 representing the model’s estimated probability of a yes answer. A value of 0.95 means Jev estimates a 95 percent probability that the answer is yes; a value near 0 favors no, and a value near 0.5 signals uncertainty.
That estimate still needs validation against your own examples. It does not establish that the model will be correct 95 percent of the time on your workload.
A Choice question selects from the options you supply and returns a probability for each one. For example, you could route a support ticket to billing, technical support, or sales:
{
"department": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": {
"technical": 0.85,
"sales": 0.0,
"billing": 0.15
}
}
}
The selected option is technical, with a probability of 0.85. The separate confidence value summarizes the concentration of the distribution; it is not the selected option’s probability. TypeSafe explains this distinction in its confidence documentation.
A Score question evaluates an input against ordered levels, such as low, medium, and high. The levels are numbered from zero. In the frustration example, the scale runs from 0 for calm to 2 for very angry.
The returned score is a weighted average: multiply each level’s number by its probability, then add the results. In the example above, the calculation is 0 × 0 + 1 × 1 + 2 × 0 = 1.
A score of 1.5 could result from equal probabilities at levels one and two. Other distributions can produce the same average, so check probabilities when you need to understand how the model arrived at a score.
Open the TypeSafe console to check access, try the playground, and generate an API key. The playground is a useful place to experiment with questions before adding them to your application.
TypeSafe provides its own SDK, and Jev is also available through Vercel AI Gateway. The examples below use the Vercel AI SDK. Keep the integration paths separate: installing @typesafe-ai/sdk does not install the ai package used by evaluate.
For the AI SDK examples, install:
npm install ai
Set your Gateway key in your server-side environment:
AI_GATEWAY_API_KEY=your_gateway_api_key
Make sure your runtime loads the environment file. Then send a request:
import { experimental_evaluate as evaluate } from "ai";
const result = await evaluate({
model: "typesafe-ai/jev",
state: "Customer: Where is my order?",
questions: {
request_type: {
type: "choice",
instructions: "Classify the customer's request.",
criteria: {
greeting: "A greeting or social opener",
faq: "A general question answered by the business FAQ",
order: "A question about an existing order",
complex: "A request that does not fit the other categories",
},
},
needs_llm: {
type: "boolean",
instructions: "Does this message require a generative LLM to answer?",
},
},
});
console.log(result);
The current AI SDK evaluation API uses boolean questions and a probability field for yes-or-no answers. TypeSafe’s native API calls this primitive noul. The evaluation API is experimental, so pin the package version you test.
I integrated Jev into a WhatsApp chatbot I was building for potential customers. The original version sent every incoming message to an LLM, which would compose a reply and help the customer place an order.
This was a hobby project with a limited budget. I wanted to check whether a message needed a generated response before paying for an LLM call.
Jev gave me a way to classify messages first. The application could then select a predefined reply or pass the request to the LLM. This is a smaller version of the tradeoff involved in routing requests between models: the extra routing step needs to justify its own cost and latency.
TypeSafe’s model documentation lists Jev v1.13 at $0.042 per million input tokens, with no charge for output tokens. End-to-end latency depends on the request and deployment, so measure the full routing workflow in your own application.
The plan was to place Jev before the LLM, send it a compact conversation state, and use its answers to decide whether to generate a response and which context to include.
For this chatbot, a one- or two-second response would probably be acceptable. The more useful question was whether the routing step could reduce unnecessary LLM calls without sending customers irrelevant replies.
Other applications have tighter latency requirements. A remote model call can exceed the time available for each frame in a game loop running dozens of times per second. In that setting, model calls would need to run asynchronously, outside the frame-critical path.
Start by defining the decision your application needs and what it should do with each answer. The API call is only one part of the integration; the thresholds and fallback behavior determine how useful the result is.
For a simple moderation example, the model can estimate whether a comment violates a specific policy. A negative product review alone should not trigger removal:
const result = await evaluate({
model: "typesafe-ai/jev",
state: {
comment: "This product is terrible and completely useless.",
},
questions: {
shouldHide: {
type: "boolean",
instructions:
"Does this comment contain threats, targeted harassment, or spam? " +
"Negative product feedback alone does not violate the policy.",
},
},
});
// Illustrative threshold, not a validated moderation policy.
if (result.answers.shouldHide.probability > 0.8) {
hideComment();
} else {
showComment();
}
This excerpt assumes evaluate has been imported and that the application defines hideComment and showComment. Test the threshold against labeled comments and decide how to handle uncertain results before using it for automatic moderation.
The chatbot needs several decisions: which category the message belongs to, whether it needs an LLM, and whether the answer depends on conversation history.
The following function builds those questions. It uses the same evaluate import and Gateway setup as the earlier example:
const categoryCriteria = {
greeting: "A greeting or simple social opener",
faq: "A basic FAQ, such as opening hours or products offered",
product_question: "A question about choosing, comparing, or using products",
order_question: "A question about an order, delivery, or an account-specific issue",
general_question: "A question that needs a helpful natural-language answer",
complex_request: "A multi-part, ambiguous, or otherwise complex request",
} as const;
type MessageCategory = keyof typeof categoryCriteria;
type JevDecision = {
category: MessageCategory;
needsLlm: boolean;
needsHistory: boolean;
};
export async function decideMessage(
message: string,
recentMessages: string[],
currentTopic?: string,
): Promise<JevDecision> {
const result = await evaluate({
model: process.env.JEV_MODEL || "typesafe-ai/jev",
state: {
message,
recentMessages,
...(currentTopic !== undefined ? { currentTopic } : {}),
},
questions: {
category: {
type: "choice",
instructions:
"Classify the customer's latest WhatsApp message. Choose the closest category.",
criteria: categoryCriteria,
},
needs_llm: {
type: "boolean",
instructions: "Does this need a language model to compose a useful response?",
},
needs_history: {
type: "boolean",
instructions: "Does answering this depend on the recent conversation context?",
},
},
});
// Illustrative thresholds: validate these against labeled messages.
return {
category: result.answers.category.choice,
needsLlm: result.answers.needs_llm.probability >= 0.5,
needsHistory: result.answers.needs_history.probability >= 0.5,
};
}
The message handler then uses the decision. This excerpt shows the error and quick-response branches:
let decision;
try {
decision = await decideMessage(
message,
state.lastMessages,
state.currentTopic,
);
} catch (error) {
console.error("Jev request failed:", error);
saveTurn(state, message, jevFallback);
return jevFallback;
}
if (!decision.needsLlm) {
const reply =
quickResponses[decision.category] ??
"Thanks for your message! Could you tell me a little more about what you need?";
saveTurn(state, message, reply);
state.currentTopic = decision.category;
return reply;
}
// Continue to the application's LLM response branch.
For a simple workflow, the application may only need one question and a conditional. More involved workflows need explicit rules for uncertain answers, missing data, and API failures. The surrounding application logic deserves as much attention as the prompt. When designing these kinds of AI features that deliver real user value, it pays to define success criteria before writing any code.
Jev’s usefulness depends on the task and how you define the questions. A high probability or confidence value should not be treated as proof that an answer is correct.
Consider this customer message:
I paid for my order yesterday, but I haven’t received it yet. Can you check what happened?
Classifying this as a general question and sending it directly to an LLM would miss a necessary step: retrieving the customer’s order information. Jev can help identify the intent, but the application still needs to fetch the relevant records before answering.
The opposite error is also possible. The model might classify an account-specific request as an FAQ, causing the application to return a predefined response that does not answer the question.
Categories such as general_question and complex_request can overlap. Define their boundaries, test ambiguous messages, and provide a fallback when the available options do not clearly fit. Account-specific questions should also have a route to data retrieval, regardless of whether an LLM is needed to phrase the response.
TypeSafe documents limitations in Jev v1.13, including tasks involving counting, dates, and multistep reasoning. Use ordinary application logic when the answer can be calculated from known data.
For example, the chatbot should calculate an order’s age directly. If the business’s delivery policy defines orders older than three days as delayed, the application could use:
const daysSinceOrder = getDaysSince(order.createdAt);
if (daysSinceOrder > 3) {
handleDelayedOrder();
}
The helper functions and delivery rule belong to the application. A real implementation should account for the relevant business days, time zones, and promised delivery date.
Jev can help determine whether the customer is asking about a delayed order. It does not need to calculate how many days have passed.
Jev is worth testing when your application makes frequent judgments over unstructured text and can express the outcomes as categories, scores, or yes-or-no questions. Message routing is one example; ticket triage and document classification follow a similar pattern. Teams building more sophisticated pipelines may also want to look at multi-turn AI agents with Genkit’s Agents API for cases where a single classification step is not enough.
It adds less value when a rule or calculation already answers the question reliably, or when nearly every request still needs the same LLM call afterward. In the latter case, you may be adding a request without reducing much work.
Before adopting it, compare the proposed workflow with your existing one using representative messages. Track incorrect routes, fallback frequency, total cost, and end-to-end latency. The relevant result is whether the whole application improves. Tools for comparing AI agent sandbox platforms can help you evaluate and isolate individual routing steps in a controlled environment.
My takeaway from the chatbot integration is that Jev’s value depends on the decisions you ask it to make and the logic you build around its answers. The same principle applies to other AI routing workflows: define the routes clearly and make failures visible. If you are managing the product side of this work, understanding the security trade-offs PMs must own before launch is equally important.
Jev gives applications typed answers they can use for classification and routing. In my WhatsApp chatbot, that meant evaluating incoming messages before deciding whether to ask an LLM to write a reply.
Start with one well-defined decision, test it against messages you can label, and compare the results with your current approach. That will tell you more about whether Jev fits your application than a demo or a low per-token price alone. You might also find it useful to review when to use skills vs. MCP tools for AI agents as your integration grows more complex.

Tailwind CSS component libraries provide pre-built components to streamline the process of developing aesthetic, user-friendly interfaces.

Vibe coding makes building apps faster, but speed without engineering can create serious failures. Here are five real-world examples and practical checks to help you ship AI-generated code safely.

Compare React ViewTransition and Motion across four animation patterns. The native build saved ~38.6kB of gzipped JS, but required 179 more lines of CSS.

Learn how to run Cline locally with Ollama to build a secure, private AI coding agent that keeps your proprietary code off cloud servers.
Hey there, want to help make our blog better?
Join LogRocket’s Content Advisory Board. You’ll help inform the type of content we create and get access to exclusive meetups, social accreditation, and swag.
Sign up now