Picture a support ticket that reads, “Hi, my order ORD-1002 still hasn’t arrived, and the tracking page hasn’t moved in days. What’s going on?” If you send that ticket to Claude Haiku 4.5 with no tools, the model can’t answer it. It asks the customer for details your store already has:
To best assist you, I'll need a bit more information: 1. **When was the order placed?** And **when was it supposed to arrive?** 2. **Where is it being shipped to?** (city/state is fine) 3. **What carrier is it using?** (UPS, FedEx, USPS, etc.) 4. **What does the tracking status currently show?** (in transit, out for delivery, stuck, etc.)
A longer prompt won’t fix this, because the order data lives behind your API and has to be fetched for each ticket as it arrives. To answer properly, the model has to call that API, read the result, decide what to call next, and stop once it has enough to act on. That cycle is the agent loop. An agent framework is the library that runs the loop for you, along with the tool definitions, conversation state, limits, and streaming that surround it.
In this article, you’ll explore a support triage agent by hand first, so you can see exactly what a framework takes off your plate. Then you’ll rebuild the same agent in five TypeScript frameworks: the Vercel AI SDK, Mastra, LangGraph.js, the OpenAI Agents SDK, and the Claude Agent SDK. Every version uses the same two tools, the same order data, and the same model, Claude Haiku 4.5, so any difference you see comes from the framework itself. You can clone the companion repository and run every example yourself.
Your agent reads a ticket, finds the order reference in it, and looks up that order. It then checks the order against your escalation policy and returns a decision. Both tools return canned in-memory data, so the only thing you need to run the examples is an Anthropic API key.
Each tool is a plain object in shared/tools.ts, and every framework in this article wraps the same object instead of redefining it. Here is the order lookup:
// shared/tools.ts
...
export const lookupOrder = {
name: "lookup_order",
description:
"Look up the live status of a customer order by its order reference (format ORD- followed by digits). Returns status, carrier, days in transit, and order value in USD, or found: false.",
inputSchema: z.object({
orderId: z.string().describe("Order reference exactly as the customer wrote it, for example ORD-1001"),
}),
async execute({ orderId }: { orderId: string }) {
toolLog.push({ tool: "lookup_order", input: { orderId } });
const order = ORDERS[orderId.trim().toUpperCase()];
return order ? { found: true as const, ...order } : { found: false as const, orderId };
},
};
The toolLog line records every call so the comparison script can report which tools each framework ran. The second tool, get_escalation_policy, takes the status that lookup_order returned and gives back the rule for that status. Whichever framework runs the loop, every run has to end with a decision in this shape:
// shared/agent-config.ts
...
export const Decision = z.object({
action: z.enum(["resolve", "escalate", "ask_customer"]),
orderId: z.string().nullable(),
reason: z.string(),
customerReply: z.string(),
});
A single model call can’t produce that decision, because the second tool needs what the first one returns, and you can’t know which branch applies until the lookup comes back. A delivered order resolves, and an order that has been in transit for more than seven days escalates. A reference that doesn’t exist means the agent should ask the customer instead of guessing, and any order worth $1,000 or more escalates no matter its status.
To compare the frameworks fairly, the repository runs every version against the same set of tickets, three times each. The set covers a late order, a delivered order, and a high-value order. It also covers a ticket with no order reference, a reference that doesn’t exist, and a ticket that mentions two orders. One conversation sends a follow-up message that corrects the order number, and a final run repeats the late-order ticket with the agent limited to two model calls. You’ll see those tickets come up again as each framework is tested.
Before you reach for a framework, it helps to see what the loop looks like without one. The baseline in 01-baseline/agent.ts uses only the Anthropic SDK. It starts by turning the shared tools into two things the loop needs: a lookup table of handlers and the tool definitions the API expects:
// 01-baseline/agent.ts
import Anthropic from "@anthropic-ai/sdk";
import { z } from "zod";
import { Decision, INSTRUCTIONS, MAX_MODEL_CALLS, MODEL_ID } from "../shared/agent-config";
import { getEscalationPolicy, lookupOrder } from "../shared/tools";
const client = new Anthropic();
const handlers: Record<string, (input: unknown) => Promise> = {
lookup_order: (input) => lookupOrder.execute(lookupOrder.inputSchema.parse(input)),
get_escalation_policy: (input) => getEscalationPolicy.execute(getEscalationPolicy.inputSchema.parse(input)),
};
const tools: Anthropic.Tool[] = [lookupOrder, getEscalationPolicy].map((t) => ({
name: t.name,
description: t.description,
input_schema: z.toJSONSchema(t.inputSchema, { target: "draft-7" }) as Anthropic.Tool.InputSchema,
}));
Each handler validates the model’s arguments with the tool’s Zod schema before running it. The loop itself sends the conversation to the model and adds the reply to the history. When the model stops asking for tools, the reply is the final answer, and output_config makes the API return it as JSON that matches Decision:
// 01-baseline/agent.ts
...
export async function triage(messages: Anthropic.MessageParam[], maxModelCalls = MAX_MODEL_CALLS): Promise {
for (let call = 0; call < maxModelCalls; call++) { const response = await client.messages.create({ model: MODEL_ID, max_tokens: 1024, system: INSTRUCTIONS, messages, tools, output_config: { format: { type: "json_schema", schema: z.toJSONSchema(Decision, { target: "draft-7" }) } }, }); messages.push({ role: "assistant", content: response.content }); if (response.stop_reason !== "tool_use") { const text = response.content.find((block) => block.type === "text")?.text ?? "";
return Decision.parse(JSON.parse(text));
}
If the model did ask for tools, the loop runs each one and sends the results back as the next message. When the loop uses up its model calls without reaching an answer, it throws:
// 01-baseline/agent.ts
...
const results: Anthropic.ToolResultBlockParam[] = [];
for (const block of response.content) {
if (block.type !== "tool_use") continue;
const output = await handlers[block.name](block.input);
results.push({ type: "tool_result", tool_use_id: block.id, content: JSON.stringify(output) });
}
messages.push({ role: "user", content: results });
}
throw new Error(`No decision after ${maxModelCalls} model calls`);
}
This baseline works. It got the late, delivered, high-value, and no-reference tickets right on every attempt, and the whole file comes to 39 lines. So writing the loop isn’t the hard part. The hard part is everything this file leaves to you once real customers start sending tickets:
inputSchema.parse, the exception ends the whole run, and an unknown tool name crashes on the handlers lookupmessages, so storing the history between requests, trimming it as it grows, and repairing it after a failed turn are all up to youEach framework below takes over part of that list, and each one takes over a different part. As you read through them, keep an eye on three things. The first is how much code it takes to declare the agent. The second is where the conversation history lives. The third is what you get back when the agent runs out of steps.
The AI SDK’s ToolLoopAgent accepts the shared tools almost unchanged, because its tool() helper takes a Zod inputSchema and an execute function directly:
// 02-ai-sdk/agent.ts
import { anthropic } from "@ai-sdk/anthropic";
import { isStepCount, Output, tool, ToolLoopAgent } from "ai";
import { Decision, INSTRUCTIONS, MAX_MODEL_CALLS, MODEL_ID } from "../shared/agent-config";
import { getEscalationPolicy, lookupOrder } from "../shared/tools";
// stopWhen is a constructor setting, so a different cap means a different agent instance.
export function createTriageAgent(maxModelCalls = MAX_MODEL_CALLS) {
return new ToolLoopAgent({
model: anthropic(MODEL_ID),
instructions: INSTRUCTIONS,
tools: {
lookup_order: tool({
description: lookupOrder.description,
inputSchema: lookupOrder.inputSchema,
execute: lookupOrder.execute,
}),
get_escalation_policy: tool({
description: getEscalationPolicy.description,
inputSchema: getEscalationPolicy.inputSchema,
execute: getEscalationPolicy.execute,
}),
},
output: Output.object({ schema: Decision }),
stopWhen: isStepCount(maxModelCalls),
});
}
export const triageAgent = createTriageAgent();
Two settings at the bottom do most of the work. Output.object sends the Decision schema to Anthropic’s native structured output, so you get a typed object back instead of text you have to parse. stopWhen: isStepCount(...) caps the loop, where each step is one model call, and the default cap is 20. You can only set that cap when you create the agent, so if you need a different cap for a particular request, you have to build a new agent. That’s why the file exports a factory function.
The agent doesn’t remember anything between calls, so for the follow-up ticket you keep the history yourself. You add the customer’s new message, pass the full list to generate(), and then append whatever the agent produced:
// 02-ai-sdk/runner.ts
...
messages.push({ role: "user", content: turn });
const result = await agent.generate({ messages });
messages.push(...result.response.messages);
The step limit behaves in a way worth knowing before you ship. When the agent runs out of steps, generate() doesn’t throw. It resolves normally with finishReason: "tool-calls", and the error only appears when you read result.output, which throws AI_NoOutputGeneratedError. That happened on all three limited runs, so you’ll catch the failure as long as your code reads the output before acting on it.
Streaming into a React app is where the AI SDK makes your life easiest. The API route takes six lines in total:
// app/api/ai-sdk/route.ts
...
export async function POST(request: Request) {
const { messages } = await request.json();
return createAgentUIStreamResponse({ agent: triageAgent, uiMessages: messages });
}
On the client, useChat receives typed tool parts such as tool-lookup_order, so your components know the exact shape of each tool’s input and output. Mastra is the only other framework here that gives you that, because it’s built on the AI SDK.
Mastra builds on the AI SDK and adds a central registry plus its own memory layer. You create tools with createTool, which takes the same Zod schema along with an id:
// 03-mastra/agent.ts
...
const lookupOrderTool = createTool({
id: lookupOrder.name,
description: lookupOrder.description,
inputSchema: lookupOrder.inputSchema,
execute: (input) => lookupOrder.execute(input),
});
The policy tool is built the same way. The agent names its model with a plain string, anthropic/claude-haiku-4-5, and Mastra’s model router resolves it using your ANTHROPIC_API_KEY. Memory is configured on the agent, but it only works once you register storage on a Mastra instance:
// 03-mastra/agent.ts
...
export const triageAgent = new Agent({
id: "triage-agent",
name: "Triage Agent",
instructions: INSTRUCTIONS,
model: `anthropic/${MODEL_ID}`,
tools: { lookup_order: lookupOrderTool, get_escalation_policy: escalationPolicyTool },
memory: new Memory({ options: { lastMessages: 20 } }),
defaultOptions: { maxSteps: MAX_MODEL_CALLS },
});
// Memory needs storage, and storage is registered on the Mastra instance.
export const mastra = new Mastra({
agents: { triageAgent },
storage: new LibSQLStore({ id: "triage-storage", url: ":memory:" }),
});
That extra registration makes Mastra’s agent file 32 lines long, against 25 for the AI SDK. In return, you no longer carry the conversation around yourself. You give each conversation a thread ID, pass it on every call, and Mastra loads and saves the history for you:
// 03-mastra/runner.ts
...
const memory = { thread: randomUUID(), resource: "customer-1" };
...
const response = await agent.generate(turn, {
maxSteps: maxModelCalls,
memory,
structuredOutput: { schema: Decision },
});
You ask for structured output on each call rather than on the agent, and the decision arrives on response.object. Unlike the AI SDK, Mastra also lets you change maxSteps per call, and if you leave it unset, the default is only five steps.
The step limit is where Mastra needs the most care from you. When the agent ran out of steps, generate() returned finishReason: "tool-calls" with response.object set to undefined, and nothing was thrown. That happened on all three limited runs, which makes Mastra the only framework in this comparison that fails silently. If your code only checks response.object, it can’t tell a run that hit the limit from one where the model simply returned nothing, so check finishReason before you trust the result.
To stream into React, you use handleChatStream from @mastra/ai-sdk. It defaults to the AI SDK v5 message format, so you need to pass version: "v7" to match the current SDK, and the decision schema goes in through defaultOptions. With that in place, your client gets the same typed tool parts as the AI SDK.
LangGraph’s older prebuilt createReactAgent is deprecated and now points you to createAgent in the langchain package. Under the hood, createAgent builds a LangGraph graph with a node that calls the model, a node that runs tools, and the edges that move between them. Its tool() helper takes your function first and the metadata second:
// 04-langgraph/agent.ts
...
const lookupOrderTool = tool(lookupOrder.execute, {
name: lookupOrder.name,
description: lookupOrder.description,
schema: lookupOrder.inputSchema,
});
You pass the decision schema as responseFormat, and you add a checkpointer, which is the component that saves the conversation between calls:
// 04-langgraph/agent.ts
...
export const triageAgent = createAgent({
model: new ChatAnthropic({ model: MODEL_ID }),
tools: [lookupOrderTool, escalationPolicyTool],
systemPrompt: INSTRUCTIONS,
responseFormat: Decision,
checkpointer: new MemorySaver(),
});
At 22 lines, this is the shortest agent file of the five. The checkpointer keys each conversation by a thread_id, so a follow-up message only needs the same ID, and the step limit is set per call:
const result = await triageAgent.invoke(
{ messages: [{ role: "user", content: turn }] },
{ configurable: { thread_id: threadId }, recursionLimit: 11 },
);
recursionLimit counts graph steps, not model calls. Calling the model is one step and running the tools is another, so six model calls with tools in between take 11 steps, and limiting the agent to two model calls means setting the limit to 3. The default is 25, and when the agent reaches it, LangGraph throws GraphRecursionError.
Structured output is where LangGraph behaved differently from the rest. With Claude Haiku 4.5, responseFormat doesn’t use the API’s native structured output. LangChain adds a hidden tool named extract-N and forces the model to call a tool on every request, so the final decision comes back as a call to that tool. Each turn therefore leaves three extra messages in your saved history. If you later switch to Claude Sonnet 5.5, keep in mind that this model returns a 400 error for any request that forces a tool call.
The same mechanism broke the ticket that mentions two orders on all three attempts. The model returned one decision per order as two extract calls in a single response. LangChain’s default retry answered only the first of those calls before asking the model again, so Anthropic’s API rejected the request because the second call never got a result:
400 messages.6: `tool_use` ids were found without `tool_result` blocks immediately after
If your tickets can produce more than one answer, test that case before you ship. The error comes from inside LangChain’s structured output handling, not from your code.
To stream into React, you use toBaseMessages and toUIMessageStream from @ai-sdk/langchain. Tool calls reach useChat as generic dynamic-tool parts with their results as JSON strings. The decision itself appears only as the extract tool call and never as text, so your UI has to read it from that tool part.
The OpenAI Agents SDK isn’t limited to OpenAI models. Through the aisdk() adapter in @openai/agents-extensions, it can run any model the AI SDK supports, including Claude. Its tools take the Zod schema as parameters:
// 05-openai-agents/agent.ts
...
const lookupOrderTool = tool({
name: lookupOrder.name,
description: lookupOrder.description,
parameters: lookupOrder.inputSchema,
execute: lookupOrder.execute,
});
The SDK requires strict mode when your parameters are Zod schemas, so every tool goes to the API with strict: true, which guarantees the model’s arguments match the schema exactly. None of the other frameworks turned that on by default. The agent wraps the Claude model in the adapter and takes the decision schema as outputType:
// 05-openai-agents/agent.ts
...
export const triageAgent = new Agent({
name: "Triage Agent",
instructions: INSTRUCTIONS,
model: aisdk(anthropic(MODEL_ID)),
tools: [lookupOrderTool, escalationPolicyTool],
outputType: Decision,
});
The SDK’s documentation still labels the adapter as beta and recommends its default provider if you’re using OpenAI models. With Claude, the adapter caused no failures in any test, and the decision came back through native structured output.
You pass conversation state and the step limit to run:
// 05-openai-agents/runner.ts
...
const result = await run(triageAgent, turn, { session, maxTurns: maxModelCalls });
The session here is a MemorySession, which stores the history for you, so the follow-up ticket needs no extra code. Each turn is one model call, and the default limit is 10. When the agent ran out of turns, it threw MaxTurnsExceededError on all three limited runs, which makes this the most obvious failure of the five. If you’d rather return a fallback decision than throw, you can supply one through errorHandlers.maxTurns.
One default will catch you out if you’re not using OpenAI. Tracing is on by default in Node.js and sends traces to OpenAI, so every run printed No API key provided for OpenAI tracing exporter. To turn it off, set OPENAI_AGENTS_DISABLE_TRACING=1.
For streaming, createAiSdkUiMessageStreamResponse converts a streamed run into the format useChat expects. You still have to convert the incoming chat messages into the SDK’s own input format yourself, using its user() and assistant() helpers and a small textOf function that pulls the text out of each message:
// app/api/openai-agents/route.ts
...
const input = messages.map((m) => (m.role === "user" ? user(textOf(m)) : assistant(textOf(m))));
const stream = await run(triageAgent, input, { stream: true });
return createAiSdkUiMessageStreamResponse(stream);
Tool calls arrive in your UI as dynamic-tool parts with their results as objects, and the decision streams as JSON text.
The Claude Agent SDK works differently from the other four. Instead of calling the model API directly, it runs the Claude Code binary as its engine, and that binary’s platform package takes 234 MB on disk. Your custom tools reach that engine through an MCP server that runs inside your process:
// 06-claude-agent-sdk/agent.ts
...
const supportServer = createSdkMcpServer({
name: "support",
version: "1.0.0",
alwaysLoad: true,
tools: [
tool(lookupOrder.name, lookupOrder.description, lookupOrder.inputSchema.shape, async (args) => ({
content: [{ type: "text", text: JSON.stringify(await lookupOrder.execute(args)) }],
})),
tool(getEscalationPolicy.name, getEscalationPolicy.description, getEscalationPolicy.inputSchema.shape, async (args) => ({
content: [{ type: "text", text: JSON.stringify(await getEscalationPolicy.execute(args)) }],
})),
],
});
Every tool result has to be wrapped in an MCP content array, which is why each handler turns its output into a text block. alwaysLoad: true switches off tool search, which is on by default and costs an extra round trip each time the model looks up a tool. Claude Code is built as a coding agent, so most of the remaining configuration turns off defaults your triage agent doesn’t need:
// 06-claude-agent-sdk/agent.ts
...
export const triageOptions: Options = {
model: MODEL_ID,
systemPrompt: INSTRUCTIONS,
tools: [], // drop every built-in tool (Bash, Read, Edit, WebFetch and the rest)
settingSources: [], // ignore CLAUDE.md and settings on the host machine
mcpServers: { support: supportServer },
allowedTools: ["mcp__support__*"],
permissionMode: "dontAsk",
thinking: { type: "disabled" }, // the other four run without extended thinking
outputFormat: { type: "json_schema", schema: z.toJSONSchema(Decision, { target: "draft-7" }) },
maxTurns: MAX_MODEL_CALLS,
};
Without the thinking setting, the SDK requested a 31,999-token thinking budget on every call to Haiku 4.5, which you’d pay for on every ticket. It also made a separate model call on every run just to generate a session title. You have to authenticate with an API key, because Anthropic doesn’t allow third-party products to offer claude.ai login.
The decision comes back as a call to a built-in StructuredOutput tool, and that call counts against maxTurns. Because it’s a tool, the model can also choose not to call it. On one of the three no-reference attempts, the model asked for the order number in plain text instead, so the run reported success with no decision attached. Before you trust a successful run, check that structured_output is actually there.
When the agent runs out of turns, the SDK first gives you a result with subtype: "error_max_turns" and then throws. For follow-up tickets, you pass the previous run’s session_id as resume, and the SDK reloads the conversation from the transcript it saved on disk.
There’s no ready-made adapter between this SDK and useChat, so the streaming route in the repository converts the SDK’s messages into chat UI chunks by hand. That route is 42 lines long, compared with six for the AI SDK.
Now that you’ve seen all five, here is how they compare on the same agent:
| AI SDK | Mastra | LangGraph.js | OpenAI Agents SDK | Claude Agent SDK | |
|---|---|---|---|---|---|
| Agent file (lines) | 25 | 32 | 22 | 24 | 32 |
| Tool schema | Zod inputSchema |
Zod inputSchema plus id |
Zod schema, function first |
Zod parameters, sent strict |
Zod shape, MCP content result |
| Step limit | isStepCount, set on the agent |
maxSteps, set per call |
recursionLimit, set per call |
maxTurns, set per call |
maxTurns, set per call |
| What the limit counts | model calls | model calls | graph steps | model calls | tool-use turns |
| Default limit | 20 | 5 | 25 | 10 | none |
| Follow-up state | you keep the messages | Memory plus storage |
checkpointer and thread_id |
MemorySession |
session resume from disk |
| Structured output | native | native | forced extract tool |
native | StructuredOutput tool |
| Streaming route | 6 lines, built in | 15 lines, @mastra/ai-sdk |
12 lines, @ai-sdk/langchain |
13 lines, agents-extensions |
42 lines, by hand |
Tool parts in useChat |
typed | typed | dynamic | dynamic | dynamic |
Every framework declares the agent in 22 to 32 lines, so the amount of setup code shouldn’t decide your choice. The rows further down matter more, as three frameworks use the API’s native structured output, while LangGraph and the Claude Agent SDK route the decision through a tool the model has to call. The step limit counts something different in three of the frameworks, so the same number gives your agent a different budget depending on which one you use. And if your frontend is React, only the AI SDK and Mastra give your components typed tool parts without extra work.
A happy path tells you little about a framework, so the repository also sends tickets designed to go wrong. Each one ran three times against every framework:
| Ticket | AI SDK | Mastra | LangGraph.js | OpenAI Agents SDK | Claude Agent SDK |
|---|---|---|---|---|---|
| No order reference | asked, 3/3 | asked, 3/3 | asked, 3/3 | asked, 3/3 | asked 2/3, once no decision |
| Reference not found | asked, 3/3 | asked, 3/3 | asked, 3/3 | asked, 3/3 | asked, 3/3 |
| Two orders | asked once, escalated ORD-1002 twice | escalated ORD-1002, 3/3 | 400 error, 3/3 | escalated ORD-1002, 3/3 | escalated twice, asked once |
| Limit of two model calls | returns, then output throws |
returns empty, no error | throws GraphRecursionError |
throws MaxTurnsExceededError |
error_max_turns, then throws |
The unknown reference worked everywhere, because the lookup’s found: false leaves the model only one sensible move. The missing reference failed only once, when the Claude Agent SDK’s model answered in text instead of calling StructuredOutput.
The two-order ticket shows what happens when your schema is smaller than the question. Every framework that finished looked up both orders, but Decision only has room for one, so the agents disagreed about which order to report. LangGraph never got that far, for the reason covered in its section.
The step-limit row is the one to remember for production. Three frameworks throw as soon as the limit is reached, and the AI SDK throws when you read the output. Mastra hands back a run that looks finished but has no decision in it, so with Mastra you have to check finishReason yourself.
Python is a better choice either when the rest of your stack is already Python, or when you need a capability that has no TypeScript equivalent. These four frameworks are where you’d start:
You built the same triage agent six times, and the setup code barely changed between frameworks. What changed was how each one behaves when things go wrong, and that’s the part you’ll deal with in production:
finishReason, because a run that hits the step limit fails silently
Tailwind CSS component libraries provide pre-built components to streamline the process of developing aesthetic, user-friendly interfaces.

Jev is a TypeSafe AI decision model that classifies messages using typed questions. Learn how to use it for routing to reduce unnecessary LLM calls.

Vibe coding makes building apps faster, but speed without engineering can create serious failures. Here are five real-world examples and practical checks to help you ship AI-generated code safely.

Compare React ViewTransition and Motion across four animation patterns. The native build saved ~38.6kB of gzipped JS, but required 179 more lines of CSS.
Hey there, want to help make our blog better?
Join LogRocket’s Content Advisory Board. You’ll help inform the type of content we create and get access to exclusive meetups, social accreditation, and swag.
Sign up now