Skip to content
Private preview The Kodeus SDK and demo app are not public yet. Get early access
Kodeus
Observability

AI agent observability

Down to the tool call. When an agent gets something wrong, the answer is rarely in the answer. AI agent observability means seeing the calls it made, the results it got and the checks it ran.

A chat log is not a trace

Reading the final message tells you what the agent said, not what it did. The useful record is the sequence underneath: which tool was called, with what arguments, what came back, what a guardrail refused, and where the time went.

Every call and result

Tool calls and their results recorded as structured events rather than free text, so they can be queried instead of read.

Refusals are visible

When a guardrail blocks an action, the block is part of the record. Silence is the worst possible audit trail.

Attributable to a person

Runs carry the identity they acted for, so a trace answers who as well as what.

Rehearse before you run

A dry run plans the turn and executes nothing, which is the cheapest way to see what an agent intended to do.

What agent observability is for

Three audiences want the same data for different reasons, which is why it has to be structured rather than printed.

Debugging

Find the call that returned the wrong thing, instead of guessing from the summary the model wrote afterwards.

Audit

Show a reviewer exactly what the agent did on a given day, for a given person, with the approvals attached.

Improvement

Recurring failures and repeated successes are only findable once runs are recorded in a shape you can compare.

What a traced turn contains

Each turn produces an ordered stream rather than a paragraph. That is what makes AI agent observability queryable instead of anecdotal.

Recorded Why it matters
The tool called, and its argumentsThe single most common cause of a wrong answer is a right tool called wrongly.
The result that came backDistinguishes a bad decision from bad data, which need different fixes.
Any guardrail refusalShows what the agent tried to do, not only what it managed to do.
Validation checksRecords whether the answer was grounded in what the tools actually returned.
The identity it acted forTurns a trace into an audit record instead of a debug log.
Timing per stepLatency is nearly always one slow call, not a uniformly slow agent.

The same record is what makes improvement possible later. Patterns across many runs are only visible once the runs are comparable, which is the argument for structure over prose. See the developer surface for how the stream is consumed.

How to use AI agent observability on a real failure

Start with the call, not the paragraph the model wrote afterwards. AI agent observability is useful only if the first screen answers four questions: which tool ran, with which arguments, what came back, and whether a guardrail stopped it. If any of those is missing, you are reading a story about the run, not the run.

A wrong answer usually falls into one of three piles. The tool was right and the arguments were wrong. The arguments were right and the data that came back was wrong. Or the call should never have left, and the policy did not refuse it. The first two are product bugs. The third is AI agent governance, and the refusal, when it works, is an event in the same stream. You should not need a second system to see a block.

Timing belongs in that stream. A slow turn is almost always one slow call. Cost belongs there too, so a spend cap is a number you can tie to a step instead of a surprise on an invoice. Identity belongs there because a trace that cannot name the person is a debug log, not an audit record. Those are the fields a reviewer asks for, and they are the same fields you want on an ordinary Tuesday when something looks off.

Keep the record where the run lives. The runtime is self-hostable against your database, in your VPC or airgapped, so AI agent observability does not mean shipping prompts to someone else's dashboard. A dry run is the rehearsal: the agent plans the turn and executes nothing, which is how you inspect intent before a tool touches a real system. When you are ready to put that run somewhere a user can reach, follow how to deploy AI agents in production. The infrastructure those traces sit on is AI agent infrastructure.

Look inside a real run

We will walk you through a traced turn, including the calls a guardrail refused.

Frequently asked questions

What is AI agent observability?

It is the ability to see what an agent actually did: the tools it called, the arguments it used, the results it received, the guardrails that refused it and the time each step took. It is the difference between reading an agent's summary and checking its work.

How is it different from LLM logging?

Prompt and completion logging tells you what the model was asked and what it said. Agent observability covers the actions in between, which is where the consequences live.

Can I see a run without executing it?

Yes. A dry run has the agent plan the turn and execute nothing, so you can inspect the intended tool calls before anything touches a real system.

What should I open first when an answer is wrong?

The tool call and the result, then any refusal. A wrong summary is usually a wrong argument, bad data, or a check that should have blocked the turn. The chat text is the last place to look.

Who is a trace for?

The person debugging today's failure, the reviewer who asks what happened last Tuesday, and the team improving the agent from patterns across runs. One record, three uses, which is why it has to be structured.

How does observability relate to governance?

A guardrail refusal is an event in the same stream as a successful call. AI agent governance decides. AI agent observability shows the decision. Silence after a block is the failure mode both exist to avoid.

Can I keep the traces in my own environment?

Yes. The runtime is self-hostable and works against a database you own, in your VPC or airgapped. The trace is part of that run, not a feed you must send to Kodeus.

Does this replace metrics on the model?

No. Latency and cost per step sit beside the call record. You still want to know which call was slow. You do not need a separate logging project to get there.