Skip to content
Private preview The Kodeus SDK and demo app are not public yet. Get early access
Kodeus
Guide

How to deploy AI agents in production

Deployment is the same agent you already ran locally, placed on a boundary you chose. The policy engine does not get weaker on the way to production.

How to deploy AI agents in production

Deploying is not a second rewrite. How to deploy AI agents in production, on Kodeus, means taking the spec you already ran locally and choosing where that same runtime executes. The policy engine does not get swapped for a looser one on the way out of your laptop. If the local run refused a call, the deployed run refuses it too.

You have two ways to describe the agent. Create it in the Console, or author it as kodeus.yaml with the SDK. The file holds the model, the tools, the skills, the memory, the policy and the limits. The SDK validates it, scaffolds skills and prompts, and launches the runtime locally. You own the database the artifacts land in. The SDK is in private preview, so the first step for most teams is early access rather than a public install.

Do this before you pick a region. Write down the outcome, the tools that outcome needs, the identity model, and which actions must wait for a person. That is the scope conversation. An agent that can call anything, as anyone, is not ready for a VPC or for Kodeus Cloud.

Four places the same agent can run

The runtime is one. What changes is who operates it and what is allowed to leave the network. Enterprise deployment uses the same four.

LocalYour cloudKodeus CloudAirgapped
Operated byYouYouKodeusYou
DatabaseYoursYours, in your VPCOurs, per tenantYours
Network egressModel providerModel provider onlyModel providerNone
ModelsAny providerAny provider, your keysAny providerHosted inside the perimeter
Best forDevelopmentProduction under your network policyFastest startSites where nothing may leave

The sequence we actually use

Run it locally against the same policies, with no account required. Read a trace. If the agent called the wrong tool, fix the spec or the skill before you talk about infrastructure. A dry run is available when you want the plan without the effect: the agent proposes the calls and executes nothing.

Put identity on the request. Anonymous shared service accounts make the trace useless the first time someone asks who did this. Credentials go in the per-user vault, AES-256-GCM, with a revocation path. Tools attach over MCP, and only the ones this outcome needs. The notes on the best MCP servers are the filter.

Turn policy on where it matters. Rails can start in monitor mode so you see what the agent tries. Then gate the actions with real consequences. Gating everything trains people to click approve. Human approval parks the turn until someone decides, and the decision is stored with the run.

Choose the boundary from the table. Your VPC when the data already lives there and egress should be the model provider only. Kodeus Cloud when you want the fastest start and a database per tenant is acceptable. Airgapped when the cable is pulled and the model is hosted inside. Then set spend caps per agent and per user. Cost should be visible before the invoice.

That is how to deploy AI agents in production without discovering tenancy in week six. The infrastructure list, if you want it as a checklist rather than a sequence, is on AI agent infrastructure. The control catalogue and the evidence a reviewer can ask for are on the enterprise AI agent platform page.

What done looks like

A real user can open a session. The run carries their identity. A tool call uses their credential. A guarded action either proceeds or waits. The trace shows the call, the result, any refusal, the timing and the cost. You can point at that record the next day without reading a chat log.

If you cannot do that yet, you do not have a deployment problem. You have a missing control. Add the control in the spec and run it locally again. Production is the same agent afterwards, in a place you chose on purpose.

Walk through a deployment

Thirty minutes on one workflow: spec, identity, tools, policy, and where it should run.

Frequently asked questions

What does deploy mean here?

The same agent, described in the Console or as a kodeus.yaml spec, running somewhere other than your laptop. Local, your VPC, Kodeus Cloud, and airgapped all use the same policy engine.

Do I need an account to run locally?

No. A local run uses the same runtime and the same policies with no account. An account matters when you want Kodeus Cloud or a managed path.

What has to be true before the first real user?

An identity on every request, credentials scoped to that user, the tools you actually intend, a policy that can refuse, and a trace you can read. Spend limits belong in that list if the agent can incur cost.

Can the model stay inside our network?

Yes. In your VPC, egress can be limited to the model provider. Airgapped means no outbound dependency, with models hosted inside the perimeter.

What changes between local and production?

Where it runs, who operates it, and what is allowed to leave the network. The spec, the policy, and the trace shape stay the same.

How do I rehearse a turn?

A dry run has the agent plan the turn and execute nothing. You inspect the intended tool calls before anything touches a real system.

Who approves a risky action?

A person you name. The action waits, the decision is recorded, and the trace shows who approved it and what changed.

Where do I read the rest?

Developers covers the spec and the SDK. Enterprise covers deployment models. AI agent infrastructure covers what arrives with the runtime.