How to deploy AI agents in production
Deployment is the same agent you already ran locally, placed on a boundary you chose. The policy engine does not get weaker on the way to production.
How to deploy AI agents in production
Deploying is not a second rewrite. How to deploy AI agents in production, on Kodeus, means taking the spec you already ran locally and choosing where that same runtime executes. The policy engine does not get swapped for a looser one on the way out of your laptop. If the local run refused a call, the deployed run refuses it too.
You have two ways to describe the agent. Create it in the Console, or author it as kodeus.yaml with the SDK. The file holds the model, the tools, the skills, the memory, the policy and the limits. The SDK validates it, scaffolds skills and prompts, and launches the runtime locally. You own the database the artifacts land in. The SDK is in private preview, so the first step for most teams is early access rather than a public install.
Do this before you pick a region. Write down the outcome, the tools that outcome needs, the identity model, and which actions must wait for a person. That is the scope conversation. An agent that can call anything, as anyone, is not ready for a VPC or for Kodeus Cloud.
Four places the same agent can run
The runtime is one. What changes is who operates it and what is allowed to leave the network. Enterprise deployment uses the same four.
| Local | Your cloud | Kodeus Cloud | Airgapped | |
|---|---|---|---|---|
| Operated by | You | You | Kodeus | You |
| Database | Yours | Yours, in your VPC | Ours, per tenant | Yours |
| Network egress | Model provider | Model provider only | Model provider | None |
| Models | Any provider | Any provider, your keys | Any provider | Hosted inside the perimeter |
| Best for | Development | Production under your network policy | Fastest start | Sites where nothing may leave |
The sequence we actually use
Run it locally against the same policies, with no account required. Read a trace. If the agent called the wrong tool, fix the spec or the skill before you talk about infrastructure. A dry run is available when you want the plan without the effect: the agent proposes the calls and executes nothing.
Put identity on the request. Anonymous shared service accounts make the trace useless the first time someone asks who did this. Credentials go in the per-user vault, AES-256-GCM, with a revocation path. Tools attach over MCP, and only the ones this outcome needs. The notes on the best MCP servers are the filter.
Turn policy on where it matters. Rails can start in monitor mode so you see what the agent tries. Then gate the actions with real consequences. Gating everything trains people to click approve. Human approval parks the turn until someone decides, and the decision is stored with the run.
Choose the boundary from the table. Your VPC when the data already lives there and egress should be the model provider only. Kodeus Cloud when you want the fastest start and a database per tenant is acceptable. Airgapped when the cable is pulled and the model is hosted inside. Then set spend caps per agent and per user. Cost should be visible before the invoice.
That is how to deploy AI agents in production without discovering tenancy in week six. The infrastructure list, if you want it as a checklist rather than a sequence, is on AI agent infrastructure. The control catalogue and the evidence a reviewer can ask for are on the enterprise AI agent platform page.
What done looks like
A real user can open a session. The run carries their identity. A tool call uses their credential. A guarded action either proceeds or waits. The trace shows the call, the result, any refusal, the timing and the cost. You can point at that record the next day without reading a chat log.
If you cannot do that yet, you do not have a deployment problem. You have a missing control. Add the control in the spec and run it locally again. Production is the same agent afterwards, in a place you chose on purpose.
Walk through a deployment
Thirty minutes on one workflow: spec, identity, tools, policy, and where it should run.
Frequently asked questions
What does deploy mean here?
The same agent, described in the Console or as a kodeus.yaml spec, running somewhere other than your laptop. Local, your VPC, Kodeus Cloud, and airgapped all use the same policy engine.
Do I need an account to run locally?
No. A local run uses the same runtime and the same policies with no account. An account matters when you want Kodeus Cloud or a managed path.
What has to be true before the first real user?
An identity on every request, credentials scoped to that user, the tools you actually intend, a policy that can refuse, and a trace you can read. Spend limits belong in that list if the agent can incur cost.
Can the model stay inside our network?
Yes. In your VPC, egress can be limited to the model provider. Airgapped means no outbound dependency, with models hosted inside the perimeter.
What changes between local and production?
Where it runs, who operates it, and what is allowed to leave the network. The spec, the policy, and the trace shape stay the same.
How do I rehearse a turn?
A dry run has the agent plan the turn and execute nothing. You inspect the intended tool calls before anything touches a real system.
Who approves a risky action?
A person you name. The action waits, the decision is recorded, and the trace shows who approved it and what changed.
Where do I read the rest?
Developers covers the spec and the SDK. Enterprise covers deployment models. AI agent infrastructure covers what arrives with the runtime.