You closed the tab but the agent is still on.
That’s the promise of always-on AI-agents. Last month, they got mainstream. OpenAI launched Dots on September 29th during DevDay. Meta launched Muse three weeks early Already, self-hosted boxes are continuously running open-source autonomous AI agents such as OpenClaw and Hermes Agent.
The pitch is straightforward. Always on agents reduce prompting and increase the work while you sleep. For enterprise AI agents, this is not straightforward at all in terms of engineering.
What are always-on AI agents? Persistent, not perpetual
An always-on AI agent doesn’t need a model that runs 24/7 It requires permanent AI agent memory between runs.
A June 2026 arXiv survey of 435 works define these agents by durable state, not simply memories that can be retrieved. Task logbooks, authorizations, credentials, commitments, audit logs, can trigger conditions. An agent that wakes once a day still counts, as long as it carries that condition forward and acts upon it.
One work is in three parts as follows:
- A wake-up process. What triggers the agent: timer, schedule or event?
- The State. What it recalls between starts.
- A leash. What it can accomplish alone. What it needs permission to do. What it prevents.
It is a significant change from the assistants we have now. Chatbots, copilots and coding agents are in a session and take on the persona of a single individual. Ambient agents (or background agents) run as bounded workers in real processes.
Always-on AI agents shipping now: Dots, Muse, OpenClaw and more
There are two camps with two distinct offerings. One, hosted agents you rent. The other, self-hosted agents you run.
| Agent | Runs on | Model | Starts at |
| OpenAI Dots | OpenAI cloud | GPT-6 Astra | $100/mo |
| Grok Bot | Cursor cloud | Undisclosed | $20/mo |
| Meta Muse | Meta cloud | Muse Spark | Free |
| Hermes Agent | Self-hosted | Your choice | Free + server and model |
| OpenClaw | Self-hosted | Your choice | Free + server and model |
Prices as listed by AIMultiple on September 30, 2026.
Dots sets the template for hosted camping. Each dot has its own cloud computer and browser, connects to 4,000+ apps, and lives in Slack, Teams and ChatGPT. It runs “proactive research” between jobs via read-only connections. One early tester’s Dot noticed a missed invoice, which was produced and sent out following approval. OpenAI also showed off specialty dots with their own identities and credentials for duties like procurement and invoicing processing.
The self-hosted camp sacrifices some convenience for control. NVIDIA’s NemoClaw stack runs OpenClaw inside an OpenShell sandbox on a DGX Spark with Nemotron 3 Super 120B delivered locally. No data ever leaves the device. People have to authorize any requests to go outside.
AI agent cost: what wakes the agent sets the bill
This is where the problem gets more troubling for organizations, because listening is cheap, thinking is not.
An always-on AI agent has three ways to wake up. A heartbeat timer tells the model to look around. A schedule does a specific job. It only starts when anything occurs – a fresh email, a webhook. That difference shows on the bill.
Heym costed one inbox agent on gpt-6.1-sol. Same model, same 2,800-token prompt:
- Five-minute heartbeat, cold cache: $122.69/month
- Five-minute heartbeat, warm cache: $30.76/month
- Event trigger plus a cheap decision gate: $2.14/month
Difference of 57 times. The only policy that was changed was the wake-up policy.
Defaults make it worse. OpenClaw’s heartbeat performs a full agent turn every 30 minutes and can resend around 100,000 tokens of history each time. AIMultiples values that at roughly $19 per day for a premium model. Before the agent can accomplish anything beneficial.
As the fix, what enterprises require is architecture. Imagine a funnel, with four phases, each more expensive than the one before:
- Plain code is listening. A trigger such as a mailbox poll, a webhook, or a queue waits for something to occur. No model is involved hence the waiting cost is 0 tokens.
- A plain filter drops the obvious stuff. Simple filters filter out the obvious junk: newsletters, no-reply mail, a web page that hasn’t updated, a status that didn’t move. Check for accuracy and free.
- A small model scores what is left. A lightweight decision model poses a small question such as “does this email need a reply?” and gives a probability. Each check is about 300 tokens, meaning it costs about one one-thousandth of a cent.
- The huge model only runs on what passes. Events that cross the threshold get to the complete agent, which uses its tools and memory to accomplish the real work. “Anything that has side effects, such as sending a reply, then goes to a person for approval.
The aim is to get work out the door as soon as feasible. In Heym’s inbox, only one in four emails reaches the big model. And that’s how the monthly bill goes from $122.69 to $2.14.
AI agent governance and security: the missing layer
The field is good at remembering, not ruling. The arXiv survey revealed that research focuses on state storage and retrieval, much less about ruling it, or recovering from poor state, or letting it go. Its answer is a pilot assessment protocol, AOEP-v0, that assesses how agents modify and recover state, not just answer quality.
Every source keeps getting three controls:
- Identity of its own. Enterprise AI needs service accounts and least privilege, not a stolen user session. Every action is accountable and auditable.
- Side effects first, then approval. Dots always leaves duties like password updates to the user and sends dangerous actions through auto-review. NemoClaw prevents network calls before they are approved.
- Alarms and switches off. Execution-count and cost notifications catch a rogue trigger. When the job is done, a stop condition turns the agent off.
There is no optimal security model for AI agents. As NVIDIA puts it succinctly: None completely guards against sophisticated quick injection. Review relevant work.
How to deploy always-on AI agents in the enterprise
Always-on is an operations problem in AI clothing. Treat it as such. The model is the easy bit. The tricky part is everything surrounding it. What it can touch when the agent runs, how you clean up when it gets something wrong.
Five stages to begin:
- Choose a standing job, not a demo. Recurring work such as invoice processing, helpdesk triage or release checks justify always-on agents. Pick work with a clear owner who will examine the output and a clear definition of done so you can identify if the agent is helpful.
- Wake on events first. If the system can tell you something has changed (i.e. new email, webhook, ticket update), let that event wake the agent. Use schedules for work that is due at a specific time, like a weekly report. Nothing else can communicate change. Save heartbeats, where the agent wakes on a timer to glance about. A tick costs you a model call.
- Gate the model. Run the above-mentioned funnel. Plain code filters the obvious, a tiny model scores what is left and the giant model runs only on what passes. You still pay for meaningful work, not for idle time.
- Give it an identity of its own. An always-on agent shouldn’t be using someone’s login. Give it a service account with scoped credentials and only the access it needs to do its job. Log every action so it is accountable and auditable. Do not let it inherit admin rights
- Design for recovery. Persistent agents will build up state, some of which will be incorrect. Before go-live, understand how to audit what the agent did, roll back a negative change, and wipe what is learnt. According to the arXiv study, this is the least researched area.
Companies will rename their products. Dots, Muse and OpenClaw could seem very different a year from now. The three design decisions below will not: the wake-up policy which chooses when the agent thinks, the state model which decides what it remembers, and the leash which decides what it can do alone. Those are up to you to make.