Hermes Agent vs OpenClaw: Which One Should Run Your Store's Back Office?

Picture two stores running the same 7am low-stock report to Slack. Store A uses Hermes Agent. Store B uses OpenClaw. At the end of the month the reports look identical. The model bills can differ by hundreds of dollars, and the security exposure differs too, because of defaults most owners never open.
Both agents are free, open source and self-hosted. Both chat through Slack, run scheduled jobs, write their own reusable procedures and connect to store systems through MCP. So the choice comes down to details: how each one learns, what it runs without asking, what it costs and how each project has handled security problems.
This comparison covers those details for a store operator. It builds the same three back-office jobs in both agents, prices a month of running them, and ends with a "choose this if" list and a rollout plan that keeps the agent away from anything it can break.
If you have never used Hermes, start with our plain-English guide to what Hermes Agent is and how a store can use it. OpenClaw newcomers should read our guide to OpenClaw for ecommerce operations first.
What's in this guide
- Side by side: the 16 differences that matter to a store
- How each one learns, remembers and talks to you
- Setup effort and where each one runs
- Models and running costs, with a monthly example
- Security defaults and track record
- Connecting to Shopify, marketplaces and an order system
- The same three store jobs, built in each
- Choose Hermes if, choose OpenClaw if
- What to do this week, and the rest of the month
Side by side: the 16 differences that matter to a store
Checked against each project's documentation as of September 22, 2026. Both ship changes weekly, so re-check anything you rely on.
| What you care about | Hermes Agent | OpenClaw |
|---|---|---|
| Who makes it | Nous Research, a venture-backed AI lab | The OpenClaw Foundation, a donor-funded nonprofit |
| Licence and price | MIT, free | MIT, free |
| Hosted option from the maker | Yes, Hermes Cloud, billed per day | No; you host it (guides cover many cloud hosts) |
| Team permissions | Everyone you authorize on a chat app gets the same trust | Per-person roles and per-agent tool profiles |
| Chat apps | Slack, Teams, Telegram, WhatsApp, Signal, Google Chat, email and more | Slack, Teams, WhatsApp, Telegram, iMessage, Discord and 20+ more |
| Memory it loads every session | Two small capped files (2,200 and 1,375 characters) | USER.md and MEMORY.md, plus daily notes it can search |
| Writes its own skills | Yes, by default; optional approval gate | Yes, by default since late July 2026; "propose" mode for review |
| Commands it runs without asking | "Smart" mode: a second model auto-approves low-risk commands and escalates the rest | On the main host, full access with no prompts unless you pick a stricter preset |
| Scheduled jobs when a risky command appears | Blocked by default | Blocked if nobody answers the prompt, but prompts are off by default |
| Jobs that skip the AI model | Yes, script-only "no-agent" jobs | Yes, command and script jobs, plus condition watchers |
| Always-on check-in that costs tokens | None by default; background learning reviews run after conversations | Heartbeat, every 30 minutes by default |
| Skill marketplace | Skills Hub, which also pulls from ClawHub and other registries | ClawHub, scanned with VirusTotal since February 2026 |
| Public vulnerability record | 39 CVEs filed by a third party (Apr to Sep 2026); no project-published advisories | More than 700 advisories published by the project itself |
| MCP (store connections) | Yes; per-server tool allowlist | Yes; per-server tool allowlist |
| Models | Nous Portal, OpenRouter, OpenAI, Anthropic, local models and more | Anthropic, OpenAI, Google, OpenRouter, local models and more |
| Release pace | Very fast; weekly tags with thousands of commits | Fast; several release channels with fixed version numbers |
Two rows surprise people who read older comparisons, including our own: OpenClaw now writes its own skills by default, and it runs host commands without prompting by default.
How each one learns, remembers and talks to you
The gateway: one process, many chat apps
Both agents run a long-lived program called a gateway. It holds your chat connections and scheduled jobs, so the agent keeps working after you close your laptop.
Both reach Slack, Teams, Telegram and WhatsApp, so the chat app rarely decides it. The bigger difference is permissions. OpenClaw lets you give different people different role limits and run several agents with separate tool profiles on one gateway. Hermes treats everyone you authorize on a chat app as equally trusted, so a teammate who can message it can use every tool it has.
Memory: small on purpose
Hermes loads two files into every session: MEMORY.md for facts about your setup and USER.md for your preferences, capped at 2,200 and 1,375 characters. It can also search past conversations. OpenClaw uses the same two file names plus dated daily notes, which a background "dreaming" pass distils into long-term memory.
For a store, the difference is small. Both remember conventions, such as "a stuck order is one unshipped after 48 hours." Neither should be trusted with stock levels or order counts. Those must come from a live tool call every time, or the agent reports yesterday's numbers with today's confidence.
Skills: both write their own now
A skill is a plain-text instruction file: when to use it, the steps and how to check the result. This is where many comparisons are out of date.
Hermes has always written skills on its own. The Hermes skills documentation says the agent "writes skills freely" by default, including from a background review that runs after a conversation. Setting skills.write_approval: true stages every new or edited skill until you approve it. Memory has the same switch under memory.write_approval.
OpenClaw used to stage learned skills as proposals for a person to approve. That changed in late July 2026. The OpenClaw self-learning documentation now lists auto as the default mode. After a long conversation (at least 10 model steps), a background reviewer can edit the agent's skill collection directly, and a weekly review tidies it. The docs note that this direct route skips the proposal scanner and rollback snapshots, and tell you to keep backups. Setting the mode to propose brings back human review.
So the old line "Hermes learns by itself, OpenClaw asks first" no longer holds. For anything that touches money, turn on the review gate in whichever one you run.
Release pace
Hermes moves fast. Its September 21 release notes count 1,812 merged pull requests in the single week since the previous tag. New bugs can arrive as fast as features. OpenClaw's fixed-version release channels make a known-good version easier to hold. On either, pin a version and upgrade on purpose.
Setup effort and where each one runs
Installing is a one-liner on both: a script for macOS, Linux and Windows installs the runtime (Python for Hermes, Node.js for OpenClaw) and starts a setup wizard that asks for a model key and a chat app. Expect an hour to get a first reply in Slack, and most of a day to connect store data safely.
Where the agent lives is the bigger decision:
- Your own server. Both run on a small rented server (a VPS), which the Hermes README pitches at about $5 a month. You handle updates, backups and access. OpenClaw has no hosted service of its own, so this is its usual home.
- A desktop app. Both have one (Hermes added its own in June 2026). It is handy for trying things out, but a laptop sleeps and travels, and scheduled 7am jobs need a machine that is awake at 7am.
- Hermes Cloud. Nous Research hosts the agent for you. The Hermes Cloud pricing page lists a Medium instance (2 GB of memory, 10 concurrent sessions) at $0.56 a day while running and $0.03 a day while stopped, with model usage billed separately. Optional Nous Portal plans start at $20 a month and include $22 of model credits.
Switching later is cheaper than it looks. The Hermes README documents a hermes claw migrate command that imports OpenClaw settings, memories, skills and keys, with a dry-run preview. Because both agents reach store data through MCP, the store connection you build carries over either way.
The backers differ. TechCrunch reported in July 2026 that Nous Research was raising at least $75 million at a $1.5 billion valuation, with Hermes at roughly 214,000 GitHub stars. OpenClaw is run by a nonprofit foundation with no paid tier.
Models and running costs, with a monthly example
You pay for hosting and for the AI model. Both agents work with the big model providers, routers like OpenRouter and local models. For where the model makers are heading in commerce, see our overview of AI shopping agents, OpenClaw, Hermes and the new commerce protocols and our note on Meta's Muse agent and what it means for sellers.
Model prices are quoted per million tokens. A token is a chunk of text, roughly three-quarters of a word. Each model call re-sends the instructions, tool list and conversation so far, so input tokens pile up. Anthropic's API pricing page lists Claude Sonnet 4.6 at $3 per million input tokens and $15 per million output tokens, and Claude Haiku 4.5 at $1 and $5. Cached input, text the provider has seen recently, costs a tenth of the normal input price on both.
The setting that can cost more than everything else
OpenClaw runs a heartbeat: a full agent turn in your main conversation every 30 minutes by default (hourly on some Anthropic sign-in methods). The OpenClaw heartbeat documentation says that running it in an isolated session cuts each run from about 100,000 tokens of conversation history to 2,000 to 5,000.
In plain English: leave the heartbeat on a long-running main chat and it can become your biggest bill without doing any store work.
Heartbeats per month = 48 a day x 30 days = 1,440
Full history: 1,440 x 100,000 tokens = 144,000,000 tokens
144 x $3 (Sonnet 4.6) = $432 before caching
Isolated: 1,440 x 5,000 tokens = 7,200,000 tokens
7.2 x $3 (Sonnet 4.6) = $21.60
Caching would shrink the first figure, and 100,000 tokens is the docs' estimate for a long history, not a fixed charge. For a store agent, either turn the heartbeat off (heartbeat.every: "0m") or set isolatedSession: true and give it a cheaper model.
Hermes runs no heartbeat by default. Its equivalent background cost is the self-improvement review that can run after a conversation, which replays that conversation to decide what to save. You can point it at a cheaper model or cap its input tokens.
Worked example: one month for a small store
Example: a store running four agent jobs, assuming 20,000 input and 500 output tokens per model call. Swap in your provider's real usage figures after a week.
| Job | Runs a month | Model calls per run | Model calls |
|---|---|---|---|
| Daily low-stock check | 30 | 4 | 120 |
| Stuck-order triage, twice a day | 60 | 6 | 360 |
| Weekly supplier follow-up drafts | 4 | 10 | 40 |
| Ad hoc questions, 10 a day | 300 | 3 | 900 |
| Total | 1,420 |
Input = 1,420 calls x 20,000 = 28,400,000 tokens Output = 1,420 calls x 500 = 710,000 tokens Sonnet 4.6: 28.4 x $3 + 0.71 x $15 = $85.20 + $10.65 = $95.85 Haiku 4.5: 28.4 x $1 + 0.71 x $5 = $28.40 + $3.55 = $31.95 Move the low-stock check to a script job (no model): 1,300 calls Sonnet 4.6: 26.0 x $3 + 0.65 x $15 = $78.00 + $9.75 = $87.75 Haiku 4.5: 26.0 x $1 + 0.65 x $5 = $26.00 + $3.25 = $29.25
Now add hosting and the settings discussed above:
| Setup | Hosting | Model | Monthly total |
|---|---|---|---|
| Hermes on Hermes Cloud Medium, Haiku 4.5 | $16.80 | $29.25 | $46.05 |
| Hermes on a $5 VPS, Sonnet 4.6 | $5.00 | $87.75 | $92.75 |
| OpenClaw on a $5 VPS, Sonnet 4.6, heartbeat off | $5.00 | $87.75 | $92.75 |
| OpenClaw, Sonnet 4.6, isolated 30-minute heartbeat | $5.00 | $109.35 | $114.35 |
| OpenClaw, Sonnet 4.6, heartbeat left on full history | $5.00 | $519.75 | $524.75 |
Configured the same way, the two cost the same. The model moves the bill about three-fold; one default setting moves it more than five-fold. Prompt caching lowers every row, so treat these as upper bounds.
Security defaults and track record
Two things decide how much damage a mistake or an attacker can do: what the agent may run without asking, and what you install into it.
What each one runs without asking
The Hermes security guide describes a "smart" approval mode by default. A second model auto-approves low-risk commands, auto-denies clearly dangerous ones and asks you about the rest. Scheduled jobs default to cron_mode: deny, so an unattended job that hits a risky command is blocked. A hardline blocklist refuses commands like wiping a disk in every mode. One catch: when Hermes runs commands inside a container, it skips these checks and treats the container as the safety wall.
OpenClaw starts more open. The OpenClaw exec approvals documentation lists the default for commands on the gateway host as security: full with ask: off, meaning no prompts. Sandboxing, which runs commands in a sealed-off container, is off unless you configure it. OpenClaw offers a one-line cautious preset that switches to an allowlist and asks about anything else, and its openclaw security audit command checks a deployment for risky settings. Both support pairing codes, so a stranger's direct messages don't reach the agent until you approve them.
So Hermes starts cautious and lets you loosen it, while OpenClaw starts permissive and expects you to tighten it. Either can end up safe; the risk is an owner who never changes the defaults.
The skill supply chain
A marketplace skill is someone else's instructions running with your agent's access. Palo Alto Networks Unit 42's June 2026 analysis recounts Koi Security's February disclosure of 341 malicious ClawHub skills and the marketplace's move to VirusTotal scanning. It then describes five more malicious skills that stayed unblocked between February and May, including two macOS password stealers and one that slipped affiliate links into the agent's product advice.
This touches Hermes users too: the Hermes Skills Hub lists ClawHub as a community source alongside other registries. Hermes scans every hub install for data theft, prompt injection (hidden text that hijacks the agent's instructions) and destructive commands, and it won't install a skill rated "dangerous" even with --force. A scanner catches known patterns, not intent, so write your own short store skills and install nothing you haven't read.
The public vulnerability record
A security advisory is a notice a project publishes about a flaw, usually after fixing it; a CVE is a public ID number for a reported flaw. OpenClaw publishes its own advisories. By September 22, 2026, OpenClaw's security advisory list held 722 entries: 14 rated critical, 249 high, 390 medium and 69 low. The early ones were serious, such as a January 31 "1-click" flaw that could leak the gateway's login token and let an attacker run code.
Hermes has no advisories published on its GitHub repository as of the same date. The US National Vulnerability Database lists 39 CVEs for hermes-agent published between April 27 and September 3, 2026, most rated medium. They were filed through a third-party database, and 28 of the records state that the vendor was contacted and did not respond.
Neither count is a safety score. A long advisory list can mean a project finds and discloses its own bugs; a short one can mean reports arrive elsewhere. OpenClaw has a busy, documented disclosure process, while Hermes's public record is mostly outside researchers' reports. On either one, update regularly, keep the gateway off the open internet and start with read-only store credentials.
Connecting to Shopify, marketplaces and an order system
Both agents connect to business software through MCP (Model Context Protocol), a common standard for plugging an AI agent into other systems. The system publishes a list of tools, such as "list orders" or "get stock level," and the agent calls them. If MCP is new to you, read our explainer on what MCP is and why it matters for ecommerce.
The configuration is close to identical. Hermes uses mcp_servers in its config file and OpenClaw uses mcp.servers. Both support OAuth sign-in (the standard "grant access" flow) and a per-server tool allowlist. Our step-by-step OpenClaw MCP configuration guide covers the OpenClaw side, and the setup section of our Hermes guide covers the Hermes side.
Three store-specific points apply to both:
- Shopify's official MCP servers mostly serve shoppers and developers. They search a catalog, manage carts and answer API questions. Back-office reads like stock levels usually go through the Admin API with a custom app token. Shopify's access scope reference notes that
read_ordersonly covers the last 60 days of orders, and older history needsread_all_orders, which Shopify must approve. Plan your supplier and returns reports around that window. - Marketplaces each need their own credentials. Amazon, Walmart, eBay and TikTok Shop run separate seller APIs. Four connections mean four sets of keys for the agent to hold and four stock numbers that can disagree.
- One order and inventory system is the cleaner connection. If your channels already feed one system, the agent reads one stock figure per product and one order list tagged by channel. An order management system (OMS) such as Nventory exposes an MCP server for orders and inventory that either agent can point at, with the channel sync handled underneath.
Whichever agent you pick, use a read-only token on the server side and an allowlist on the agent side.
The same three store jobs, built in each
Endpoints, file paths and channel IDs below are placeholders.
1. Daily low-stock check (no AI needed)
A low-stock list is arithmetic, so both agents should run it as a plain script and spend zero tokens. Set each product's reorder point with the reorder point calculator rather than a round number.
# low-stock.sh (prints nothing when nothing is low) curl -s -H "Authorization: Bearer $STORE_READ_TOKEN" \ "https://your-oms.example.com/api/inventory?below_reorder_point=true" \ | jq -r '.items[] | [.sku, .available, .reorder_point] | @tsv'
In Hermes, the script lives in ~/.hermes/scripts/ and runs as a "no-agent" job. The Hermes scheduled-tasks documentation says these jobs use no tokens and no model, that empty output sends nothing, and that a failed script sends an error alert.
hermes cron create "0 7 * * 1-5" \ --no-agent \ --script low-stock.sh \ --deliver slack \ --name "low-stock-7am"
In OpenClaw, the same script runs as a command job. The OpenClaw automation payloads documentation says command jobs run on the gateway host "without starting a model-backed turn" and post their output to the channel you choose.
openclaw automations create "0 7 * * 1-5" \ --name "Low stock 7am" \ --command "scripts/low-stock.sh" \ --command-cwd "/srv/agent" \ --announce \ --channel slack \ --to "channel:C0123456789"
Both post the same report, with no model bill for this job.
2. Stuck-order triage (AI helps)
Sorting stuck orders by likely cause takes judgment, so this one uses the model. The goal: a message only when something needs attention, with no repeat alerts.
hermes cron create "0 10,15 * * *" \ "List paid orders older than 48 hours with no tracking number. Group them by likely cause and flag marketplace orders near their ship-by date. If there are none, reply with only [SILENT]." \ --continuity \ --deliver slack \ --name "stuck-orders"
In Hermes, the [SILENT] reply suppresses the message on quiet runs. --continuity feeds each run its previous report so it can skip orders it already flagged.
openclaw automations create "0 10,15 * * *" \ "List paid orders older than 48 hours with no tracking number. Group them by likely cause and flag marketplace orders near their ship-by date." \ --name "Stuck orders" \ --session isolated \ --announce \ --channel slack \ --to "channel:C0123456789"
In OpenClaw, --session isolated keeps each run's context small. To call the model only when stuck orders exist, add a condition watcher: a short script that counts stuck orders and fires the agent turn only when the count is above zero. It takes more setup than Hermes's [SILENT] line but skips the model call entirely on quiet days.
3. Supplier follow-ups (AI drafts, a person sends)
Every Monday, the agent lists purchase orders (POs) past their promised ship date and posts a draft follow-up email per supplier to #purchasing. Use the same prompt on both, ending with "Do not send email or change any purchase order."
That sentence is a request, not a lock. The lock is the tool list: read access to POs and nothing that sends email or edits a PO. On Hermes, list only read tools in the server's tools.include; on OpenClaw, use toolFilter.include. The purchase order generator is a simple way for the buyer to turn an approved draft into a clean PO.
After two or three good runs, ask the agent to save the procedure as a skill. With the review gate on (skills.write_approval on Hermes, propose mode on OpenClaw), you read the steps before they become permanent. That is when you catch a wrong assumption, such as counting a partly shipped PO as late.
Choose Hermes if, choose OpenClaw if
Choose Hermes Agent if
- You want a hosted option from the maker and don't want to run a server (Hermes Cloud).
- You prefer cautious defaults: approval checks on by default and risky commands blocked in scheduled jobs.
- One or two people will talk to the agent, so per-person permission limits don't matter.
- You want to try many models through one account (Nous Portal).
- You can tolerate a very fast release pace and will pin versions yourself.
Choose OpenClaw if
- Several teammates will message the agent and should not all have the same powers.
- You want separate agents with separate permissions, such as a read-only ops agent and a sandboxed research agent.
- You value a project that publishes its own security fixes and offers a built-in security audit.
- You are comfortable hardening defaults on day one: the cautious exec preset, sandboxing, and the heartbeat off or isolated.
- You already run it. With the defaults tightened, the store-side benefits of switching are small.
If none of these tip it, pick the one someone on your team already knows.
What to do this week, and the rest of the month
This plan keeps the agent read-only until it earns trust.
This week: install and lock down
- Put the agent on a machine used only for it (a small VPS, a spare computer or Hermes Cloud), pin a version and turn off automatic updates.
- Tighten defaults. On Hermes, set
skills.write_approval: trueandmemory.write_approval: true, and keepcron_mode: deny. On OpenClaw, runopenclaw exec-policy preset cautious, set self-learning topropose, turn the heartbeat off or isolate it, then runopenclaw security audit. - Turn on DM pairing or an allowlist so only your team can message the agent.
- Create a read-only store token and connect one MCP server with a tool allowlist.
- Set up the low-stock script job and compare its list with your own for five days.
Weeks 2 and 3: add judgment, still read-only
- Add stuck-order triage and check every flagged order by hand for a week, noting what the agent misreads (time zones, cancellations, bundles).
- Add the supplier draft job. Have the buyer send the emails.
- Check your model provider's usage page twice a week, and drop any job that costs more than the time it saves.
Week 4: decide on writes
Pick the one write action you most want to hand over, such as creating a draft PO. Add only that tool, with a token that can create drafts but not submit or send them, and practise revoking it.
The agent is only as good as the stock number it reads. If your channels don't yet share one inventory count, fix that before you connect either agent. Nventory's Free plan connects one channel with unlimited orders and no card, and its pricing page lists the plans for more channels, or you can start free here.
Frequently Asked Questions
Mostly, yes. Hermes ships a hermes claw migrate command that imports OpenClaw settings, memories, skills and API keys, with a dry-run option to preview the move. Your store connection carries over too, because both agents talk to business systems through MCP servers. Plan to re-test every scheduled job after a move, since schedule syntax and delivery settings differ.
The software is free for both. Your bill is hosting plus model usage, and model usage depends far more on settings than on which agent you pick. The two big levers are the model you choose and how many unnecessary model calls you allow, such as OpenClaw's default 30-minute heartbeat or background learning passes. Checks that need no judgment can run as plain scripts in both, at zero model cost.
Hermes keeps two small files that load into every session: MEMORY.md, capped at 2,200 characters, and USER.md, capped at 1,375. It can also search past conversations. OpenClaw keeps USER.md, MEMORY.md and dated daily notes, and a background dreaming pass moves useful notes into long-term memory. For a store, both should remember conventions only. Stock and order numbers should always come from live tools.
Yes. The Hermes Skills Hub lists ClawHub as a community source, so the malware history of that marketplace matters to Hermes users as well. Hermes runs its own scanner on every hub install and refuses skills it rates dangerous, even with the force flag. Read any third-party skill line by line before installing it, whichever agent you run.
They are close. Each installs with a one-line script on macOS, Linux or Windows, then walks you through choosing a model and connecting a chat app. Hermes adds a paid hosted option, Hermes Cloud, if you don't want to manage a server. OpenClaw has no hosted tier of its own but documents many hosting routes. Budget an afternoon for either, plus two weeks of read-only testing.
