Hermes in the Enterprise: A Security Team's Threat Model
Hermes in the Enterprise: A Security Team's Threat Model
Hermes Agent is heading for enterprise networks. On 7 October 2026 the Wall Street Journal reported that Nous Research had raised $90 million at a $1.5 billion valuation to take its open-source agent from individual developers into corporate environments — the article notes more than 22 million downloads since February. This is no longer a tool your most curious engineer runs on a spare VM. It will be bought, piloted and integrated by organisations that answer to auditors, insurers and customers.
If your team is evaluating Hermes or any agentic AI of this class, what should your security function ask first? Not "is it safe?" — that never gets a useful answer. The right questions: what can this thing do without me, who can talk to it, and how would I know when it did something I would not have approved?
What Hermes actually does
Hermes Agent is an autonomous agent from Nous Research, the open-weights lab behind the Hermes model family — not a chatbot bolted onto an API. It uses tool calling to read and write files, run shell commands and code, browse the web, call external APIs and model context protocol servers, and spawn isolated subagents. It runs on terminal backends from your own machine to Docker containers, remote SSH hosts or serverless sandboxes.
Two features matter most for a threat model. First, the learning loop: Hermes creates and updates its own skills — reusable procedural documents it loads when a task matches — and the official docs state the agent can modify or delete any skill, with human-approved skill writes as an optional gate that is off by default. Second, the messaging gateway: the same agent is reachable from Telegram, Discord, Slack, WhatsApp, Microsoft Teams and more, with built-in scheduling so it runs unattended jobs around the clock. You can message it from your phone while it executes on a server you never log into.
That combination — persistent autonomy, self-modifying procedure, conversational remote control — is genuinely useful, and a prompt-injection class in production form.
The threat model
Prompt injection via web content and files. Hermes reads what you point it at: web pages, documents, repositories, emails. Instructions hidden in that content are indistinguishable from instructions you gave it. This is not hypothetical — in July 2026 researchers published a proof of concept, reported by The Hacker News, showing leading coding agents tricked into executing attacker-supplied binaries via a poisoned README in an open-source repository. Hermes has mitigations — context-file scanning, always-on request forgery protection blocking fetches to private networks, a configurable website blocklist — but the scanning is explicitly heuristic, not semantic. Treat web and file content as untrusted input to a privileged process.
Malicious or drifting self-created skills. Skills are the agent's procedural memory, drawn from bundled sets, community hubs, or the agent writing its own. A skill is a markdown document that shapes future behaviour — instructions to ignore approvals, exfiltrate data or quietly widen scope would persist until someone reads the file. Default configuration writes skills freely, and a skill that was correct in March can drift as the environment changes.
Chat account takeover. Hermes treats whoever can message the gateway as the operator. An attacker who compromises a linked Telegram or Discord account inherits the agent's access. The gateway has layered authorisation — per-platform user allowlists, pairing codes, deny-by-default when nothing is configured — but consumer chat platforms are not your identity provider.
Secrets exposure. The agent holds model API keys and can be handed cloud, email and code-hosting credentials. Redaction exists for tool error output, but anything the agent can read — environment files, notes, logs — can end up in a model prompt, a chat message, or a file it writes somewhere public.
Over-broad permissions. With dozens of built-in tools and shell execution, a default install is a small operating system with your name on it. YOLO mode exists and disables approval prompts entirely; so does flipping approvals off in configuration. Headless surfaces default to denying dangerous commands, but defaults are a floor, not a control set.
Weak audit trail. Hermes logs locally and keeps session history, but if your security operations team cannot query it, you do not have an audit trail — you have files. An agent acting every few minutes needs central logging before it needs a use case.
Controls that actually apply
Most of these exist in Hermes today. The gap in most pilots is that nobody configures them.
Sandbox execution. Run command execution on the container backend — which drops Linux capabilities, blocks privilege escalation and limits processes — or on a dedicated host behind an SSH backend. Never a domain-joined workstation.
Gate skill writes and new skills. Turn on the skill write-approval gate so every skill the agent creates or edits is staged for human review. Review staged skills with the seriousness you give code changes; functionally they are.
Least-privilege credentials. Give the agent scoped keys with hard spend caps, separate from any human admin account. Never one credential that unlocks both infrastructure and production data.
Disable consumer chat channels, or fence them properly. If the pilot needs messaging, use a dedicated workspace with explicit user allowlists and pairing codes; if not, run CLI-only. Do not bridge an enterprise agent to a personal Telegram account.
Central logging. Ship gateway and session logs to your existing log platform on day one. Log authorisation denials, approval decisions and skill writes at minimum.
A kill switch that is really a switch. Hermes does not ship a single named panic button; build one. Stopping the gateway service, revoking platform tokens and disabling API keys are three independent ways to stop it — documented and tested, and placed outside the agent, since Hermes deliberately blocks self-restart from inside its own supervised process.
Keep it updated, and say no to the easy paths. Run updates for security patches, keep allow-all-user flags off in production, and treat "just disable approvals for the pilot" as the decision it really is.
Field notes
I run Hermes myself across my own operations, so these come from real deployments, not lab speculation. Tool names stay; internal infrastructure stays out.
- Secrets do end up in transcripts. A dedicated scanning and rotation process found hundreds of credential-shaped findings in Hermes session transcripts over a few months — API keys and bearer tokens reaching cloud model contexts through normal tool use. Hermes redacts tool error output, but that is one layer; what the agent reads can still reach a model prompt. Scanning after the fact is no substitute for not feeding it secrets.
- Skill self-management is double-edged. The agent's ability to create and edit its own skills is useful — and it also meant every profile loaded the full skill catalogue regardless of relevance, wasting tokens and polluting context with instructions for domains that did not apply. Per-profile isolation was needed to fix it. If the agent writes its own procedures, someone has to review what accumulated.
- Approval gates create friction — and that friction is the point. My approval gate blocking routine cleanup commands (think
rm -rfon temp directories after agent work) caused files to pile up. The fix is not to weaken the gate; it is to document an approved alternative path. If your team starts asking to disable approvals to "get things moving", you have found the real control. - Monitoring the agent matters as much as the agent itself. My own budget-and-spend guard cron fired a false critical alert for days because of a logic quirk — reading the wrong figure as budget exhaustion. Automated agents generate automated noise; if nobody owns the monitoring that watches the agent, false alarms and real ones look identical.
Pre-deployment checklist
Before Hermes (or anything like it) touches enterprise data:
- Written threat model covering injection via every content source the agent reads.
- Command execution sandboxed — container or dedicated host, non-root, resource-limited.
- Skill write-approval enabled; a named human reviews staged skills.
- Credential inventory complete; agent keys scoped, capped and separable.
- Chat channels either disabled or on allowlists behind a dedicated workspace.
- Central logging live and queried at least once before go-live.
- Kill switch documented, owned and tested from outside the agent.
- YOLO and approval-off modes blocked by configuration review.
- Update and patch cadence assigned to a named owner.
- A rollback: what you will do when the first incident happens.
Most of this is unglamorous. None of it is optional. Agentic AI rewards teams that treat the control plane as the product.
If your team is evaluating agentic AI and wants a second pair of eyes on the threat model before a pilot starts, my services page covers security and security strategy work for UK businesses — or get in touch and we can talk through where you are.
Sources
- WSJ Pro, "Nous Research Scores $90 Million to Bring Open-Source AI Assistant to Enterprises", 7 October 2026 (funding figure, valuation, enterprise strategy, download count).
- Hermes Agent official documentation, Nous Research: overview, skills system, security user guide, messaging gateway.
- The Hacker News, "Top AI Agents Built to Catch Malicious Code Can Be Tricked Into Running It", 9 July 2026, covering the AI Now Institute "Friendly Fire" proof of concept.
Flagged, not stated as fact: "kill switch" is operational framing, not a documented Hermes feature; prompt-injection evidence cited is from competing agents; download counts are vendor-reported via WSJ.