ในบทความนี้
- ห้าประตูที่งานเดินเข้ามาหา Gateway
- กายวิภาคของ Cron Job
- Zero-Token Automation — งานที่ไม่ต้องปลุก LLM
- Cron ที่จำได้ และวงจร Ops
- Inbound Webhooks ที่ตรวจลายเซ็นก่อนเชื่อ
- ส่งเหตุการณ์ออกไป — Outbound Webhooks, Shell Hooks และ Deliverable Mode
- Blueprints — Skill ที่มีตารางเวลา
- Hermes เป็น Backend ให้แอปที่องค์กรมีอยู่แล้ว
- Hardening และการ Scale ระนาบ Automation
In this post
- Five Doors Into the Gateway
- Anatomy of a Cron Job
- Zero-Token Automation: Work That Never Wakes the Model
- Cron That Remembers — and the Ops Loop
- Inbound Webhooks Done Properly
- Pushing Events Out: Outbound Webhooks, Shell Hooks, and Deliverable Mode
- Blueprints: A Skill With a Schedule
- Hermes as a Backend for the Apps You Already Run
- Hardening and Scaling the Automation Plane
🤔 ลองนึกภาพเช้าวันจันทร์ครับ — ก่อนที่ใครในทีมจะเปิด laptop มีสรุป PR ที่ค้างรีวิวมาทั้งสุดสัปดาห์รออยู่ใน Slack แล้ว dependency ที่เพิ่งมี CVE ใหม่เมื่อคืนก็ถูกเปิด ticket ไว้แล้ว และไม่มีใครพิมพ์คำสั่งอะไรเลยตั้งแต่เย็นวันศุกร์ คำถามของบทความนี้ไม่ใช่ว่า agent ทำแบบนั้น ได้ไหม แต่คือ งานพวกนี้ควรเดินเข้ามาหา agent ทาง ประตูไหน — ตามเวลา ตามเหตุการณ์ หรือให้แอปอื่นเรียกเข้ามา — และรอบไหนบ้างที่ไม่จำเป็นต้องจ่ายค่า token สักบาท
พื้นฐานของ cron อยู่ใน #2 Agent Teams หัวข้อ "Always-On Work" ครบแล้ว — action ของ tool cronjob รูปแบบตารางเวลา session ที่แยกขาดกันทุกรอบ continuity/context_from ปลายทางส่ง [SILENT] no_agent และ "cron ที่จำได้" ของ v0.21.0 — ส่วนใน #5 Integrations API server กับ A2A ได้ที่ทางเพียงหัวข้อสั้น ๆ บทความนี้จึงไม่แนะนำสิ่งเหล่านั้นซ้ำ แต่เป็นชั้น ops ที่ต่อจากตรงนั้น: ไฟล์และ ledger เบื้องหลังระบบ cron ในตัวของ Hermes Agent ค่า default ที่ควรรู้ วงจร incidents/doctor/runs webhook ขาเข้าที่ตรวจลายเซ็นก่อนเชื่อ hook และ webhook ขาออกที่ส่งเหตุการณ์ไปหาระบบที่องค์กรมีอยู่แล้ว blueprint ที่เป็นแค่ skill พร้อมตารางเวลา และ API server ที่ให้แอปภายในเรียก Hermes ได้เหมือนเรียก model ตัวหนึ่ง
สำหรับคนที่ตามซีรีส์ OpenClaw ของบล็อกนี้มา รูปแบบ heartbeat ที่ผมเล่าใน OpenClaw #5: Integrations คือจุดเปรียบเทียบที่ดี — แนวคิด "ปลุก agent เป็นระยะให้ไปดูว่ามีอะไรต้องทำไหม" ยังอยู่ แต่ Hermes แยกมันออกเป็นชิ้นส่วนที่มีชื่อ ตรวจสอบได้ และที่สำคัญ ปิด LLM ทิ้งได้ในรอบที่ไม่มีอะไรเปลี่ยน ผมจะเปรียบเทียบเชิงคุณภาพเท่านั้น ไม่มีตัวเลขใดในบทความนี้มาจากการวัดของผมเอง ทุกเวอร์ชัน ค่า default และคำพูดที่ยกมา อ้างอิงเอกสารทางการและ release notes บน GitHub ของ Nous Research ณ วันที่ 1 กันยายน 2026 (v0.21.0, tag v2026.8.31)
ห้าประตูที่งานเดินเข้ามาหา Gateway
สิ่งแรกที่ควรเข้าใจคือ ระบบ automation ทั้งหมดของ Hermes อาศัย gateway daemon ตัวเดียวกับที่รับข้อความจาก Telegram หรือ Slack — เครื่องเดียวกับ VPS หรือ Docker container ที่ผมเล่าใน #7 Production ไม่มี scheduler แยก ไม่มี worker pool แยก งานเดินเข้ามาหา daemon ตัวนี้ได้ห้าทาง:
| ประตู | กลไก | เหมาะกับ | มาตั้งแต่ |
|---|---|---|---|
| Cron | gateway tick ทุก 60 วินาที อ่าน ~/.hermes/cron/jobs.json | รายงานประจำ, watchdog, งานตามรอบ | มีตั้งแต่ release แรกที่มองเห็นบน GitHub (v0.2.0, 12 มี.ค. 2026) |
| Inbound webhook | HTTP POST เข้าพอร์ต 8644 พร้อมลายเซ็น HMAC | GitHub/GitLab events, ระบบ monitoring, n8n | v0.4.0 (มี.ค. 2026) |
| Hooks / outbound webhooks | shell script ในจังหวะ lifecycle + HTTP POST ขาออกที่ลงลายเซ็น | veto คำสั่งอันตราย, ส่ง audit event ไป SIEM | shell hooks v0.11.0 (เม.ย. 2026), outbound v0.20.0 (ส.ค. 2026) |
| API server | OpenAI-compatible endpoint ที่พอร์ต 8642 | Open WebUI, แอปภายในองค์กร, hermes peer | v0.4.0 |
| A2A / peer | JSON-RPC ตามมาตรฐาน A2A v1.0 ที่พอร์ต 9900 / peer ผ่าน API server | agent ต่าง host คุยกัน | A2A v0.20.0, hermes peer v0.21.0 |
คำแนะนำสำหรับองค์กรที่เพิ่งเริ่ม: เริ่มจาก cron เพราะมันเป็นประตูเดียวที่คุณกำหนดจังหวะเอง ไม่มีใครภายนอกยิงเข้ามา และตั้งแต่ v0.13.0 มีโหมดที่ไม่แตะ model เลย ถัดมาคือ inbound webhook เมื่อคุณอยากให้ GitHub หรือระบบ monitoring เป็นคนปลุก ส่วน API server และ A2A เก็บไว้ท้ายสุด — ทั้งคู่เปิดพื้นผิวเครือข่ายที่มีประวัติ CVE เป็นของตัวเอง ซึ่งผมเก็บไว้เล่าในหัวข้อสุดท้าย
💡 ถ้าอ่านตอน #2 มาแล้ว ให้มองอย่างนี้ครับ: heartbeat คือ "agent ที่ตื่นมาถามว่ามีอะไรไหม" ส่วน cron job ของ Hermes คือ "หน่วยงานที่มีชื่อ มีตารางเวลา และมี ledger ของตัวเอง" — และในสองหัวข้อถัดไปคุณจะเห็นว่ามันแยก "การตื่น" ออกจาก "การคิด" ได้
กายวิภาคของ Cron Job
Cron ของ Hermes ไม่ใช่ crontab ของ Linux ที่ยิงคำสั่งเปล่า ๆ แต่เป็น job ที่ห่อ prompt (หรือ script) ไว้พร้อม metadata ทั้งหมดถูกเก็บเป็นไฟล์ธรรมดาที่คุณเปิดอ่านได้เอง:
~/.hermes/cron/
├── jobs.json # นิยาม job ทุกตัว — เขียนแบบ atomic
├── executions.db # ledger: claimed → running → completed | failed | unknown
├── .tick.lock # file lock กัน tick ซ้อนกัน
└── output/
└── {job_id}/
└── {timestamp}.md # ผลลัพธ์ทุกรอบ ย้อนอ่านได้
tick ทุก 60 วินาทีที่ตอน #2 เล่าไว้ มีนัยเชิง ops สองข้อ: ความละเอียดต่ำสุดของตารางเวลาคือหนึ่งนาที และ .tick.lock คือสิ่งที่กันไม่ให้สอง process (เช่น gateway สองตัวที่เผลอ mount volume เดียวกัน) รัน job เดียวกันซ้ำ
สร้าง job ได้สามทาง — CLI (hermes cron create), tool ชื่อ cronjob ที่ agent เรียกเองใน session และ dashboard บนพอร์ต 9119 (ซึ่งตั้งแต่ v0.4.0 ก็มี REST API /api/jobs หนุนอยู่ข้างหลัง) ผมขอแสดงรูปแบบของ tool เพราะเป็นรายการพารามิเตอร์ที่เอกสารระบุครบที่สุด — ใน chat คุณเพียงบอก agent ว่าอยากได้ job แบบไหน มันจะเรียกสิ่งนี้ให้ ส่วน flag ฝั่ง CLI ให้ดูจาก hermes cron create --help:
# tool เดียวรวมทุก action — release v0.3.0 (17 มี.ค. 2026) รวมหลายคำสั่งเป็นตัวเดียว
cronjob(
action="create", # list | update | pause | resume | run | remove
schedule="0 9 * * 1-5", # cron 5 ช่อง — เช้าวันทำงาน
name="morning-brief",
prompt="สรุป issue ที่เปิดใหม่ใน repo ทีมเมื่อวาน จัดกลุ่มตาม label",
deliver="slack",
skills=["github-triage"], # แนบ skill จากตอน #6
workdir="/srv/repo", # scope file/terminal tools + โหลด AGENTS.md ของ repo นั้น (absolute path เท่านั้น)
enabled_toolsets=["web", "file", "terminal"],
continuity=True, # v0.21.0 — จำผลรอบก่อนของตัวเอง
context_from="deps-scan", # v0.12.0 — ต่อจากผลสำเร็จล่าสุดของ job อื่น
reasoning_effort="high", # v0.21.0 — pin ต่อ job
no_agent=False,
attach_to_session=False,
)
รูปแบบตารางเวลาอยู่ในตอน #2 แล้ว (หน้า feature ระบุห้าแบบ: relative, interval, ภาษาธรรมชาติ, cron 5 ช่อง และ ISO timestamp) มีข้อเดียวที่ควรเสริม: หน้า feature ของ cron ลิสต์รูปแบบภาษาธรรมชาติอย่าง daily at 7am ไว้ และตอน #2 ก็ใช้ weekdays at 9am เป็นตัวอย่าง แต่คู่มือ automate-with-cron ซึ่งเป็นเอกสารทางการอีกหน้าหนึ่ง เขียนว่าภาษาธรรมชาติอย่าง "daily at 9am" ไม่รองรับ — ผมยังไม่ได้ทดสอบเองว่าหน้าไหนตรงกับโค้ด ในงานจริงจึงแนะนำให้ยึด cron 5 ช่องเป็นหลัก และถ้าอยากใช้ภาษาธรรมชาติ ให้ทดสอบบน build ของคุณเอง (v0.21.0) ก่อนพึ่งพา
ค่า default ในบล็อก cron: ของ config.yaml มีหลายตัวที่องค์กรควรรู้จักก่อนคิดจะแก้:
| Key | Default | ความหมาย |
|---|---|---|
script_timeout_seconds | 3600 | script รันได้นานสุดหนึ่งชั่วโมง (override ด้วย HERMES_CRON_SCRIPT_TIMEOUT) |
failure_nudge_threshold | 3 | ล้มติดกันกี่ครั้งจึงจะสะกิดเตือน |
misfire_grace_minutes | 10 | gateway ดับแล้วกลับมาภายในสิบนาที job ที่พลาดไปยังได้รัน |
model_drift_guard | true | job ที่ไม่ได้ pin model จะ snapshot provider/model ตอนสร้าง — ถ้า default ทั้งระบบเปลี่ยน job จะข้ามรอบพร้อมแจ้งเตือนครั้งเดียว |
allow_agent_scheduling | false | session ที่ cron ปลุกขึ้นมา สร้าง cron job เพิ่มไม่ได้ |
preflight | true | ตรวจ API key, env ของ skill ที่แนบ และปลายทางส่ง ก่อน dispatch |
mirror_delivery | false | ถ้า true ผลทุกรอบจะไปโผล่เป็น thread ที่คุยต่อได้ |
max_parallel_jobs | (ไม่ตั้ง) | เพดานจำนวน job ที่รันพร้อมกัน |
cleanup_timeout_seconds | 10 | เพดานเวลาเก็บกวาดหลังจบรอบ |
media_send_timeout_seconds | 300 | เพดานเวลาส่งไฟล์สื่อไปยังปลายทาง |
bot_chat_delivery_timeout_seconds | 600 | เพดานเวลาส่งผลลัพธ์เข้า Bot Chat |
allow_agent_scheduling: false ที่ตอน #2 อธิบายไว้แล้ว prompt ทุกตัวยังถูกสแกนตอน create/update หา prompt injection, Unicode ที่มองไม่เห็น, SSH backdoor และรูปแบบขโมย credential — ถ้าเข้าข่ายจะถูกบล็อกก่อนได้รันด้วยซ้ำ นี่คือบทเรียนโดยตรงจากเรื่อง malicious skill ที่ผมเล่าไว้ใน #4 Security
เรื่อง ปลายทางส่ง — ตอน #2 ยกตัวอย่างไว้บางส่วน รายการเต็มตามหน้า feature คือ origin, local, telegram, discord, slack, whatsapp, signal, matrix, mattermost, email, sms, homeassistant, dingtalk, feishu, wecom, weixin, bluebubbles, qqbot, bot-chat และ all ระบุ chat และ thread เจาะจงได้ (telegram:-100123:17585) หรือ fan-out หลายปลายทางคั่นด้วย comma กติกา [SILENT] อยู่ในตอน #2 แล้ว สิ่งที่ควรเพิ่มคือสองข้อ: ผลลัพธ์ที่ถูกเงียบยังถูกบันทึกลง output/ และ job ที่ล้มเหลวจะส่งเสมอ ไม่ว่าคุณจะขอให้เงียบแค่ไหน
ฝั่ง tool ที่ job ใช้ได้ — toolsets.py นิยาม toolset hermes-cron ให้สะท้อน hermes-cli แล้วค่อยถูกทับด้วย enabled_toolsets ต่อ job และ config ของ platform cron ใน hermes tools ตามลำดับ ส่วน workdir นอกจากจะจำกัดขอบเขต file/terminal tools แล้ว ยังโหลด AGENTS.md, CLAUDE.md หรือ .cursorrules จาก directory นั้นเข้ามาด้วย — cron job ที่ทำงานกับ repo จึงได้กติกาของ repo นั้นติดมาโดยไม่ต้องยัดลง prompt
Zero-Token Automation — งานที่ไม่ต้องปลุก LLM
นี่คือหัวข้อที่ทำให้ผมมองระบบ cron ของ Hermes ต่างจาก heartbeat ยุค OpenClaw ที่สุด ในรูปแบบ heartbeat ทุกรอบที่ตื่นคือหนึ่ง LLM call แม้ไม่มีอะไรเปลี่ยน — Hermes มี สามชั้น ที่ตัดค่าใช้จ่ายส่วนนี้ทิ้ง
ชั้นที่ 1 — script-only job (--no-agent, v0.13.0)
ตอน #2 เอ่ยถึง no_agent=True ไว้ประโยคเดียว นี่คือกติกาฉบับเต็มของมัน — Release v0.13.0 "The Tenacity Release" (7 พฤษภาคม 2026) เพิ่ม watchdog mode: job ที่รัน script ธรรมดาและไม่มี agent อยู่ในวงจรเลย
# script ต้องอยู่ใน $HERMES_HOME/scripts/ เท่านั้น — path traversal ถูกปฏิเสธ
cat > ~/.hermes/scripts/watchdog.sh <<'EOF'
#!/usr/bin/env bash
# shebang ถูกเพิกเฉย — .sh/.bash รันด้วย bash เสมอ ไฟล์อื่นรันด้วย Python ตัวปัจจุบัน
code=$(curl -s -o /dev/null -w '%{http_code}' https://intranet.example.ac.th/health)
if [ "$code" != "200" ]; then
echo "⚠️ intranet health = $code" # มี stdout → ข้อความนี้ถูกส่งตรง ๆ
exit 0
fi
# exit 0 + stdout ว่าง = tick เงียบ ไม่ส่งอะไร ไม่เสีย token
EOF
hermes cron create "every 5m" --no-agent --script watchdog.sh --deliver telegram
กติกาของ watchdog สรุปได้สี่บรรทัด: exit 0 + stdout ว่าง = เงียบ; มี stdout = ส่งข้อความนั้นออกไปตรง ๆ; exit ไม่เป็นศูนย์หรือ timeout = แจ้งเตือนข้อผิดพลาด; และถ้าบรรทัดสุดท้ายของ stdout เป็น {"wakeAgent": false} ระบบก็ถือว่าเงียบเช่นกัน ที่ผมชอบเป็นพิเศษคือ environment ของ subprocess ถูก ล้าง credential ของ provider ออก — script ที่คุณเขียน (หรือที่ agent เขียนให้) จะไม่ได้ API key ของ model ติดมือไปด้วย
ชั้นที่ 2 — monitor-mode ที่ข้าม LLM เมื่อไม่มีอะไรเปลี่ยน (v0.21.0)
v0.21.0 "Pantheon" (31 สิงหาคม 2026) เพิ่ม hash-suppressed change detection (PR #81138): job ในโหมด monitor จะเทียบ hash ของสิ่งที่มันเฝ้าดูกับรอบก่อน และถ้าเท่ากัน จะ ไม่เรียก LLM เลย — นี่คือคำตอบตรง ๆ ของปัญหา "heartbeat ที่จ่ายเงินให้กับการบอกว่าไม่มีอะไรเกิดขึ้น" ที่ผมบ่นไว้ในซีรีส์ก่อน ข้อควรระวังเรื่องแหล่งอ้างอิง: ณ วันที่เขียน โหมด monitor ปรากฏใน release notes ของ v0.21.0 และ PR #81138 เท่านั้น หน้า cron ในเอกสารที่ render แล้วยังไม่ได้บรรยายมัน ผมจึงยังไม่พบตัวเลขว่าประหยัดกี่เปอร์เซ็นต์และจะไม่เดาให้ แต่ตรรกะชัด: รอบที่ไม่มีการเปลี่ยนแปลงมีต้นทุน token เป็นศูนย์
ชั้นที่ 3 — deliver_only และ hermes send
ฝั่ง webhook ก็มีของแบบเดียวกัน — route ที่ตั้ง deliver_only: true (v0.11.0, PR #12473) จะส่ง payload ที่แปลงแล้วไปยังช่องทางปลายทางโดยไม่ผ่าน agent และสำหรับ script ใด ๆ บนเครื่อง มี hermes send จาก v0.15.0 "The Velocity Release" (28 พฤษภาคม 2026, PR #27188) ที่ release note สรุปว่า "pipe any script's output to any messaging platform":
# ส่งข้อความหนึ่งครั้งไปยังช่องทางที่ตั้งค่าไว้ — ไม่มี agent loop ไม่มี token
hermes send --to slack "Backup เสร็จแล้ว: 2.1 GB"
hermes send --to telegram --file /var/log/nightly-report.pdf
hermes send --list # ดูปลายทางที่ใช้ได้
# ต่อท้าย script เดิมขององค์กรได้ทันที
./run-etl.sh 2>&1 | tail -n 5 | xargs -0 hermes send --to slack
ข้อละเอียดที่มีผลกับ script บน server: สำหรับ platform ที่ใช้ bot token — Telegram, Discord, Slack, Signal, SMS และ WhatsApp Cloud API — hermes send ส่งได้เองโดยไม่ต้องมี gateway รันอยู่ ส่วน platform ที่มาเป็น plugin ยังต้องมี gateway ทำงานอยู่
Cron ที่จำได้ และวงจร Ops
สะพานข้ามรอบทั้งสามอยู่ในตอน #2 แล้ว — context_from (v0.12.0, 30 เมษายน 2026, PR #15606) ที่นำผลสำเร็จล่าสุดของ job อื่น มาต่อหน้า prompt, continuity=True ที่ฉีดผลรอบก่อนของ job เดียวกัน และ persistent memory กับ notepad ต่อ job ของ v0.21.0 สองข้อที่ตอน #2 ไม่ได้บอก: memory ที่ cron agent โหลดคือ MEMORY.md ตัวเดียวกับที่มีขอบเขตและเพดานตาม #3 Memory (สิ่งที่ job เรียนรู้ตอนตีสองจึงอยู่กับ session ตอนเช้าของคุณด้วย) และ ณ วันที่เขียน ทั้ง memory และ notepad ของ cron ปรากฏใน release notes ของ v0.21.0 กับ PR #81138 เท่านั้น — หน้า cron ที่ render แล้วยังบรรยายแค่ continuity และ context_from นี่คือหน้าตาของ chain ในงานจริง:
# chain สองขั้น: สแกนตอนตีหนึ่ง แล้วให้ job ตีสองอ่านผลนั้นต่อ
cronjob(action="create", name="deps-scan", schedule="0 1 * * *",
prompt="รัน pip-audit และ npm audit ใน repo สรุปเฉพาะรายการ CVSS ≥ 7",
workdir="/srv/repo", deliver="local")
cronjob(action="create", name="deps-triage", schedule="0 2 * * *",
context_from="deps-scan", # ผลสำเร็จล่าสุดของ deps-scan ถูกต่อหน้า prompt
continuity=True, # และจำว่ารอบก่อนตัวเองเปิด issue อะไรไปแล้ว
prompt="จากผลสแกนที่แนบมา เปิด issue ให้ทีม ข้ามรายการที่เคยเปิดแล้ว",
deliver="slack")
ของใหม่อีกสองชิ้นใน v0.21.0 ที่คนทำ ops จะชอบ: ผลลัพธ์ส่งเข้า Bot Chat ประจำตัวของ bot ได้ (แนวคิด Bot Mode อยู่ในตอน #2 — ผมจะไม่เล่าซ้ำ) และปุ่ม "Trigger now" (#70638) ใน Desktop app และ dashboard ที่รัน job ทันทีเพื่อทดสอบโดยไม่ต้องแก้ตารางเวลา — บน CLI คำสั่งเดียวกันคือ hermes cron run <job>
วงจร ops: incidents → doctor → runs
ระบบ cron ที่รันตอนไม่มีใครดู ต้องตอบคำถาม "เมื่อคืนมีอะไรพังไหม" ได้ในสามคำสั่ง:
# 1) เหตุการณ์ที่ระบบตรวจพบ — detected → alerted → closed
hermes cron incidents
hermes cron incidents ack <incident-id> # v0.21.0: ack แล้ว signature เดิมจะไม่ ping ซ้ำ (#95017)
# 2) ตรวจสุขภาพแบบ read-only — exit 1 ถ้ามีปัญหา เหมาะใส่ใน monitoring
hermes cron doctor
# ตรวจ: รอบล่าสุดล้ม / ส่งไม่ถึง / next_run_at ค้าง / script หาย / workdir หาย
# 3) ประวัติการรัน
hermes cron runs deps-triage --limit 20
รายละเอียดที่ผมให้คะแนน: ข้อความ error ใน incident ถูก redact ก่อนเก็บ เพราะ stack trace มักลาก credential ติดมาด้วย และ hermes cron doctor ถูกออกแบบให้ exit code มีความหมาย — เอาไปวางใน monitoring ของตอน #7 ได้ทันที ส่วน model-drift guard คือความรอบคอบที่ผมยังไม่เคยเห็นในระบบ cron ของ agent ตัวอื่น: ถ้าคุณเปลี่ยน model default ทั้งระบบด้วย hermes model job ที่ไม่ได้ pin model จะข้ามรอบและแจ้งเตือนครั้งเดียว แทนที่จะรันเงียบ ๆ ด้วย model ที่คุณไม่เคยทดสอบกับ prompt นั้น (ลำดับการเลือก model คือ pin ต่อ job → cron.model → default ทั้งระบบ)
สำหรับผลลัพธ์ที่อยากคุยต่อ — ตั้ง cron.mirror_delivery: true ทั้งระบบ หรือ attach_to_session=True ต่อ job แล้วบน platform ที่มี thread (Telegram topics, Discord/Slack threads) ทุกรอบจะเปิด thread ใหม่ที่คุณตอบกลับได้เหมือนคุยกับ agent ปกติ Slack แบบ in-channel ต้องตั้ง slack.cron_continuable_surface: in_channel พร้อม reply_in_thread: false เพิ่ม
Inbound Webhooks ที่ตรวจลายเซ็นก่อนเชื่อ
Webhook adapter มาพร้อม v0.4.0 (tag v2026.3.23, เผยแพร่ 24 มีนาคม 2026) ในชุดเดียวกับ Signal, DingTalk, SMS, Mattermost และ Matrix — Hermes มองมันเป็น "platform" หนึ่งเหมือน Telegram เพียงแต่ผู้ส่งคือเครื่องจักร ไม่ใช่คน กติกาพื้นฐานทั้งหมดอยู่ในตารางเดียว:
| รายการ | ค่า |
|---|---|
| พอร์ต default | 8644 (env WEBHOOK_PORT; เปิดด้วย WEBHOOK_ENABLED) |
| Endpoint | POST /webhooks/<route-name> และ GET /health |
| ลายเซ็นที่รับ | GitHub X-Hub-Signature-256 · GitLab X-Gitlab-Token · Generic V2 X-Webhook-Signature-V2 + X-Webhook-Timestamp (HMAC-SHA256 ของ <timestamp>.<body> ยอมรับภายใน ±300 วินาที) · Generic V1 เลิกใช้แล้ว ไม่กัน replay |
| Secret | ทุก route ต้องมี — ไม่มีข้อยกเว้น |
| Rate limit | 30 req/นาที ต่อ route |
| ขนาด body | 1 MB |
| Idempotency | cache 1 ชั่วโมงบน X-GitHub-Delivery / X-Request-ID |
| Response codes | 200 · 400 · 401 · 404 · 413 · 429 · 502 |
Route แบบ static อยู่ใน config.yaml ใต้ platforms.webhook.extra.routes.<name> — นี่คือโครงของ recipe รีวิว PR ที่เอกสารทางการเผยแพร่ ปรับข้อความเป็นภาษาไทย:
# config.yaml — โครงตามคู่มือ webhook-github-pr-review ทางการ
platforms:
webhook:
enabled: true
extra:
port: 8644
rate_limit: 30
routes:
pr-review:
secret: "<long-random-secret>" # ใส่ค่าเดียวกันฝั่ง GitHub
events: [pull_request]
prompt: |
รีวิว PR "{pull_request.title}" ใน {repository.full_name}
payload ไม่มี diff — รัน: gh pr diff {number} --repo {repository.full_name}
สรุปความเสี่ยงและข้อเสนอแนะเป็นภาษาไทย ไม่เกิน 10 ข้อ
deliver: github_comment
alerts:
secret: "<another-secret>"
deliver_only: true # ไม่ผ่าน agent — zero token
deliver: telegram
สิ่งที่ template ทำได้: dot-notation อ้างฟิลด์ใน payload ({pull_request.title}) และ {__raw__} ที่ dump ทั้ง payload เป็น JSON ตัดที่ 4,000 ตัวอักษร นอกจากนี้แต่ละ route ยังมี filters (exists / missing / equals / not_equals / contains / in / in_file / regex ประกอบกันด้วย all / any / not) เพื่อคัดเหตุการณ์ก่อนถึง agent และ script สำหรับ transform ด้วย script ใน ~/.hermes/scripts/ — stdout ที่เป็น JSON จะแทน payload, stdout ว่างหรือ [SILENT] หรือ {"__hermes_ignore__": true} จะตอบ 200 แล้วทิ้งเหตุการณ์นั้นไปเงียบ ๆ ปลายทาง deliver ของ webhook มี log และ github_comment เพิ่มจากรายการช่องทางปกติ
web_search, web_extract, vision_analyze และ clarify (_HERMES_WEBHOOK_SAFE_TOOLS ใน toolsets.py ระบุเหตุผลไว้ตรง ๆ ว่า "to prevent prompt injection attacks") ไม่มี terminal ไม่มี file เอกสารสรุปหลักคิดไว้ประโยคเดียวที่ควรจำขึ้นใจ — "HMAC validation authenticates the sender, not the content" ลายเซ็นบอกว่า GitHub เป็นคนส่ง แต่ชื่อ PR นั้นใครก็เขียนได้ และคู่มือ PR-review ทางการก็เตือนซ้ำว่า "PR titles and descriptions are attacker-controlled" — ถ้าเปิดสู่อินเทอร์เน็ต ให้รัน gateway ใน Docker หรือ VM
ถ้า route ไหนต้องการ tool มากกว่านั้นจริง ๆ — recipe PR-review ข้างบนให้ agent รัน gh pr diff ซึ่งอยู่นอกสี่ tool นี้ — key toolsets ของ route จะ แทนที่ ค่า default ของ platform ทั้งชุด ไม่ใช่รวมเพิ่ม ดังนั้นเปิดเท่าที่ต้องใช้ และเปิดโดยรู้ตัวว่ากำลังยกกำแพงลง ฝั่ง GitHub คุณเพียงชี้ webhook ไปที่ https://<host>/webhooks/pr-review เลือก content type เป็น JSON ใส่ secret เดียวกัน และเลือก event Pull requests — คู่มือบอกว่า comment จะปรากฏบน PR ภายในราว 30–90 วินาที และเหตุการณ์ซ้ำจาก GitHub ถูก dedupe ไว้หนึ่งชั่วโมง
Subscribe แบบไม่ต้อง restart
นอกจาก route แบบ static ยังมี dynamic subscription ที่มีผลทันที:
hermes webhook subscribe ci-failed \
--events workflow_run \
--prompt "อ่าน log ของ workflow ที่ล้ม สรุปว่าสาเหตุน่าจะเป็นอะไร" \
--deliver telegram --deliver-chat-id -100123 \
--skills ci-triage
hermes webhook list
hermes webhook test ci-failed --payload ./sample.json # ยิง payload ตัวอย่างใส่ route
hermes webhook remove ci-failed
# เก็บที่ ~/.hermes/webhook_subscriptions.json
# ถ้าชื่อชนกับ route ใน config.yaml — route static ชนะ
เอกสารเขียนไว้ชัด: "No gateway restart required — subscribe and it's immediately live" และมีจุดที่ผมอยากให้สังเกต — ตัว agent เอง ไม่มี tool สำหรับ webhook เมื่อคุณบอกใน chat ว่า "ช่วยตั้ง webhook รับ event จาก GitLab" มันจะใช้ terminal tool ยิงคำสั่ง hermes webhook subscribe ข้างบนนี้ โดยมี skill ชื่อ webhook-subscriptions ที่มาพร้อมระบบเป็นคู่มือ (โครงสร้าง SKILL.md อยู่ใน #6 Skills) แปลว่าถ้า session นั้นไม่มี terminal tool agent ก็ทำไม่ได้ — ซึ่งถูกต้องแล้วสำหรับ session ที่ webhook เป็นคนปลุก
ตลาดรอบ ๆ พอร์ต 8644 ก็เริ่มก่อตัว — Hookdeck เขียนคู่มือวาง relay ของตัวเองไว้หน้า Hermes และ fast.io เขียนคู่มือให้ n8n เรียก hermes webhook subscribe พร้อม node คำนวณ HMAC ทั้งสองเป็นเนื้อหาการตลาดของผู้ขาย ผมยกมาเพียงเป็นหลักฐานว่าองค์กรกำลังเอาเครื่องมือ workflow ที่มีอยู่แล้วมาต่อหน้าประตูนี้ ไม่ใช่คำแนะนำ
ส่งเหตุการณ์ออกไป — Outbound Webhooks, Shell Hooks และ Deliverable Mode
ประตูฝั่งขาออกมีสามแบบที่ต่างกันโดยสิ้นเชิง และผมอยากให้เห็นความต่างชัด ๆ เพราะเอกสารเรียกหลายอย่างว่า "hook" เหมือนกันหมด
1) Outbound webhooks ที่ลงลายเซ็น (v0.20.0)
PR #69406 (merge 2 สิงหาคม 2026 และออกใน v0.20.0 "Herald" วันถัดมา) เพิ่ม hooks.outbound: ทุกครั้งที่เหตุการณ์ใน lifecycle เกิดขึ้น gateway จะ POST ไปยัง URL ที่คุณกำหนดพร้อม HMAC — นี่คือวิธีที่หัวข้อ observability ในตอน #7 จะได้เห็น lifecycle event โดยไม่ต้อง scrape log:
# config.yaml
hooks:
outbound:
- name: siem
url: https://siem.example.ac.th/hermes
events: [on_session_end, subagent_stop, post_tool_call] # event ใดก็ได้ในแคตตาล็อก plugin hook
secret_env: HERMES_SIEM_SECRET # อ่านจาก ~/.hermes/.env — แนะนำกว่า secret: ตรง ๆ
timeout: 10 # 1–60 วินาที
matcher: "^(terminal|write_file)$" # regex คัดเฉพาะ tool ที่สนใจ สำหรับ event ระดับ tool
# สิ่งที่ปลายทางได้รับ
POST /hermes HTTP/1.1
Content-Type: application/json
X-Hermes-Event: post_tool_call
X-Hermes-Delivery: <delivery-id>
X-Hermes-Signature-256: sha256=<hex ของ HMAC-SHA256 บน raw body>
{ "hook_event_name": "post_tool_call", "tool_name": "terminal", "session_id": "…",
"delivery_id": "…", "timestamp": "…", … } # โครงเดียวกับ stdin ของ shell hook + สองฟิลด์ท้าย
นโยบายส่งเป็นแบบ fire-and-forget ผ่าน bounded queue และ daemon thread ตัวเดียว: ลองซ้ำหนึ่งครั้งพร้อม backoff เมื่อ connection error หรือ 5xx, ไม่ลองซ้ำเมื่อ 4xx, ไม่ตาม 3xx และการแก้รายการ target มีผลใน CLI session ถัดไปหรือหลัง restart gateway — พูดตรง ๆ คือมันไม่ใช่ message queue ที่รับประกันการส่ง ถ้าองค์กรต้องการ at-least-once ให้เอา relay ที่รับผิดชอบเรื่องนั้นมาคั่น hermes hooks list แสดง outbound target ทั้งหมดรวมกับ shell hook
2) Shell hooks ที่ veto ได้ (v0.11.0)
Shell hooks (v0.11.0, 23 เมษายน 2026, PR #13296) เป็นคนละเรื่องกับ webhook — มันคือ script บนเครื่องที่รันในจังหวะของ agent และบางจังหวะ ยับยั้งการกระทำได้:
# config.yaml
hooks:
pre_tool_call:
- matcher: "^terminal$"
command: ~/.hermes/agent-hooks/guard-terminal.sh
timeout: 30 # default 60 สูงสุด 300
fail_closed: true # ถ้า script พัง = บล็อก (default คือ fail-open!)
#!/usr/bin/env bash
# ~/.hermes/agent-hooks/guard-terminal.sh — อ่าน JSON จาก stdin
payload=$(cat)
cmd=$(printf '%s' "$payload" | jq -r '.tool_input.command // ""')
if printf '%s' "$cmd" | grep -Eq 'rm -rf /|mkfs|DROP TABLE'; then
echo '{"action":"block","reason":"คำสั่งอันตราย — ต้องให้คนอนุมัติ"}'
exit 2 # exit 2 = บล็อกเช่นกัน
fi
exit 0
ฟิลด์บน stdin ได้แก่ hook_event_name, tool_name, tool_input, session_id, cwd, extra คำตอบของ pre_tool_call รับทั้ง {"action":"block"|"modify"} และรูปแบบของ Claude Code {"decision":"block"} — ตั้งใจให้ hook ที่องค์กรเขียนไว้กับ Claude Code ย้ายมาได้ ครั้งแรกที่คู่ (event, command) ใดถูกเรียก Hermes จะขอ consent และบันทึกใน ~/.hermes/shell-hooks-allowlist.json (ข้ามได้ด้วย --accept-hooks หรือ HERMES_ACCEPT_HOOKS=1 สำหรับเครื่อง headless)
hermes hooks list # shell hooks + outbound targets
hermes hooks test pre_tool_call # ยิง event จำลอง — ทำทุกครั้งหลังแก้ config
hermes hooks doctor
hermes hooks revoke # ถอน consent ที่เคยให้
นอกจาก plugin-level hook 26 ตัวที่เอกสาร plugins ระบุ (กลุ่ม directive: pre_tool_call, pre_llm_call, pre_verify, pre_gateway_dispatch; กลุ่ม transform; และ observer อย่าง on_session_start/end, subagent_start/stop, kanban_task_*) ยังมี gateway-only event แยกอีกชุด — gateway:startup, session:start/end/reset/compress, agent:start/step/end, reaction:added/removed, command:* hook ที่มีขอบเขตเวลาจะ timeout ตาม plugins.hook_callback_timeout (default 30 วินาที สูงสุด 600) โดย pre_tool_call เป็นตัวเดียวที่ fail closed เมื่อ timeout ตัวอื่น fail open
3) Deliverable Mode — ไฟล์ที่ agent ทำเสร็จ ไปถึงมือคนใน chat
ประตูขาออกแบบสุดท้ายไม่ใช่ event แต่เป็นไฟล์ — ใน Deliverable Mode gateway จะสแกนคำตอบของ agent หา path แบบ absolute (/tmp/report.pdf) หรือ home-relative (~/out/chart.png) ที่มีนามสกุลที่รองรับ แล้วอัปโหลดเป็นไฟล์แนบจริงบน Slack/Discord/Telegram/WhatsApp/Signal: รูปแสดง inline, เสียงเป็น voice note, ที่เหลือเป็นไฟล์ path ที่อยู่ใน code block ถูกข้าม และไฟล์ source อย่าง .py ถูกตัดออกโดยตั้งใจ worker ของ Kanban ก็แนบ deliverable ไปกับ kanban_complete ได้ เอกสารเตือนว่า "the agent doesn't reach for artifacts by default" — ถ้าอยากให้ cron job ทุกเช้าส่ง PDF แทนข้อความยาว ต้องบอกใน prompt หรือใส่ไว้ใน AGENTS.md/SOUL.md
💡 คำว่า "artifact" ใน Hermes หมายถึงสามสิ่งที่ไม่เกี่ยวกัน: Deliverable Mode ข้างบนนี้, แผงแสดงผลด้านขวาใน Desktop app (PR #72345, v0.20.0 — ฟีเจอร์ UI ล้วน ๆ เก็บใน localStorage ของเครื่องนั้น) และคำว่า "task artifact" ในโปรโตคอล A2A อย่าคาดหวังว่าจะมี "artifact store" กลางให้ pipeline ดึงใช้ — วันนี้ยังไม่มี และ issue #19489 ที่ขอ artifact builder skill ยังเปิดอยู่
Blueprints — Skill ที่มีตารางเวลา
ถ้า cron คือเครื่องยนต์ Automation Blueprints (v0.17.0, 19 มิถุนายน 2026, PR #41309) คือหน้าปัดที่ทำให้คนที่ไม่อยากเห็น cron syntax ใช้มันได้ — release note สรุปว่า "Pick an automation by name and Hermes asks you for what it needs — no cron syntax" แต่สิ่งที่ผมอยากให้จำจริง ๆ คือประโยคใน tools/blueprints.py:
"A blueprint is NOT a new object type. It is an ordinary skill that additionally declares an automation schedule in its frontmatter."
นั่นแปลว่าทุกอย่างที่คุณรู้จาก #6 Skills ใช้ได้ทันที — SKILL.md เดิม เพิ่มบล็อก metadata.hermes.blueprint ก็กลายเป็น blueprint ที่ไหลผ่าน pipeline เดียวกับ skills hub และตามตรรกะของเอกสาร กลไกทุกอย่างที่ครอบ skill (รวมถึงการสแกนตอนติดตั้งที่เล่าไว้ในตอน #6) ก็ครอบ blueprint ด้วย:
---
name: nightly-backlog-triage
description: จัดลำดับ issue ค้างทุกคืน แล้วสรุปเข้า Slack ของทีม
metadata:
hermes:
blueprint: # โครงตามฟิลด์ของ BlueprintSpec — ตรวจกับ tools/blueprints.py ก่อนใช้จริง
schedule: "0 2 * * *" # skill_name, schedule, deliver (default origin), prompt,
deliver: slack # no_agent, model, provider, enabled_toolsets
prompt: "รัน skill นี้กับ repo ของทีม"
enabled_toolsets: [web, file, terminal]
---
# Nightly Backlog Triage
1. ดึง issue ที่เปิดอยู่ทั้งหมดด้วย gh …
2. จัดกลุ่มตาม label และอายุ …
นิยามเดียวถูก render ออกมาสี่หน้าตา: ฟอร์มใน dashboard, คำสั่ง /blueprint ใน CLI/TUI และ messenger, บทสนทนาที่ agent ถามสิ่งที่ต้องรู้ทีละข้อ และรายการใน catalog ของเอกสาร หน้า Automation Blueprints สำหรับนักพัฒนาแสดงตัวอย่างที่องค์กรสาย software จะรู้สึกคุ้น:
| Blueprint | สิ่งที่ทำ |
|---|---|
| Nightly Backlog Triage | จัดกลุ่มและลำดับ issue ค้าง — ตัวอย่างในหน้า catalog ใช้ 0 2 * * * |
| Automatic PR Code Review | รีวิว PR ใหม่และ comment กลับ — จุดชนวนคือ webhook ไม่ใช่ cron |
| Docs Drift Detection | ตรวจว่าเอกสารตามโค้ดทันไหม |
| Dependency Security Audit | รายงานช่องโหว่ CVSS ≥ 7.0 |
| Issue Auto-Labeling | ติด label ให้ issue ใหม่ |
| CI Failure Analysis | วิเคราะห์สาเหตุ workflow ที่ล้ม |
| Auto-Port Changes Across Repos | ย้าย patch ข้าม repo |
| Security Audit Pipeline · Deploy Verification · Alert Triage · Uptime Monitor | ชุดงาน ops |
หน้าคู่มือ Automation Blueprints (hermes-agent.nousresearch.com/docs/guides/automation-blueprints) แสดง blueprint ทั้งหมด 17 ตัว ณ วันที่ 1 กันยายน 2026 — นอกจากในตารางแล้วยังมี Competitive Repository Scout, AI News Digest, Paper Digest, Stripe Payment Monitoring, Daily Revenue Summary และ Content Pipeline (ตัวเลขนี้อ้างอิงหน้าคู่มือเท่านั้น หน้า reference catalog เป็นแบบ client-rendered ผมนับจากมันไม่ได้) ข้อดีเชิงองค์กรที่ตามมาจากนิยาม "blueprint = skill": คุณจัดการมันด้วยวิธีเดียวกับ skill ทุกตัว — ผ่าน skills hub, ผ่านการ export profile — และเมื่อต้องดูแลหลายเครื่อง ตอน #10 Desktop & Fleet จะเล่าต่อ
Hermes เป็น Backend ให้แอปที่องค์กรมีอยู่แล้ว
ประตูที่ห้าคือการกลับด้าน — แทนที่ Hermes จะไปเรียกระบบอื่น ให้ระบบอื่นเรียก Hermes ตั้งแต่ v0.4.0 gateway มี OpenAI-compatible API server ในตัว (release note ระบุว่ามาพร้อม "input limits, field whitelists, SQLite-backed response persistence, and CORS origin protection" ตั้งแต่วันแรก):
# ~/.hermes/.env (หรือ gateway.api_server.* ใน config.yaml — env ชนะ)
API_SERVER_ENABLED=true
API_SERVER_KEY=<long-random-key> # บังคับ แม้ฟังแค่ loopback
# เริ่ม gateway ตามปกติ — API server ฟังที่ http://127.0.0.1:8642
hermes gateway
# ใช้ client ของ OpenAI ตัวไหนก็ได้
curl http://127.0.0.1:8642/v1/chat/completions \
-H "Authorization: Bearer $API_SERVER_KEY" \
-H "X-Hermes-Session-Key: helpdesk-user-4471" \
-H "Content-Type: application/json" \
-d '{"model":"hermes","messages":[{"role":"user","content":"สรุป ticket ที่ค้างของผม"}],"stream":true}'
# ชื่อ model ตั้งได้ด้วย gateway.api_server.model_name — ดูค่าจริงจาก GET /v1/models
| Endpoint | ลักษณะ |
|---|---|
POST /v1/chat/completions | stateless; stream ผ่าน SSE พร้อม event hermes.tool.progress ให้ UI แสดงว่า agent กำลังใช้ tool อะไร (v0.7.0) |
POST /v1/responses | ประวัติฝั่ง server ผ่าน previous_response_id หรือ conversation ที่ตั้งชื่อ; เก็บใน SQLite แบบ LRU 100 responses |
POST /v1/runs · GET /v1/runs/{id} · /events | งานยาวแบบ asynchronous พร้อม SSE; หยุดได้ด้วย POST /v1/runs/{id}/stop (v0.12.0) |
GET /v1/models · GET /v1/capabilities | ให้ client ตรวจว่า server นี้ทำอะไรได้ |
X-Hermes-Session-Id / X-Hermes-Session-Key | Id = transcript; Key (≤256 ตัวอักษร) = ขอบเขต memory ที่คงที่ต่อผู้ใช้ |
/p/<profile>/… | หลาย profile บน server เดียว แต่ละ profile มี key ของตัวเอง |
ข้อจำกัดที่ต้องรู้ก่อนออกแบบ: ไม่มี file upload, เพดาน concurrent run default 10 (max_concurrent_runs) และ toolset hermes-api-server ตัด clarify, TTS, computer_use และ send_message ออก — เพราะไม่มีมนุษย์อยู่ปลายสายให้ถามกลับ
แอปตัวแรกที่องค์กรมักเสียบคือ Open WebUI — และเอกสารของ Open WebUI เองมีหน้า "Connect an agent → Hermes Agent": base URL http://localhost:8642/v1 (จาก Docker ใช้ host.docker.internal และต้องมี /v1 ต่อท้าย) API key คือ API_SERVER_KEY และ gateway ต้องรันค้างไว้ หน้า Open WebUI ฝั่ง Hermes เสริมว่า tool รันบนเครื่องที่ API server อยู่ Open WebUI คุยแบบ server-to-server จึงไม่ต้องตั้ง CORS และถ้าอยากแยกผู้ใช้ ให้ใช้ profile คนละพอร์ต
hermes peer, A2A และสิ่งที่ hermes mcp serve ไม่ได้ทำ
สามสิ่งนี้มักถูกเข้าใจสลับกัน:
hermes peer(v0.21.0, PR #88725) ให้ Hermes สองตัวคุยกันข้าม host — แต่มันไม่ใช่ A2A: PR ระบุว่ามัน "uses the peer's existing api_server platform as the wire" (POST/api/sessions/{id}/chatด้วย bearer key) ตั้งค่า peer ใต้bot_peersเก็บ key เป็นHERMES_PEER_<NAME>_KEYใน.envแล้วใช้hermes peer add|dm|run --idempotency-key|status|stopคำตอบไปโผล่ใน Bot Chat ของแต่ละฝั่ง- A2A (plugin ที่มาพร้อมระบบ, A2A v1.0, PR #77109, v0.20.0) เปิดที่พอร์ต 9900 มี
/.well-known/agent-card.jsonและ JSON-RPC 2.0 (SendMessage, SendStreamingMessage, GetTask, ListTasks, CancelTask) — สำหรับคุยกับ agentต่างค่ายที่พูดมาตรฐานเดียวกัน toolset ฝั่ง client ปิดโดย default, bind 127.0.0.1 โดย default และการ bind ออกนอกต้องมี token มี audit log ที่~/.hermes/a2a_audit.jsonlเอกสารเองบอกว่างานบนเครื่องเดียวกันควรใช้ delegation หรือ Kanban (ตอน #2) แทน hermes mcp serve(v0.6.0, 30 มีนาคม 2026) ไม่ได้เปิด agent ให้ใครเรียก — docstring ของmcp_serve.pyบอกว่ามัน "expose messaging conversations as MCP tools": 10 tool ฝั่ง messaging bridge (conversations_list,messages_read,messages_send,events_wait,permissions_respondฯลฯ) ผ่าน stdio เท่านั้น ไม่มี tool "ask the agent" ถ้าคุณอยากให้ Claude Desktop อ่านและตอบข้อความในช่องทางที่ Hermes ดูแล นี่คือของสำหรับสิ่งนั้น แต่ถ้าอยากให้แอปเรียก Hermesคิด ให้กลับไปที่ 8642
# Claude Desktop config — เข้าถึง "กล่องข้อความ" ของ Hermes ไม่ใช่สมองของมัน (gateway ต้องรันอยู่)
{ "mcpServers": { "hermes-messaging": { "command": "hermes", "args": ["mcp", "serve"] } } }
hermes serve คือ backend แบบ headless ของ Desktop app ไม่ใช่วิธีเปิด API server — คำสั่งทางการคือ hermes gateway พร้อม API_SERVER_ENABLED=true
Hardening และการ Scale ระนาบ Automation
ทุกประตูที่เปิดในบทความนี้เคยมีคนเดินเข้ามาแบบไม่ได้รับเชิญ ใน #4 Security ผมเล่าบันทึก CVE ของ Hermes ไว้ในภาพรวม สี่รายการที่เป็นของหัวข้อนี้โดยตรง:
| CVE | องค์ประกอบ | สาระ |
|---|---|---|
| CVE-2026-7113 GHSA-mq9g-m88j-mmvx · 27 เม.ย. 2026 · Moderate · CWE-287 | gateway/platforms/webhook.py | อาร์กิวเมนต์ _INSECURE_NO_AUTH ปิดการตรวจลายเซ็นบน /webhooks/{route} ได้ — issue #6440 ของชุมชน (9 เม.ย.) บรรยายไว้ก่อนแล้ว รายการ GHSA เป็น mirror ที่ GitHub รีวิวแล้วของ record จาก VulDB (CNA) ไม่ใช่การวิเคราะห์อิสระ และระบุ "Patched versions: Unknown" |
| CVE-2026-7112 GHSA-r7hr-pvjh-r4p3 · 27 เม.ย. 2026 · Moderate | gateway/platforms/api_server.py — ฟังก์ชัน _check_auth | improper authentication บนตัวจัดการ API_SERVER_KEY — GHSA เป็น mirror ของ VulDB เช่นกัน และระบุ "Patched versions: Unknown" |
| CVE-2026-14628 NVD · 4 ก.ค. 2026 · CVSS 3.1 5.3 (Medium) | gateway/platforms/base.py — ฟังก์ชัน extract_media (องค์ประกอบ "Live Webhook Endpoint") | path traversal; ระบุ hermes-agent ถึง 2026.5.16 |
| CVE-2026-14626 NVD · 4 ก.ค. 2026 · CVSS 3.1 4.3 (Medium) | run_agent.py — AIAgent.run_conversation ผ่าน HTTP API | denial of service ผ่านอาร์กิวเมนต์ todos; ระบุ hermes-agent ถึง 2026.4.30 |
สองรายการแรกระบุ hermes-agent 0.8.0 เรื่อง INSECURE_NO_AUTH ต้องเล่าให้ครบ: หน้า webhooks ในเอกสารทางการเองอนุญาตให้ใช้มันสำหรับทดสอบบน loopback — มันคือสวิตช์ทดสอบที่มีเอกสารรองรับ ไม่ใช่ backdoor ที่ซ่อนไว้ — สิ่งที่ CVE บันทึกคือสวิตช์นั้นปิดการตรวจลายเซ็นทั้ง route ได้ และคู่มือ PR-review ก็มีบรรทัดที่ควรพิมพ์ติดผนัง: "Never use INSECURE_NO_AUTH in production" อีกสามรายการบนพื้นผิวเดียวกันตอนนี้อยู่บน NVD แล้ว (ตรวจ 1 กันยายน 2026): CVE-2026-10224 (เผยแพร่ 1 มิ.ย. 2026, CVSS 3.1 5.3 — _handle_webhook_request ใน gateway/platforms/feishu.py นำไปสู่ resource consumption โจมตีจากระยะไกลได้ ระบุ 2026.4.0–2026.4.30; แก้ด้วย "webhook body-cap sweep" PR #59215), CVE-2026-53869 (17 มิ.ย. 2026, CVSS 3.1 7.5 HIGH — DNS rebinding บน WebSocket endpoint ข้ามการตรวจ Host/Origin ได้, <0.16.0) และ CVE-2026-53870 (17 มิ.ย. 2026, CVSS 3.1 5.5 — response_store.db และ webhook_subscriptions.json ถูกสร้างด้วยสิทธิ์ 0o644 เปิดเผยประวัติสนทนาและ HMAC secret ให้ผู้ใช้ในเครื่อง, <0.16.0) หลังจากนั้น ระหว่าง 3 กรกฎาคมถึง 6 สิงหาคม ยังมีรายการความรุนแรงต่ำจาก VulDB ทยอยออกมาอีกสิบเอ็ดรายการบนพื้นผิวอื่นของ Hermes — บันทึกรวมอยู่ใน #4 Security
เช็กลิสต์ hardening ที่ผมกลั่นจากเอกสารทั้งหมดในบทความนี้:
- ☐ อยู่บน v0.21.0 หรือใหม่กว่า — ทุกรายการข้างบนระบุเวอร์ชันก่อน 0.16.0 ทั้งสิ้น
- ☐ พอร์ต 8644 อยู่หลัง reverse proxy ที่ทำ TLS; ทุก route มี secret ยาวสุ่ม; ใช้ Generic V2 ที่มี timestamp ไม่ใช่ V1
- ☐ ไม่แตะ
toolsetsของ route ถ้าไม่จำเป็น — สี่ tool ปลอดภัยคือ default ที่ถูกแล้ว - ☐ 8642 และ 9900 ฟังที่ 127.0.0.1 เว้นแต่มีเหตุผล และมี key/token เสมอ
- ☐ shell hook ที่ทำหน้าที่ guard ตั้ง
fail_closed: true— default fail-open หมายความว่า script พัง = ไม่มีการ์ด - ☐
cron.allow_agent_schedulingยังเป็น false และpreflightยังเป็น true - ☐ ทดสอบ hook จริงหลังทุกการอัปเดต — ข้อนี้มีที่มา
ประวัติเรื่องความน่าเชื่อถือของ hook ไม่ราบรื่นนัก: issue #2817 (hook ไม่ยิง) มาก่อน release note v0.5.0 (28 มีนาคม 2026) ที่ประกาศว่า lifecycle hooks "now fire" และ issue #44582 (เปิด 12 มิถุนายน 2026 บน 0.16.0: "pre_tool_call plugin hook not invoked during agent tool execution") ถูกปิดแบบ not planned โดยไม่มีคำตอบจาก maintainer ที่มองเห็นได้ ผมไม่ได้บอกว่า hook ไม่ทำงานวันนี้ — ผมบอกว่าอย่าเชื่อจนกว่าจะเห็นเอง:
# หลัง hermes update ทุกครั้ง — ให้เป็นส่วนหนึ่งของ runbook
hermes hooks test pre_tool_call
hermes hooks test on_session_end
hermes plugins doctor --ci # exit code ใช้ใน CI ได้
hermes cron doctor
Chronos — cron สำหรับ gateway ที่หลับได้
ปัญหาเชิงสถาปัตยกรรมข้อเดียวของ cron ในตัวคือ มันต้องการ gateway ที่ตื่นตลอดเวลาเพื่อ tick ทุก 60 วินาที — ซึ่งขัดกับ hosting แบบ scale-to-zero ที่ตอน #7 พูดถึง v0.17.0 (PR #48275) จึงทำให้ scheduler ถอดเปลี่ยนได้และเพิ่ม provider ชื่อ Chronos ตาม wire spec ในเอกสาร docs/chronos-managed-cron-contract.md ของ repo: แทนที่จะ tick เอง gateway ลงทะเบียน job กับบริการภายนอกที่ตั้ง one-shot timer ให้ เมื่อถึงเวลา บริการนั้นเรียกกลับมาที่ /api/cron/fire พร้อม JWT อายุสั้น (ราว 60–120 วินาที, aud=agent:{instance_id}, purpose=cron_fire) gateway ตอบ 202 {"status":"accepted"} แล้วรัน job ใน background — at-most-once ผ่าน compare-and-set ฝั่ง store
# config.yaml — เปิดใช้ Chronos; ถ้าไม่ตั้ง ระบบใช้ ticker ในตัวตามเดิม
cron:
provider: chronos
chronos:
portal_url: …
callback_url: …
expected_audience: …
nas_jwks_url: …
นี่คือกลไกสำหรับ gateway แบบ hosted ที่ scale to zero — แบบที่ Hermes Cloud โฆษณาไว้ — ให้ยังตั้ง cron ได้โดยไม่ต้องจ่ายค่าเครื่องตอน idle ผมจะไม่พิมพ์ราคาหรือขนาดเครื่องของ Cloud เพราะไม่มีตัวเลขทางการ สำหรับ self-host บน VPS ตัวเดียว ticker ในตัวก็เพียงพอ และ Chronos คือทางเลือกเมื่อคุณย้ายไป hosting ที่คิดเงินตามเวลาทำงานจริง
ปิดท้ายด้วยข้อสังเกตสำหรับคนที่ย้ายจาก OpenClaw: ตามที่เล่าไว้ในตอน #7 hermes claw migrate ไม่ได้พา cron job และ hook ข้ามมาด้วย — ซึ่งผมมองว่าเป็นโอกาสมากกว่าภาระ heartbeat แต่ละตัวในระบบเดิมควรถูกถามใหม่ว่ามันคือ watchdog (ชั้นที่ 1), monitor (ชั้นที่ 2) หรืองานที่ต้องคิดจริง (cron ปกติ) และผมพนันว่าส่วนใหญ่จะตกอยู่ในสองกลุ่มแรก
🎯 สิ่งสำคัญที่ต้องจำ
- ห้าประตู = cron · inbound webhook · hooks/outbound · API server · A2A/peer — เริ่มจาก cron แล้วค่อยขยาย
- Cron = ไฟล์ธรรมดา =
jobs.json+executions.db+output/*.md; gateway tick ทุก 60 วินาที; agent ตั้งเวลา agent ไม่ได้โดย default - Zero-token สามชั้น =
--no-agent(v0.13.0) · monitor-mode ที่ hash เท่าเดิมไม่เรียก LLM (v0.21.0) ·deliver_only/hermes send - Cron ที่จำได้ = memory + continuity + notepad (v0.21.0) และ
context_from(v0.12.0) ต่อ job เป็น chain - Ops loop =
hermes cron incidents / doctor / runs— error ถูก redact, exit code ใช้ใน monitoring ได้, model-drift guard กันรันด้วย model ที่ไม่ได้ทดสอบ - Webhook = 8644, secret ทุก route, ±300 s replay window, สี่ tool ปลอดภัย — "HMAC authenticates the sender, not the content"
- ขาออก =
hooks.outboundลงลายเซ็น HMAC (v0.20.0), shell hook veto ด้วย exit 2 (v0.11.0), Deliverable Mode แนบไฟล์จริง - Blueprint = skill ธรรมดาที่มี schedule ใน frontmatter — ทุกอย่างจากตอน #6 ใช้ได้ หน้าคู่มือแสดง 17 ตัว ณ 1 ก.ย. 2026
- 8642 = OpenAI-compatible API;
hermes mcp serveคือ messaging bridge ไม่ใช่ agent;hermes serveไม่ใช่คำสั่งเปิด API - Hardening = CVE-2026-7113/7112 บน 0.8.0 (GHSA = mirror ของ VulDB, patch ไม่ระบุ) และ 14628/14626 บน NVD ก.ค. 2026; ทดสอบ hook ด้วย
hermes hooks testหลังทุก update; Chronos สำหรับ gateway ที่ scale to zero
Picture a Monday morning. Before anyone on the team has opened a laptop, the pull requests that sat unreviewed all weekend have a summary waiting in Slack, the dependency that acquired a new CVE overnight already has a ticket, and nobody has typed an instruction since Friday evening. The question this post answers is not whether an agent can do that — it is which door the work should walk through: a schedule, an external event, or another application calling in. And, just as important for whoever pays the bill: which of those rounds never need to spend a token at all.
The introduction to cron is already in #2 Agent Teams, under "Always-On Work" — the cronjob actions, the schedule dialects, fresh isolated sessions, continuity and context_from, the delivery targets, [SILENT], no_agent, and v0.21.0's "cron jobs that remember". In #5 Integrations the API server and A2A each got a brief section. This post does not re-introduce any of that; it is the ops layer on top: the files and ledger behind the built-in cron of Hermes Agent, the defaults worth knowing, the incidents/doctor/runs loop; inbound webhooks that verify a signature before they believe anything; hooks and outbound webhooks that push events to systems your organization already owns; blueprints, which turn out to be nothing more than a skill with a schedule; and an OpenAI-compatible API server that lets internal apps call Hermes the way they would call any model.
Readers who followed this blog's OpenClaw series will recognise the heartbeat pattern from OpenClaw #5: Integrations — wake the agent periodically and let it check whether anything needs doing. The idea survives in Hermes, but it has been broken into named, inspectable pieces, and the most important of those pieces can leave the model switched off when nothing has changed. I will keep that comparison qualitative. No number in this post is my own measurement; every version, default and quotation comes from the official documentation or the GitHub release notes of Nous Research, as of 1 September 2026 (v0.21.0, tag v2026.8.31).
Five Doors Into the Gateway
The first thing to internalise is that all of Hermes's automation lives inside the same gateway daemon that receives your Telegram and Slack messages — the same VPS or Docker container from #7 Production. There is no separate scheduler process and no worker pool. Work reaches that one daemon through five doors:
| Door | Mechanism | Best for | Since |
|---|---|---|---|
| Cron | the gateway ticks every 60 seconds and reads ~/.hermes/cron/jobs.json | recurring reports, watchdogs, anything on a rhythm | present in the earliest release visible on GitHub (v0.2.0, 12 March 2026) |
| Inbound webhook | HTTP POST to port 8644 with an HMAC signature | GitHub/GitLab events, monitoring stacks, n8n | v0.4.0 (March 2026) |
| Hooks / outbound webhooks | shell scripts at lifecycle moments + signed outbound HTTP POSTs | vetoing dangerous commands, shipping audit events to a SIEM | shell hooks v0.11.0 (April 2026), outbound v0.20.0 (August 2026) |
| API server | OpenAI-compatible endpoint on port 8642 | Open WebUI, internal apps, hermes peer | v0.4.0 |
| A2A / peer | JSON-RPC per the A2A v1.0 standard on port 9900 / peer over the API server | agents on different hosts talking to each other | A2A v0.20.0, hermes peer v0.21.0 |
My advice for an organization starting out: begin with cron. It is the only door where you set the tempo, nothing external can knock, and since v0.13.0 it has a mode that never touches the model. Add inbound webhooks when you want GitHub or your monitoring stack to do the waking. Leave the API server and A2A for last — both open network surfaces with a CVE history of their own, which I keep for the final section.
💡 If you have read #2, hold the distinction this way: a heartbeat is an agent that wakes up and asks whether anything has happened; a Hermes cron job is a named unit of work with a schedule and a ledger of its own — and, as the next two sections show, it can separate "waking up" from "thinking".
Anatomy of a Cron Job
Hermes cron is not the Linux crontab firing bare commands. A job is a prompt wrapped in metadata (or a script wrapped in the same metadata), and every piece of state is a plain file you can open yourself:
~/.hermes/cron/
├── jobs.json # every job definition — written atomically
├── executions.db # the ledger: claimed → running → completed | failed | unknown
├── .tick.lock # file lock so two ticks never overlap
└── output/
└── {job_id}/
└── {timestamp}.md # every run's output, readable after the fact
The 60-second tick from #2 has two operational consequences: one minute is the finest schedule resolution you get, and .tick.lock is what stops two processes (say, two gateways accidentally mounting the same volume) from running the same job twice.
There are three ways to create a job — the CLI (hermes cron create), the cronjob tool the agent calls for itself inside a session, and the dashboard on port 9119 (backed since v0.4.0 by a /api/jobs REST API). I show the tool form because it is the parameter list the documentation spells out most completely; in chat you simply describe the job you want and the agent calls this. For the CLI flags, hermes cron create --help is the source of truth:
# one tool for every action — release v0.3.0 (17 March 2026) folded several commands into it
cronjob(
action="create", # list | update | pause | resume | run | remove
schedule="0 9 * * 1-5", # five-field cron — weekday mornings
name="morning-brief",
prompt="Summarise the issues opened in the team repo yesterday, grouped by label",
deliver="slack",
skills=["github-triage"], # attach a skill from post #6
workdir="/srv/repo", # scopes file/terminal tools and loads that repo's AGENTS.md (absolute paths only)
enabled_toolsets=["web", "file", "terminal"],
continuity=True, # v0.21.0 — remember this job's own previous output
context_from="deps-scan", # v0.12.0 — prepend another job's last successful output
reasoning_effort="high", # v0.21.0 — pinned per job
no_agent=False,
attach_to_session=False,
)
The schedule dialects are covered in #2 (the feature page lists five: relative, interval, natural language, five-field cron, and ISO timestamps). One thing worth adding: the cron feature page lists natural-language forms such as daily at 7am, and #2 uses weekdays at 9am as its example, but the automate-with-cron guide, another official page, states that natural language such as "daily at 9am" is not supported. I have not tested which page matches the code. In production I would standardise on five-field cron, and if you want natural language, test it on your own build (v0.21.0) before you depend on it.
The cron: block of config.yaml carries defaults an organization should know before touching them:
| Key | Default | Meaning |
|---|---|---|
script_timeout_seconds | 3600 | a script may run for one hour (override with HERMES_CRON_SCRIPT_TIMEOUT) |
failure_nudge_threshold | 3 | consecutive failures before you get nudged |
misfire_grace_minutes | 10 | if the gateway was down and returns within ten minutes, missed jobs still run |
model_drift_guard | true | an unpinned job snapshots provider/model at creation; if the global default later changes, the job skips its run with a one-time alert |
allow_agent_scheduling | false | a session started by cron cannot create further cron jobs |
preflight | true | checks the API key, the attached skills' env and commands, and the delivery targets before dispatch |
mirror_delivery | false | when true, every run lands as a thread you can continue |
max_parallel_jobs | (unset) | cap on concurrently running jobs |
cleanup_timeout_seconds | 10 | post-run cleanup timeout |
media_send_timeout_seconds | 300 | timeout for sending media to a delivery target |
bot_chat_delivery_timeout_seconds | 600 | timeout for delivering a result into Bot Chat |
allow_agent_scheduling: false wall that #2 explains, every prompt is scanned at create/update time for prompt injection, invisible Unicode, SSH backdoors and credential-exfiltration patterns — a suspicious prompt is blocked before it ever runs. That is a direct lesson from the malicious-skill story I told in #4 Security.
On delivery: #2 gave a partial list; the full one on the feature page is origin, local, telegram, discord, slack, whatsapp, signal, matrix, mattermost, email, sms, homeassistant, dingtalk, feishu, wecom, weixin, bluebubbles, qqbot, bot-chat and all. You can target a specific chat and thread (telegram:-100123:17585) or fan out to several targets separated by commas. The [SILENT] rule is in #2; two things to add: a suppressed result is still saved under output/, and a failed job always delivers, however quiet you asked it to be.
As for what a job may use: toolsets.py defines a hermes-cron toolset that mirrors hermes-cli, then per-job enabled_toolsets and the cron platform entry in hermes tools override it in that order. And workdir does more than fence the file and terminal tools — it also loads AGENTS.md, CLAUDE.md or .cursorrules from that directory, so a job working on a repository inherits that repository's rules without you pasting them into the prompt.
Zero-Token Automation: Work That Never Wakes the Model
This is where Hermes cron most clearly parts company with the OpenClaw-era heartbeat. In the heartbeat pattern every wake-up is an LLM call, even when nothing has changed. Hermes offers three layers that remove that cost.
Layer 1 — script-only jobs (--no-agent, v0.13.0)
#2 gave no_agent=True one sentence; this is its full contract. Release v0.13.0, "The Tenacity Release" (7 May 2026), added watchdog mode: a job that runs an ordinary script with no agent in the loop at all.
# scripts must live in $HERMES_HOME/scripts/ — path traversal is rejected
cat > ~/.hermes/scripts/watchdog.sh <<'EOF'
#!/usr/bin/env bash
# the shebang is ignored — .sh/.bash always run under bash, anything else under the current Python
code=$(curl -s -o /dev/null -w '%{http_code}' https://intranet.example.ac.th/health)
if [ "$code" != "200" ]; then
echo "⚠️ intranet health = $code" # stdout present → this text is delivered as-is
exit 0
fi
# exit 0 + empty stdout = a silent tick: nothing sent, no tokens spent
EOF
hermes cron create "every 5m" --no-agent --script watchdog.sh --deliver telegram
The watchdog contract fits in four lines: exit 0 with empty stdout is silent; any stdout is delivered verbatim; a non-zero exit or a timeout raises an error alert; and if the last line of stdout is {"wakeAgent": false} the tick is treated as silent too. The detail I like most is that the subprocess environment is scrubbed of provider credentials — a script you wrote (or one the agent wrote for you) does not walk away with your model API keys.
Layer 2 — monitor mode that skips the LLM when nothing changed (v0.21.0)
v0.21.0 "Pantheon" (31 August 2026) added hash-suppressed change detection (PR #81138): a job in monitor mode hashes what it is watching, compares it with the previous run, and if nothing changed it never calls the model. This is the direct answer to the complaint I made in the earlier series about heartbeats that pay for the privilege of reporting that nothing happened. A sourcing caveat: as of this writing, monitor mode appears in the v0.21.0 release notes and PR #81138 only — the rendered cron documentation does not describe it yet. So I have found no percentage for what it saves and will not invent one, but the logic is plain: an unchanged round costs zero tokens.
Layer 3 — deliver_only and hermes send
The webhook side has the same idea: a route with deliver_only: true (v0.11.0, PR #12473) forwards the transformed payload to a channel without involving the agent. And for any script on the box there is hermes send, from v0.15.0 "The Velocity Release" (28 May 2026, PR #27188), which the release note describes as "pipe any script's output to any messaging platform":
# one-shot delivery to a configured channel — no agent loop, no tokens
hermes send --to slack "Backup finished: 2.1 GB"
hermes send --to telegram --file /var/log/nightly-report.pdf
hermes send --list # which targets are available
# bolt it onto the scripts you already run
./run-etl.sh 2>&1 | tail -n 5 | xargs -0 hermes send --to slack
A detail that matters for scripts on a server: for the bot-token platforms — Telegram, Discord, Slack, Signal, SMS and the WhatsApp Cloud API — hermes send delivers on its own, with no running gateway required; plugin platforms still need one.
Cron That Remembers — and the Ops Loop
All three bridges across runs are already in #2 — context_from (v0.12.0, 30 April 2026, PR #15606), which prepends the most recent successful output of another job; continuity=True, which injects the same job's previous output; and v0.21.0's persistent memory and per-job notepad. Two things #2 did not say: the memory a cron agent loads is the same bounded MEMORY.md whose rules and caps are in #3 Memory (so what a job learns at 2 a.m. is there for your session at 9), and, as of this writing, cron memory and the notepad rest on the v0.21.0 release notes and PR #81138 only — the rendered cron page still documents just continuity and context_from. Here is what the chain looks like in practice:
# a two-step chain: scan at 01:00, let the 02:00 job read that result
cronjob(action="create", name="deps-scan", schedule="0 1 * * *",
prompt="Run pip-audit and npm audit in the repo; report only CVSS ≥ 7",
workdir="/srv/repo", deliver="local")
cronjob(action="create", name="deps-triage", schedule="0 2 * * *",
context_from="deps-scan", # deps-scan's last successful output is prepended
continuity=True, # and it remembers which issues it opened last time
prompt="From the attached scan, open issues for the team; skip anything already filed",
deliver="slack")
Two more v0.21.0 additions ops people will appreciate: output can land in a bot's own Bot Chat (the Bot Mode concept lives in #2 — I will not re-explain it), and a "Trigger now" button (#70638) in the Desktop app and the dashboard runs a job immediately for testing without touching its schedule — on the CLI the same verb is hermes cron run <job>.
The ops loop: incidents → doctor → runs
A cron system that runs while nobody is watching must answer "did anything break last night?" in three commands:
# 1) what the system noticed — detected → alerted → closed
hermes cron incidents
hermes cron incidents ack <incident-id> # v0.21.0: an acked failure signature stops re-pinging (#95017)
# 2) read-only health check — exits 1 on problems, so it slots straight into monitoring
hermes cron doctor
# checks: last run failed / delivery failed / stale next_run_at / missing script / missing workdir
# 3) run history
hermes cron runs deps-triage --limit 20
Details I give credit for: incident error text is redacted before storage, because stack traces drag credentials along; hermes cron doctor is designed so its exit code means something, which is exactly what the monitoring section of #7 wants. And the model-drift guard is a piece of care I have not seen in other agents' cron systems: change the global default with hermes model and every unpinned job skips its run with a one-time alert, instead of silently executing with a model you never tested against that prompt (resolution order is per-job pin → cron.model → the global default).
For output you want to keep talking to, set cron.mirror_delivery: true globally or attach_to_session=True per job; on thread-capable platforms (Telegram topics, Discord and Slack threads) each run opens its own thread you can reply in as if you were chatting with the agent. Slack's in-channel mode additionally needs slack.cron_continuable_surface: in_channel with reply_in_thread: false.
Inbound Webhooks Done Properly
The webhook adapter arrived in v0.4.0 (tag v2026.3.23, published 24 March 2026) in the same batch as Signal, DingTalk, SMS, Mattermost and Matrix. Hermes treats it as one more "platform", like Telegram — except the sender is a machine rather than a person. The ground rules fit in one table:
| Item | Value |
|---|---|
| Default port | 8644 (env WEBHOOK_PORT; enabled with WEBHOOK_ENABLED) |
| Endpoints | POST /webhooks/<route-name> and GET /health |
| Signatures accepted | GitHub X-Hub-Signature-256 · GitLab X-Gitlab-Token · Generic V2 X-Webhook-Signature-V2 + X-Webhook-Timestamp (HMAC-SHA256 over <timestamp>.<body>, accepted within ±300 seconds) · Generic V1 is deprecated and has no replay protection |
| Secret | required on every route — no exceptions |
| Rate limit | 30 requests/minute per route |
| Body size | 1 MB |
| Idempotency | a 1-hour cache keyed on X-GitHub-Delivery / X-Request-ID |
| Response codes | 200 · 400 · 401 · 404 · 413 · 429 · 502 |
Static routes live in config.yaml under platforms.webhook.extra.routes.<name>. This is the skeleton of the PR-review recipe the official docs publish:
# config.yaml — shape follows the official webhook-github-pr-review guide
platforms:
webhook:
enabled: true
extra:
port: 8644
rate_limit: 30
routes:
pr-review:
secret: "<long-random-secret>" # the same value goes into GitHub's webhook settings
events: [pull_request]
prompt: |
Review PR "{pull_request.title}" in {repository.full_name}.
The payload carries no diff — run: gh pr diff {number} --repo {repository.full_name}
Summarise risks and suggestions in at most 10 points.
deliver: github_comment
alerts:
secret: "<another-secret>"
deliver_only: true # skips the agent — zero tokens
deliver: telegram
What the template can do: dot-notation into the payload ({pull_request.title}), and {__raw__}, which dumps the whole payload as JSON truncated at 4,000 characters. Each route may also carry filters (exists / missing / equals / not_equals / contains / in / in_file / regex, composed with all / any / not) to screen events before they reach the agent, and a script transform from ~/.hermes/scripts/ — JSON on stdout replaces the payload, while empty stdout, [SILENT], or {"__hermes_ignore__": true} returns 200 and quietly drops the event. Webhook deliver adds log and github_comment to the usual channel list.
web_search, web_extract, vision_analyze and clarify and nothing else (_HERMES_WEBHOOK_SAFE_TOOLS in toolsets.py states the reason plainly: "to prevent prompt injection attacks"). No terminal, no file access. The docs compress the principle into one sentence worth memorising — "HMAC validation authenticates the sender, not the content." The signature proves GitHub sent it; anyone can write the PR title. The official PR-review guide repeats the warning — "PR titles and descriptions are attacker-controlled" — and tells you to run the gateway in Docker or a VM when it is internet-exposed.
If a route genuinely needs more — the recipe above asks the agent to run gh pr diff, which is outside those four tools — the route's toolsets key replaces the platform default wholesale; it does not merge. Open only what the job needs, and know that you are lowering a wall when you do. On the GitHub side you point the webhook at https://<host>/webhooks/pr-review, choose JSON as the content type, paste the same secret, and select the Pull requests event. The guide says the comment lands on the PR in roughly 30–90 seconds, and duplicate deliveries from GitHub are deduplicated for one hour.
Subscribing without a restart
Beyond static routes there are dynamic subscriptions that take effect immediately:
hermes webhook subscribe ci-failed \
--events workflow_run \
--prompt "Read the failed workflow's log and summarise the likely cause" \
--deliver telegram --deliver-chat-id -100123 \
--skills ci-triage
hermes webhook list
hermes webhook test ci-failed --payload ./sample.json # fire a sample payload at the route
hermes webhook remove ci-failed
# stored in ~/.hermes/webhook_subscriptions.json
# on a name collision, the static config.yaml route wins
The docs are explicit: "No gateway restart required — subscribe and it's immediately live." One thing I want you to notice: the agent itself has no webhook tool. When you say in chat "set up a webhook for GitLab events", it uses the terminal tool to run the hermes webhook subscribe command above, guided by a bundled skill called webhook-subscriptions (the SKILL.md anatomy is in #6 Skills). Which means a session without the terminal tool cannot do it — correct behaviour, for a session that was itself woken by a webhook.
A small market has already formed around port 8644 — Hookdeck publishes a guide to putting its own relay in front of Hermes, and fast.io a guide to having n8n call hermes webhook subscribe with an HMAC-computing node. Both are vendor marketing; I cite them only as evidence that organizations are wiring their existing workflow tools in front of this door, not as recommendations.
Pushing Events Out: Outbound Webhooks, Shell Hooks, and Deliverable Mode
The outbound side has three doors that are entirely different from one another, and I want the difference to be sharp, because the docs call several of them "hooks".
1) Signed outbound webhooks (v0.20.0)
PR #69406 (merged 2 August 2026, shipped in v0.20.0 "Herald" the next day) added hooks.outbound: whenever a lifecycle event fires, the gateway POSTs to a URL you choose, with an HMAC. This is how the observability section of #7 gets lifecycle events without scraping logs:
# config.yaml
hooks:
outbound:
- name: siem
url: https://siem.example.ac.th/hermes
events: [on_session_end, subagent_stop, post_tool_call] # any event from the plugin-hook catalog
secret_env: HERMES_SIEM_SECRET # read from ~/.hermes/.env — preferred over an inline secret:
timeout: 10 # 1–60 seconds
matcher: "^(terminal|write_file)$" # regex to narrow tool-scoped events to the tools you care about
# what the receiver sees
POST /hermes HTTP/1.1
Content-Type: application/json
X-Hermes-Event: post_tool_call
X-Hermes-Delivery: <delivery-id>
X-Hermes-Signature-256: sha256=<hex of HMAC-SHA256 over the raw body>
{ "hook_event_name": "post_tool_call", "tool_name": "terminal", "session_id": "…",
"delivery_id": "…", "timestamp": "…", … } # same shape as shell-hook stdin, plus the last two fields
Delivery is fire-and-forget through a bounded queue and a single daemon thread: one retry with backoff on a connection error or 5xx, no retry on 4xx, 3xx not followed, and edits to the target list take effect on the next CLI session or gateway restart. Put bluntly, this is not a message queue with delivery guarantees. If your organization needs at-least-once, put a relay that owns that responsibility in between. hermes hooks list shows outbound targets alongside shell hooks.
2) Shell hooks that can veto (v0.11.0)
Shell hooks (v0.11.0, 23 April 2026, PR #13296) are a different animal from webhooks: local scripts that run at moments in the agent's loop, and at some of those moments can block the action:
# config.yaml
hooks:
pre_tool_call:
- matcher: "^terminal$"
command: ~/.hermes/agent-hooks/guard-terminal.sh
timeout: 30 # default 60, maximum 300
fail_closed: true # if the script breaks, block (the default is fail-open!)
#!/usr/bin/env bash
# ~/.hermes/agent-hooks/guard-terminal.sh — reads JSON on stdin
payload=$(cat)
cmd=$(printf '%s' "$payload" | jq -r '.tool_input.command // ""')
if printf '%s' "$cmd" | grep -Eq 'rm -rf /|mkfs|DROP TABLE'; then
echo '{"action":"block","reason":"dangerous command — needs a human"}'
exit 2 # exit code 2 also blocks
fi
exit 0
The stdin fields are hook_event_name, tool_name, tool_input, session_id, cwd and extra. A pre_tool_call hook may answer with {"action":"block"|"modify"} or with the Claude Code form {"decision":"block"} — deliberately, so hooks an organization wrote for Claude Code can move across. The first time any (event, command) pair fires, Hermes asks for consent and records it in ~/.hermes/shell-hooks-allowlist.json (bypass with --accept-hooks or HERMES_ACCEPT_HOOKS=1 on headless machines).
hermes hooks list # shell hooks + outbound targets
hermes hooks test pre_tool_call # fire a synthetic event — do this after every config change
hermes hooks doctor
hermes hooks revoke # withdraw a consent you gave earlier
Besides the 26 plugin-level hooks the plugins doc lists (directive: pre_tool_call, pre_llm_call, pre_verify, pre_gateway_dispatch; the transform family; and observers such as on_session_start/end, subagent_start/stop, kanban_task_*), there is a separate set of gateway-only events — gateway:startup, session:start/end/reset/compress, agent:start/step/end, reaction:added/removed, command:*. Bounded hooks time out after plugins.hook_callback_timeout (30 seconds by default, 600 at most), and pre_tool_call is the only one that fails closed on timeout; the rest fail open.
3) Deliverable Mode — the files an agent produces reach the people in the chat
The last outbound door carries files, not events. In Deliverable Mode the gateway scans the agent's response for absolute (/tmp/report.pdf) or home-relative (~/out/chart.png) paths with supported extensions and uploads them as native attachments on Slack, Discord, Telegram, WhatsApp and Signal — images inline, audio as voice notes, everything else as files. Paths inside code blocks are ignored, and source files such as .py are deliberately excluded. Kanban workers can attach deliverables to kanban_complete too. The doc warns that "the agent doesn't reach for artifacts by default" — if you want the morning cron job to send a PDF instead of a wall of text, say so in the prompt or in AGENTS.md/SOUL.md.
💡 "Artifact" means three unrelated things in Hermes: Deliverable Mode above; the right-rail viewer in the Desktop app (PR #72345, v0.20.0 — a pure UI feature, persisted in that machine's localStorage); and the "task artifact" of the A2A protocol. Do not expect a central artifact store your pipelines can pull from — there is none today, and issue #19489 asking for an artifact-builder skill is still open.
Blueprints: A Skill With a Schedule
If cron is the engine, Automation Blueprints (v0.17.0, 19 June 2026, PR #41309) are the dashboard that lets people who never want to see cron syntax drive it — the release note's summary is "Pick an automation by name and Hermes asks you for what it needs — no cron syntax." But the line I want you to remember is the one in tools/blueprints.py:
"A blueprint is NOT a new object type. It is an ordinary skill that additionally declares an automation schedule in its frontmatter."
Everything you know from #6 Skills therefore applies unchanged. Take an existing SKILL.md, add a metadata.hermes.blueprint block, and it becomes a blueprint that flows through the same skills-hub pipeline — which, by the docs' own logic, means every mechanism that wraps a skill (including the install-time scanning described in #6) wraps a blueprint too:
---
name: nightly-backlog-triage
description: Rank the open backlog every night and post a summary to the team's Slack
metadata:
hermes:
blueprint: # shape follows the BlueprintSpec fields — verify against tools/blueprints.py before relying on it
schedule: "0 2 * * *" # skill_name, schedule, deliver (default origin), prompt,
deliver: slack # no_agent, model, provider, enabled_toolsets
prompt: "Run this skill against the team repository"
enabled_toolsets: [web, file, terminal]
---
# Nightly Backlog Triage
1. Fetch every open issue with gh …
2. Group by label and age …
One definition is rendered four ways: a form in the dashboard, a /blueprint command in the CLI, the TUI and the messengers, a conversation in which the agent asks for what it needs one question at a time, and an entry in the docs catalog. The developer-facing Automation Blueprints page shows examples any software organization will recognise:
| Blueprint | What it does |
|---|---|
| Nightly Backlog Triage | groups and ranks the open backlog — the catalog example runs on 0 2 * * * |
| Automatic PR Code Review | reviews new PRs and comments back — triggered by a webhook, not cron |
| Docs Drift Detection | checks whether the documentation still matches the code |
| Dependency Security Audit | reports vulnerabilities at CVSS ≥ 7.0 |
| Issue Auto-Labeling | labels incoming issues |
| CI Failure Analysis | explains why a workflow failed |
| Auto-Port Changes Across Repos | carries a patch across repositories |
| Security Audit Pipeline · Deploy Verification · Alert Triage · Uptime Monitor | the ops set |
The Automation Blueprints guide page (hermes-agent.nousresearch.com/docs/guides/automation-blueprints) lists 17 blueprints as of 1 September 2026 — beyond the table above it adds Competitive Repository Scout, AI News Digest, Paper Digest, Stripe Payment Monitoring, Daily Revenue Summary and Content Pipeline (that count is attributed to the guide page only; the reference catalog is client-rendered and I could not count it). The organizational consequence of "blueprint = skill" is the useful part: you manage them the way you manage every skill — through the skills hub, through profile export — and when you are looking after many machines, #10 Desktop & Fleet picks the story up.
Hermes as a Backend for the Apps You Already Run
The fifth door is the inversion: instead of Hermes calling other systems, other systems call Hermes. Since v0.4.0 the gateway carries an OpenAI-compatible API server (the release note says it shipped "hardened with input limits, field whitelists, SQLite-backed response persistence, and CORS origin protection" from day one):
# ~/.hermes/.env (or gateway.api_server.* in config.yaml — env wins)
API_SERVER_ENABLED=true
API_SERVER_KEY=<long-random-key> # mandatory, even on loopback
# start the gateway as usual — the API server listens on http://127.0.0.1:8642
hermes gateway
# any OpenAI client works
curl http://127.0.0.1:8642/v1/chat/completions \
-H "Authorization: Bearer $API_SERVER_KEY" \
-H "X-Hermes-Session-Key: helpdesk-user-4471" \
-H "Content-Type: application/json" \
-d '{"model":"hermes","messages":[{"role":"user","content":"Summarise my open tickets"}],"stream":true}'
# the model name is configurable via gateway.api_server.model_name — read the real value from GET /v1/models
| Endpoint | Behaviour |
|---|---|
POST /v1/chat/completions | stateless; SSE streaming with hermes.tool.progress events so a UI can show which tool the agent is using (v0.7.0) |
POST /v1/responses | server-side history via previous_response_id or a named conversation; SQLite store, LRU of 100 responses |
POST /v1/runs · GET /v1/runs/{id} · /events | long asynchronous jobs with SSE; stoppable with POST /v1/runs/{id}/stop (v0.12.0) |
GET /v1/models · GET /v1/capabilities | lets a client discover what this server can do |
X-Hermes-Session-Id / X-Hermes-Session-Key | Id = the transcript; Key (≤256 characters) = a stable memory scope per user |
/p/<profile>/… | several profiles on one server, each with its own key |
Limits to know before you design around it: no file upload, a default cap of 10 concurrent runs (max_concurrent_runs), and the hermes-api-server toolset drops clarify, TTS, computer_use and send_message — there is no human on the other end to ask.
The first app most organizations plug in is Open WebUI — and Open WebUI's own documentation carries a "Connect an agent → Hermes Agent" page: base URL http://localhost:8642/v1 (host.docker.internal from Docker; the /v1 suffix is required), the API key is API_SERVER_KEY, and the gateway must stay running. Hermes's Open WebUI page adds that tools execute on the API-server host, that Open WebUI talks server-to-server so no CORS setting is needed, and that per-user isolation is done with profiles on separate ports.
hermes peer, A2A, and what hermes mcp serve does not do
These three are routinely confused with one another:
hermes peer(v0.21.0, PR #88725) lets two Hermes instances talk across hosts — but it is not A2A. The PR says it "uses the peer's existing api_server platform as the wire" (POST/api/sessions/{id}/chatwith a bearer key). Peers are configured underbot_peers, keys live in.envasHERMES_PEER_<NAME>_KEY, and the CLI ishermes peer add|dm|run --idempotency-key|status|stop. Replies land in each side's Bot Chat.- A2A (a bundled plugin implementing A2A v1.0, PR #77109, v0.20.0) listens on port 9900, serves
/.well-known/agent-card.jsonand JSON-RPC 2.0 (SendMessage, SendStreamingMessage, GetTask, ListTasks, CancelTask) — for talking to other vendors' agents that speak the standard. The client toolset is off by default, the server binds 127.0.0.1 by default, a remote bind requires a token, and there is an audit log at~/.hermes/a2a_audit.jsonl. The doc itself says that for same-machine work you should prefer delegation or Kanban (#2). hermes mcp serve(v0.6.0, 30 March 2026) does not expose the agent to anyone. Themcp_serve.pydocstring says it exists to "expose messaging conversations as MCP tools": ten messaging-bridge tools (conversations_list,messages_read,messages_send,events_wait,permissions_respondand so on) over stdio only. There is no "ask the agent" tool. If you want Claude Desktop to read and answer messages in the channels Hermes manages, this is the thing for that; if you want an app to make Hermes think, go back to 8642.
# Claude Desktop config — this reaches Hermes's inbox, not its brain (the gateway must be running)
{ "mcpServers": { "hermes-messaging": { "command": "hermes", "args": ["mcp", "serve"] } } }
hermes serve that one third-party guide uses is the Desktop app's headless backend, not the way to start the API server — the documented path is hermes gateway with API_SERVER_ENABLED=true.
Hardening and Scaling the Automation Plane
Every door opened in this post has had an uninvited visitor. #4 Security tells the CVE record as a whole; four entries belong specifically here:
| CVE | Component | Substance |
|---|---|---|
| CVE-2026-7113 GHSA-mq9g-m88j-mmvx · 27 April 2026 · Moderate · CWE-287 | gateway/platforms/webhook.py | the _INSECURE_NO_AUTH argument could disable signature verification on /webhooks/{route} — community issue #6440 (9 April) had described it first. The GHSA entry is a GitHub-reviewed mirror of the VulDB (CNA) record, not independent analysis, and lists "Patched versions: Unknown". |
| CVE-2026-7112 GHSA-r7hr-pvjh-r4p3 · 27 April 2026 · Moderate | gateway/platforms/api_server.py — function _check_auth | improper authentication in the API_SERVER_KEY handler — likewise a GHSA mirror of VulDB, "Patched versions: Unknown" |
| CVE-2026-14628 NVD · 4 July 2026 · CVSS 3.1 5.3 (Medium) | gateway/platforms/base.py — function extract_media (component "Live Webhook Endpoint") | path traversal; names hermes-agent up to 2026.5.16 |
| CVE-2026-14626 NVD · 4 July 2026 · CVSS 3.1 4.3 (Medium) | run_agent.py — AIAgent.run_conversation via the HTTP API | denial of service through the todos argument; names hermes-agent up to 2026.4.30 |
The first two name hermes-agent 0.8.0. The INSECURE_NO_AUTH story needs its full context: the official webhooks page itself sanctions it for loopback testing — it is a documented test switch, not a hidden backdoor — and what the CVE records is that the switch can disable signature verification for a whole route. The PR-review guide carries the line worth pinning to the wall: "Never use INSECURE_NO_AUTH in production." Three further entries on the same surface are now NVD-listed (checked 1 September 2026): CVE-2026-10224 (published 1 June 2026, CVSS 3.1 5.3 — _handle_webhook_request in gateway/platforms/feishu.py leads to resource consumption, remotely triggerable, 2026.4.0–2026.4.30; fixed by the "webhook body-cap sweep" PR #59215), CVE-2026-53869 (17 June 2026, CVSS 3.1 7.5 HIGH — a DNS-rebinding vulnerability in WebSocket endpoints that bypasses Host and Origin validation, <0.16.0) and CVE-2026-53870 (17 June 2026, CVSS 3.1 5.5 — response_store.db and webhook_subscriptions.json created world-readable at 0o644, exposing conversation history and HMAC secrets to local users, <0.16.0). After that, between 3 July and 6 August, a steady trickle of eleven further low-severity VulDB entries landed on other Hermes surfaces — the consolidated record is in #4 Security.
The hardening checklist I distilled from every document in this post:
- ☐ Run v0.21.0 or later — every entry above names a version before 0.16.0.
- ☐ Port 8644 sits behind a TLS-terminating reverse proxy; every route has a long random secret; use Generic V2 with its timestamp, never V1.
- ☐ Leave a route's
toolsetsalone unless you must — the four safe tools are the right default. - ☐ 8642 and 9900 listen on 127.0.0.1 unless there is a reason, and always carry a key or token.
- ☐ A shell hook acting as a guard sets
fail_closed: true— the fail-open default means a broken script is no guard at all. - ☐
cron.allow_agent_schedulingis still false andpreflightis still true. - ☐ Test the hooks for real after every update — this one has a history.
Hook reliability has not been a smooth story: issue #2817 (hooks not firing) preceded the v0.5.0 release note of 28 March 2026 announcing that lifecycle hooks "now fire", and issue #44582 (opened 12 June 2026 on 0.16.0: "pre_tool_call plugin hook not invoked during agent tool execution") was closed as not planned with no visible maintainer response. I am not saying hooks do not work today — I am saying do not believe it until you have seen it:
# after every hermes update — make it part of the runbook
hermes hooks test pre_tool_call
hermes hooks test on_session_end
hermes plugins doctor --ci # CI-friendly exit code
hermes cron doctor
Chronos — cron for a gateway that is allowed to sleep
The one architectural problem with built-in cron is that it needs a gateway that is awake all the time to tick every 60 seconds — which contradicts the scale-to-zero hosting #7 discussed. v0.17.0 (PR #48275) therefore made the scheduler pluggable and added a provider called Chronos, specified in the repo document docs/chronos-managed-cron-contract.md: instead of ticking itself, the gateway registers its jobs with an external service that arms one-shot timers; when one fires, the service calls back to /api/cron/fire with a short-lived JWT (roughly 60–120 seconds, aud=agent:{instance_id}, purpose=cron_fire), the gateway answers 202 {"status":"accepted"} and runs the job in the background — at-most-once via a compare-and-set on the store.
# config.yaml — enable Chronos; leave it unset and the built-in ticker is used
cron:
provider: chronos
chronos:
portal_url: …
callback_url: …
expected_audience: …
nas_jwks_url: …
This is the mechanism that lets a hosted gateway that scales to zero — the kind Hermes Cloud advertises — keep its cron jobs without paying for an idle machine. I will not print Cloud prices or server sizes, because there are no official figures. For a self-hosted single VPS the built-in ticker is enough; Chronos is the option when you move to hosting that bills only for time actually worked.
A closing note for anyone migrating from OpenClaw: as #7 recorded, hermes claw migrate does not carry cron jobs or hooks across — and I see that as an opportunity rather than a chore. Every heartbeat in the old system deserves to be asked again whether it is a watchdog (layer 1), a monitor (layer 2), or work that genuinely needs thought (a regular cron job). My bet is that most of them fall into the first two.
🎯 Key Takeaways
- Five doors = cron · inbound webhook · hooks/outbound · API server · A2A/peer — start with cron and widen from there
- Cron is plain files =
jobs.json+executions.db+output/*.md; the gateway ticks every 60 seconds; agents cannot schedule agents by default - Three zero-token layers =
--no-agent(v0.13.0) · monitor mode that never calls the LLM on an unchanged hash (v0.21.0) ·deliver_only/hermes send - Cron that remembers = memory + continuity + notepad (v0.21.0), plus
context_from(v0.12.0) to chain jobs - The ops loop =
hermes cron incidents / doctor / runs— errors redacted, exit codes fit for monitoring, a model-drift guard against untested models - Webhooks = 8644, a secret on every route, a ±300 s replay window, four safe tools — "HMAC authenticates the sender, not the content"
- Outbound = HMAC-signed
hooks.outbound(v0.20.0), shell hooks that veto with exit 2 (v0.11.0), Deliverable Mode for real attachments - A blueprint = an ordinary skill with a schedule in its frontmatter — everything from #6 applies; the guide page lists 17 as of 1 September 2026
- 8642 = the OpenAI-compatible API;
hermes mcp serveis a messaging bridge, not the agent;hermes serveis not the API server command - Hardening = CVE-2026-7113/7112 on 0.8.0 (GHSA mirrors VulDB, patch unknown) plus 14628/14626 on NVD, July 2026;
hermes hooks testafter every update; Chronos for a gateway that scales to zero