ในบทความนี้
- บันไดห้าขั้นของการเลือก Model
- Portal ในฐานะ Routing Tier — ไม่ใช่ศาสนา
- สามชั้นของความทนทาน — Pool, Fallback, Auxiliary
- Token หายไปไหน — ต้นทุนคงที่ต่อการเรียกหนึ่งครั้ง
- เศรษฐศาสตร์ของ Prompt Cache
- ตัวเลขต้นทุนวัดอะไรกันแน่
- GPU ของเราเอง — vLLM, llama.cpp, Ollama และพื้น 64K
- ข้อมูลไปอยู่ที่ไหน — Telemetry, Privacy Policy และ Air Gap ที่ยังไม่มี
- PDPA และ Decision Tree สำหรับองค์กรไทย
In this post
- The Model Ladder: Five Rungs an Organization Climbs
- Portal as a Routing Tier, Not a Religion
- Three Layers of Resilience: Pools, Fallbacks, Auxiliaries
- Where the Tokens Go: The Fixed Cost of Every Call
- Prompt-Cache Economics
- What the Cost Numbers Actually Measure
- Bring Your Own GPUs: vLLM, llama.cpp, Ollama and the 64K Floor
- Where the Data Goes: Telemetry, the Privacy Policy and the Air Gap That Isn't
- PDPA and a Decision Tree for Thai Organizations
🤔 ลองนึกภาพนี้ดูครับ — ปลายเดือน บิลจาก provider มาถึงพร้อมตัวเลขที่สูงกว่าที่ dashboard ของ agent บอกไว้หลายเท่า คุณเปิด /insights ก็เห็นตัวเลขหนึ่ง เปิดหน้า billing ของ provider ก็เห็นอีกตัวเลขหนึ่ง แล้วตัวเลขไหนคือ "ความจริง"? และคำถามที่ซ่อนอยู่ใต้คำถามนั้น — ระหว่างที่ agent วิ่งอยู่ทั้งคืน ข้อมูลขององค์กรที่มันอ่านผ่านตา ไปนอนอยู่บน server ของใครบ้าง?
สองคำถามนี้ — จ่ายเท่าไร กับ ข้อมูลไปไหน — พันกันแน่นกว่าที่เห็น เพราะทุกครั้งที่เราขยับขึ้นลงบันไดของการเลือก model ตั้งแต่ Nous Portal ที่จ่ายบิลใบเดียว ไปจนถึง GPU ในห้อง server ของเราเอง เรากำลังแลกต้นทุนกับการควบคุมพร้อมกันเสมอ ในซีรีส์นี้ผมแตะเรื่องเงินไว้แล้วสองครั้ง — ตาราง tier ของ Portal และกรอบคิด "MIT ฟรี จ่ายเฉพาะค่า model" อยู่ใน #1 Hermes 101 ส่วนหัวข้อสั้น ๆ เรื่องคุมค่าใช้จ่ายและ BYO keys vs Portal อยู่ใน #7 Production — บทความนี้คือฉบับเต็มของทั้งสองหัวข้อนั้น
ผมจะไล่จากบันไดห้าขั้นของการเลือก model, สามชั้นของกลไก fallback, ที่ที่ token หายไป, เศรษฐศาสตร์ของ prompt cache, ตัวเลขต้นทุนที่ Hermes รายงานให้ (และสิ่งที่มันมองไม่เห็น), การรัน open-weight บน GPU ของเราเอง แล้วปิดด้วยคำถามที่องค์กรไทยหนีไม่พ้น — PDPA ทุกตัวเลขในบทความมาจากเอกสารทางการ, release notes, issue tracker หรือหน้า Portal ที่ผมดึงเมื่อ 1 กันยายน 2026 ตรงไหนที่แหล่งข้อมูลขัดกันเอง ผมจะบอกว่าขัดกันอย่างไร แทนที่จะเลือกตัวเลขที่ฟังดูดีกว่า
บันไดห้าขั้นของการเลือก Model
เวลาผมคุยกับหน่วยงานที่กำลังพิจารณา Hermes คำถามแรกมักเป็น "ควรใช้ Portal หรือเอา key ตัวเองมาเสียบ" ซึ่งเป็นคำถามที่ตั้งแคบไป ในความเป็นจริง Hermes ให้เราเลือกยืนได้ห้าขั้น และแต่ละขั้นตอบคำถามสามข้อต่างกัน: ใครถือ key, ข้อมูลวิ่งไปที่ไหน และใครเป็นคนส่งบิล
- Nous Portal — OAuth ครั้งเดียว บิลใบเดียว README ระบุ "300+ models — pick any of them with /model <name>" และ subscription แบบจ่ายเงินปลดล็อก Tool Gateway (web, image, TTS, cloud browser) ให้ด้วย ขั้นนี้เร็วที่สุด และเป็นขั้นเดียวที่ข้อมูลผ่านมือ Nous เพิ่มมาอีกหนึ่งทอด
- BYO keys — key ของ OpenRouter หรือ key ตรงของ provider (Anthropic, OpenAI ฯลฯ) บิลกระจายไปตามเจ้าของ key แต่เส้นทางข้อมูลสั้นลงหนึ่งทอด
- Named custom provider หลัง gateway ขององค์กร — ประกาศ provider ชื่อของเราเองใน config ให้ชี้ไปที่ LLM gateway ภายใน (LiteLLM, proxy ของฝ่าย IT หรือสัญญา Bedrock ที่มีอยู่แล้ว) พร้อม token จาก SSO ที่หมดอายุสั้น ๆ
- Self-hosted inference — vLLM, llama.cpp, SGLang, LM Studio หรือ Ollama บน GPU ของเราเอง ค่า token เป็นศูนย์ ค่าไฟและค่าคนไม่เป็นศูนย์
- Air-gapped — ขั้นที่หลายหน่วยงานราชการอยากได้ และเป็นขั้นที่ Hermes ยังไปไม่ถึงอย่างเป็นทางการ (ผมจะอธิบายในหัวข้อเรื่องข้อมูล)
สิ่งที่ทำให้บันไดนี้ใช้งานได้จริงคือทุกขั้นสลับกันด้วยคำสั่งเดียวกัน หลักการที่เอกสาร configuring-models วางไว้: hermes model คือ wizard นอก session ที่ใช้ เพิ่ม provider ส่วน /model ใน session ใช้ สลับ เฉพาะสิ่งที่ตั้งค่าไว้แล้ว และมีขอบเขตสามระดับ:
# สลับเฉพาะ session นี้ — ปิด session แล้วกลับเป็นค่าเดิม
/model <name> --provider <p>
# เขียนลง config.yaml ถาวร
/model <name> --provider <p> --global
# ใช้แค่ turn เดียว แล้วคืนค่าเดิมอัตโนมัติ
/model <name> --once
# ตั้ง model_aliases ไว้ใน config แล้วเรียกสั้น ๆ
/model fav
ข้อควรจำสำหรับคนดูแล gateway: session ที่รันผ่าน gateway จะรับค่า config ใหม่เฉพาะ session ใหม่ เท่านั้น เว้นแต่จะ restart gateway — ถ้าเปลี่ยน model แล้วบอทใน Telegram ยังตอบด้วย model เดิม นี่คือสาเหตุแรกที่ควรสงสัย
ขั้นที่สามคือขั้นที่องค์กรขนาดกลางขึ้นไปมักลงเอย และเป็นขั้นที่เอกสารรองรับดีเกินคาด providers.<name> รับ api, key_env และ transport ที่เลือกได้ระหว่าง chat_completions กับ anthropic_messages ส่วน key_cmd จะรัน CLI ของเราเพื่อ mint token อายุสั้นจาก enterprise SSO แล้วต่ออายุให้เองก่อนหมดอายุ:
# config.yaml — provider ชื่อของเราเอง หลัง gateway ขององค์กร
providers:
corp-gateway:
api: https://llm-gateway.internal.example.ac.th/v1
transport: chat_completions # หรือ anthropic_messages
key_cmd: "my-auth-cli print-token --profile prod" # token จาก SSO อายุสั้น ต่ออายุเอง
# สลับใน session
/model custom:corp-gateway:<model>
และตั้งแต่ v0.21.0 (31 สิงหาคม 2026) บุคคลที่สามส่ง provider มาเป็น pip package ที่ Hermes ค้นพบผ่าน entry points ได้ — แปลว่าฝ่าย platform ขององค์กรสามารถแพ็ก provider ภายในเป็น package หนึ่งตัวแล้วแจกให้ทุกเครื่อง โดยไม่ต้องแก้ config ทีละคน (เรื่องนี้ต่อเนื่องไปถึงการดูแล fleet ใน #10 Desktop & Fleet)
Portal ในฐานะ Routing Tier — ไม่ใช่ศาสนา
ใน #7 ผมยังไม่ยอมฟันธงว่า Free tier ให้อะไรบ้าง เพราะหน้าเอกสารกับหน้า pricing ไม่ตรงกัน วันนี้หน้า pricing ของ portal.nousresearch.com (ดึงเมื่อ 1 กันยายน 2026) ตอบชัดแล้ว: Free คือ "Free models only" กับเครดิต $0 ต่อเดือน ส่วนตาราง tier ทั้งสี่ (Free $0 · Plus $20 → เครดิต $22 · Super $100 → $110 · Ultra $200 → $220) ผมพิมพ์ไว้แล้วใน #1 Hermes 101 จึงไม่ยกมาซ้ำ — สิ่งที่ #1 ยังไม่มีคือรายละเอียดต่อไปนี้:
- Rollover cap — เครดิตที่ใช้ไม่หมดยกไปเดือนหน้าได้ แต่มีเพดาน $10 (Plus) / $50 (Super) / $100 (Ultra) องค์กรที่ใช้งานเป็นจังหวะ เช่น หนักปลายภาคเรียนแล้วเงียบช่วงปิดเทอม จะเสียเครดิตส่วนที่เกินเพดานไปเปล่า ๆ ทุกเดือนเบา
- การ์ด tier จ่ายเงิน ระบุเพิ่มว่า "Hosted tool usage" และ "High rate limits" (Free ได้ "Standard rate limits")
- Top-up — เติมเครดิตแบบ pay-as-you-go ได้ครั้งละ $10, $20, $50, $100 หรือ $200
- Copy ที่ยังไม่ตรงกันเอง — การ์ด tier เขียน "200+ Models" ขณะที่ catalog บนหน้าเดียวกันเขียน "over 300" และ README เขียน "300+" ผมยึด 300+ ตาม README และบันทึกไว้ว่าการ์ดยังตามไม่ทัน หน้าเดียวกันระบุช่วงราคา catalog ว่า "free to $24 per million input tokens"
เรื่อง "10%" ต้องแยกให้ออกว่ามีสองตัวที่คนละความหมาย ตัวแรกคือ โบนัสเครดิต บนหน้า pricing — "10% bonus" หมายถึงจ่าย $20 ได้เครดิต $22 (และ $100 → $110, $200 → $220) ตัวที่สองคือ ส่วนลด provider ในเอกสาร configuring-models — "Portal subscribers also get 10% off token-billed providers" คือราคาต่อ token ที่ถูกลงเมื่อ route ผ่าน Portal ไม่ใช่เครดิตแถม ส่วนตัวเลขที่สาม — โพสต์ของ Teknium (ผู้ร่วมก่อตั้ง Nous) บน X ตอนออก release Herald ที่พูดถึง "20% Discount on all models" — มีอยู่บน X เท่านั้น ไม่ปรากฏบนหน้า Portal หรือในเอกสาร และผมยืนยันไม่ได้ว่ายังใช้อยู่หรือเคยใช้กับอะไร ผมจะไม่พิมพ์ 20% เป็นข้อสรุป — เข้า Portal แบบ login แล้วดูตัวเลขจริงก่อนคำนวณงบ
สิ่งที่ subscription แบบจ่ายเงินให้เพิ่มจริง ๆ คือ Tool Gateway — เอกสารบอกว่า "included with every paid Nous Portal subscription" ส่วนบัญชี Free "can use Portal for inference but don't include managed tools" (โดยมีข้อยกเว้นว่า "some accounts are also entitled to a free tool pool") backend ที่ Nous จัดการให้มีห้าตัว: Firecrawl สำหรับ web, FAL สำหรับ image generation, OpenAI TTS, Browser Use สำหรับ cloud browser และ Modal cloud terminal เป็น add-on เสริม — จุดที่ผมชอบคือมัน "opt-in per tool, not all-or-nothing":
# config.yaml — เลือกใช้ managed tool ของ Nous ทีละตัว ไม่ต้องยกทั้งชุด
web:
backend: nous # Firecrawl ผ่าน Portal
image_gen:
provider: nous # FAL
tts:
provider: nous # OpenAI TTS
stt:
provider: nous
browser:
cloud_provider: nous # Browser Use — เปิดเฉพาะเมื่อต้องการ
ตั้งแต่ v0.19.0 (20 กรกฎาคม 2026) การจัดการแผนทำได้จากใน TUI/CLI โดยตรงผ่าน /subscription และ /topup — สองคำสั่งนี้ยังไม่ปรากฏในรายการ slash command ของคู่มือผู้ใช้ ผมอ้างจาก release notes — และ hermes portal info แสดงสถานะ login, subscription และ routing ปัจจุบัน เบื้องหลัง Hermes จะ mint JWT อายุสั้นจาก refresh token ของ Portal ทุกครั้งที่เรียก inference
ทีนี้ถึงคำถามที่ตอบผิดกันบ่อยที่สุด: "ก็ Nous ทำโมเดล Hermes 4 เอง ทำไมไม่รัน Hermes Agent บน Hermes 4 ให้ถูกที่สุดไปเลย" คำตอบมาจาก Nous เอง ในคู่มือ run-hermes-with-nous-portal:
💡 "Hermes-4-70B and Hermes-4-405B are available on the Portal at deep discounts, but they're chat/reasoning models, not tool-call-tuned. They will struggle with multi-step agent loops… For Hermes Agent itself, stick to the frontier agentic models." — ชื่อเดียวกัน แต่ไม่ใช่ของสำหรับงานเดียวกัน
ที่ที่ Hermes 4 เข้าที่เข้าทางคือ subscription proxy — hermes proxy start เปิด server แบบ OpenAI-compatible บนเครื่องเรา (/v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models) ให้แอปที่ไม่ใช่ agent เช่น Open WebUI ใช้ subscription เดียวกันสำหรับงาน chat/research ตอนนี้ส่งมาแค่ nous กับ xAI:
# เปิด proxy บนเครื่องนี้ แล้วชี้แอปอื่นมาที่ localhost
hermes proxy start
ปิดหัวข้อนี้ด้วยเชิงอรรถสำหรับสาย Claude ซึ่งกระทบผู้อ่านทั้งซีรีส์นี้และ ซีรีส์ OpenClaw: เส้นทาง OAuth ของ Anthropic ใน Hermes ต้องการ "Claude Max subscription + purchased extra usage credits" และกินเฉพาะเครดิตส่วนเกินนั้น ส่วน "Claude Pro subscribers cannot use the OAuth path; the supported alternative is an ANTHROPIC_API_KEY" แบบจ่ายตาม token — ถ้าคุณตั้ง OpenClaw ไว้ด้วย Claude Pro เมื่อต้นปี บันไดขั้นที่สองของคุณคือ API key ไม่ใช่ subscription
สามชั้นของความทนทาน — Pool, Fallback, Auxiliary
ใน #7 ผมสรุป fallback chain ไว้บรรทัดเดียวว่าเป็น "ทางรอดจาก 429" ฉบับเต็มมีสามชั้นซ้อนกัน และเอกสาร fallback-providers แยกไว้ชัด: credential pool หมุนหลาย key ของ provider เดียวกัน, fallback_providers ย้ายไป provider:model คนละตัว และ auxiliary task แต่ละงานมีเส้นทางหา provider ของตัวเอง ลำดับคือ pool ถูกลองก่อน fallback เสมอ
ชั้นแรก — fallback chain มาตั้งแต่ v0.6.0 (30 มีนาคม 2026) เป็น list ระดับบนสุดของ YAML และมีคำสั่งจัดการมาตั้งแต่ v0.12.0:
# config.yaml — ลำดับสำรองเมื่อ primary ล้ม
fallback_providers:
- provider: openrouter
model: <provider/model>
- provider: nous
model: <model>
# หรือจัดการผ่าน CLI โดยไม่แก้ไฟล์
hermes fallback list
hermes fallback add <provider> <model>
hermes fallback remove <provider> <model>
ค่าเดิม fallback_model แบบเอกพจน์ยังใช้ได้ แต่ถ้ามีทั้งสองแบบ list จะชนะ และไม่มี environment variable ให้ตั้ง — เอกสารบอกว่าเป็นความตั้งใจให้ควบคุมผ่าน YAML/CLI เท่านั้น สิ่งที่ต้องรู้ให้แม่นคือ อะไร ทำให้ fallback ทำงาน:
| สถานการณ์ | พฤติกรรม |
|---|---|
| HTTP 429 · 500 · 502 · 503 | retry จนครบก่อน แล้วจึง fallback |
| HTTP 401 · 403 · 404 | fallback ทันที |
| response ผิดรูปซ้ำ ๆ | fallback |
| 429 ชั่วคราวที่แจ้ง reset window มาด้วย | ไม่ fallback — Hermes รอ |
| ขอบเขต | ต่อ turn — ข้อความใหม่ทุกครั้งเริ่มที่ primary เสมอ |
ชั้นที่สอง — credential pool (v0.7.0, 3 เมษายน 2026, release ที่ Nous เรียกว่า "the resilience release") รองรับ OpenRouter, Anthropic, Nous Portal และ custom endpoint แบบ OpenAI-compatible:
# เพิ่ม key ที่สองของ provider เดิมเข้า pool
hermes auth add <provider>
hermes auth list
hermes auth reset # ล้างสถานะ cooldown
# config.yaml — กลยุทธ์การหมุนต่อ provider
credential_pool_strategies:
openrouter: round_robin # fill_first (ค่าเริ่มต้น) | round_robin | least_used | random
กติกาการหมุนละเอียดกว่าที่คิด: เจอ 429 จะ retry หนึ่งครั้ง ถ้าโดน 429 ติดกันครั้งที่สองจึงหมุน key (cooldown 1 ชั่วโมง), เจอ 402 หมุนทันที (cooldown 1 ชั่วโมง), เจอ 401 ลอง refresh token ก่อนแล้วค่อยหมุน (cooldown 5 นาที) และเมื่อทุก key ใน pool หมดสภาพ จึงตกไปที่ fallback สถานะทั้งหมดอยู่ใน ~/.hermes/auth.json — อีกไฟล์หนึ่งที่ต้องอยู่ใน backup ตามสูตร #7
ชั้นที่สาม — auxiliary tasks คืองานเบื้องหลังที่ไม่ใช่ตัว agent หลัก หน้า configuring-models แจกแจงไว้สิบเอ็ดช่อง: title generation, vision, compression, approval scoring, web extraction, skills hub, MCP tool routing, triage specifier, Kanban decomposer, profile describer และ curator (ในเอกสาร fallback-providers ช่องเหล่านี้ปรากฏเป็น key ของ YAML เช่น vision, compression, skills_hub, mcp, approval, title_generation, triage_specifier) ค่าเริ่มต้น provider: auto คือใช้ model หลักก่อน และตั้งแต่ v0.15.0 (28 พฤษภาคม 2026) มีบันไดสำรองเป็นชั้น ๆ: model หลัก → auxiliary.<task>.fallback_chain → fallback_providers → discovery chain ในตัว (OpenRouter → Nous Portal → custom endpoint → Codex OAuth → provider ที่มี API key → ยอมแพ้)
# config.yaml — งานเบื้องหลังใช้ model ฟรี/ถูก แยกจาก model หลัก (รูปแบบตามตัวอย่างในเอกสาร fallback-providers)
auxiliary:
compression:
provider: auto # ลอง model หลักก่อน
model: ""
fallback_chain:
- provider: openrouter
model: inclusionai/ring-2.6-1t:free # ตัวอย่าง :free ในเอกสารเอง
vision:
provider: auto # ใช้ model หลักถ้ามัน "มองเห็น" ได้
นี่คือรูปแบบ "auxiliary ศูนย์บาท" ที่เอกสารรับรอง — v0.8.0 (8 เมษายน 2026) เพิ่ม Xiaomi MiMo v2 Pro แบบฟรีบน Portal สำหรับ compression/vision/summarization และตัวอย่างในเอกสาร fallback-providers ก็วาง inclusionai/ring-2.6-1t:free ของ OpenRouter ไว้ใน fallback_chain ของ compression (อีกตัวอย่างหนึ่งบนหน้าเดียวกันตั้ง google/gemini-3-flash-preview เป็น model ของ compression ซึ่งไม่ใช่ตัว :free) สองข้อควรระวัง: compression ที่หา model ไม่ได้จะ "degrade to no-summary" คือเงียบ ๆ ไม่สรุปให้ และ subagent สืบทอด chain ของ parent ทั้งชุด — ผลทางต้นทุนของการ delegate ผมเล่าไว้ใน #2 Agent Teams
ราคาที่ทุกชั้นเรียกเก็บเหมือนกัน และเอกสารเขียนไว้ตรง ๆ: การ fallback หรือการหมุน key จะ reset cache key ทำให้ "the next request re-reads the entire history at full input-token price" — และเสียอีกรอบตอนสลับกลับ ผมจะกลับมาที่เรื่องนี้ในหัวข้อ cache
Token หายไปไหน — ต้นทุนคงที่ต่อการเรียกหนึ่งครั้ง
ก่อนคุยเรื่องราคาต่อล้าน token เราต้องรู้ก่อนว่า แต่ละคำขอ ส่งอะไรออกไปบ้าง เพราะ agent อย่าง Hermes ไม่ได้ส่งแค่ข้อความของผู้ใช้ — มันแนบ system prompt, schema ของทุก tool ที่เปิดอยู่, ดัชนี skills, ความจำ และ profile ไปด้วยทุกครั้ง
ตัวเลขที่ทำให้เรื่องนี้เป็นรูปธรรมมาจาก issue #4379 (เปิด 1 เมษายน 2026 โดยผู้ใช้ชื่อ Bichev, ยังเปิดอยู่, P2, label area/usage-cost) ซึ่งวัดที่ v0.6.0 ได้ overhead คงที่ราว 13,935 tokens ต่อการเรียกหนึ่งครั้ง — ประมาณ 73% ของคำขอทั่วไป — แยกเป็น system prompt ~5,176, tool definitions 31 ตัว ~8,759 และดัชนี skills ~2,200 และหนึ่งค่ำคืนของ gateway Telegram/WhatsApp/cron กินไป ~3.9 ล้าน tokens จาก 207 API calls
hermes prompt-size ซึ่งทำงาน offline และไม่เสียเงิน
# งบ prompt คงที่ต่อข้อความ — system prompt, skills index, memory, profile, tool schemas
# "Runs entirely offline — no API call" (v0.16.0, tagged 2026.6.5)
hermes prompt-size
hermes prompt-size --platform telegram --json # แยกตาม platform เพราะ toolset ต่างกัน
# toolset ไหนกิน token เท่าไร (ตัวเลขประมาณการ)
hermes tools
สูตรลดน้ำหนักที่ FAQ ทางการวางไว้เรียบง่าย: วัดก่อน แล้วปิด toolset กับ skill ที่ไม่ใช้ ตั้งแต่ v0.11.0 cron job แต่ละงานกำหนด enabled_toolsets ของตัวเองได้เพื่อ "cap token overhead + cost per job" — งานสรุปข่าวตอนเช้าไม่ต้องแบก schema ของ browser automation ไปด้วย (รายละเอียดฝั่ง scheduling อยู่ใน #8 Automation)
เครื่องมือชิ้นใหญ่กว่านั้นคือ Tool Search — โหมด opt-in ที่แทน schema ของ MCP tools และ plugin tools ที่ไม่ใช่ core ด้วย bridge tool สามตัว tool_search / tool_describe / tool_call แล้วให้ model ค้นหา tool ที่ต้องใช้เมื่อถึงเวลา แทนที่จะแบกทุก schema ไปทุก turn (progressive disclosure) กติกาสำคัญ: "Built-in Hermes tools never defer" — tool ในตัวยังอยู่ครบเสมอ ดังนั้นประโยชน์จะมากเป็นพิเศษกับองค์กรที่ต่อ MCP server หลายตัวตามที่เล่าใน #5 Integrations
คันโยกถัดมาที่เอกสาร configuration ระบุตรง ๆ ว่าเป็นเรื่องต้นทุนคือ reasoning effort — ตั้งได้ทั้งระดับ agent และแยกต่อ auxiliary task:
# config.yaml — ความลึกของการคิดเป็นเงิน
agent:
reasoning_effort: medium # none | minimal | low | medium | high | max (v0.19.0 เพิ่มชั้นบนสุด max/ultra)
auxiliary:
compression:
reasoning_effort: low # งานสรุปไม่ต้องคิดลึก
vision:
reasoning_effort: none # อ่านภาพ ไม่ต้อง reasoning
และ Nous เองก็ไล่ตัดค่าใช้จ่ายซ่อนเร้นของตัวเองมาเป็นระยะ ซึ่งบอกเราได้ว่าเงินเคยรั่วตรงไหน: v0.15.0 สร้าง session_search ใหม่แบบไม่ใช้ LLM — release notes บอกว่าของเดิม "cost ~$0.30/call and took ~30 seconds to summarize three sessions, sometimes confabulating"; v0.17.0 ทำให้ consolidation pass ของ skill Curator เป็น opt-in (curator.consolidate: true) เพื่อให้รอบปกติ "no aux-model spend"; และ v0.18.0 ย้าย self-improvement fork หลังจบ turn ไปใช้ auxiliary model ที่ย่อย context แทนการเล่นทั้งบทสนทนาซ้ำ กฎง่าย ๆ ที่ผมสรุปจากทั้งหมดนี้: งานเบื้องหลังทุกชิ้นคือการเรียก model และทุกการเรียกคือเงิน — ให้แน่ใจว่าคุณรู้ว่ามีงานเบื้องหลังอะไรเปิดอยู่บ้าง
เศรษฐศาสตร์ของ Prompt Cache
ถ้า overhead คงที่คือปัญหา prompt cache คือคำตอบครึ่งหนึ่ง — เพราะส่วนที่ซ้ำทุก turn (system prompt, tool schemas) คือส่วนที่ cache ได้ดีที่สุด กลไกของ Hermes ตาม developer guide: caching เปิดอัตโนมัติเมื่อชื่อ model เป็น Claude และ provider รองรับ cache_control ใช้กลยุทธ์ system_and_3 คือวาง breakpoint สี่จุด — system prompt บวกข้อความล่าสุดสามข้อความที่ไม่ใช่ system — และเอกสารอ้างว่า "Reduces input token costs by ~75% on multi-turn conversations"
# config.yaml
prompt_caching:
cache_ttl: "5m" # รับแค่ "5m" (ค่าเริ่มต้น) หรือ "1h" — ปรับได้ตั้งแต่ v0.12.0
สิ่งที่ทำให้ cache แตกมีสามอย่าง และสองในสามคือสิ่งที่เราเพิ่งตั้งค่าไปในหัวข้อก่อน: การแก้ไขประวัติกลางทาง, การเปลี่ยนตัวตนของ model (/model, fallback อัตโนมัติ, การหมุน credential) และ compression — ข้อสุดท้ายนี้ cache ของ system prompt รอด และส่วนที่เหลือกลับมาติดใหม่ภายใน 1–2 turn หน้า tips ทางการสรุปเป็นกฎที่ผมอยากให้ทุกทีมจำ:
💡 "an explicit /model switch, an automatic provider fallback, or a credential-pool rotation all force the next turn to re-read the entire conversation at full input price" — และใน session ยาว ๆ "it's often cheaper to start a fresh session on the other model than to bounce back and forth"
คำถามที่ตามมาคือ "แล้ว provider ไหน cache ได้บ้าง" ซึ่งเป็นจุดที่ เอกสารตามหลัง release notes — developer guide ยังเขียนว่าเฉพาะ Anthropic native หรือ OpenRouter แต่ configuration.md เพิ่ม Nous Portal ไว้แล้ว และ release notes เพิ่มไปอีกหลายราย: xAI ผ่าน header x-grok-conv-id (v0.8.0), Bedrock Converse API ผ่าน cachePoint และ DeepSeek บน OpenCode gateway (v0.20.0), Claude ผ่าน LiteLLM บน OpenAI wire (v0.20.2/v0.21.0) และ api.meta.ai ผ่าน Responses API (v0.21.0) v0.20.0 ยังทำให้ tool schema บน Anthropic native cache ได้ "without history loss" — พูดอีกแบบ ถ้าคุณอยู่ที่บันไดขั้นสาม (gateway ขององค์กร) โอกาสที่ cache จะทำงานสูงกว่าที่เอกสารหน้าหลักบอกไว้ แต่ต้องดู cache_read ใน /usage ยืนยันเอง
อีกครึ่งของสมการคือ compression ซึ่งผมเล่ากลไกไว้ใน #3 Memory วันนี้ขอเพิ่มเฉพาะมิติต้นทุน ค่าเริ่มต้นตาม developer guide:
# ค่าเริ่มต้นของ compression (developer guide) — ตัวอย่างประกอบสำหรับ model 200K
compression:
threshold: 0.50 # เริ่มบีบเมื่อใช้ context ถึง 50% → 100,000 tokens
target_ratio: 0.20 # สัดส่วนส่วนท้ายที่เก็บไว้ → tail budget 20,000
tail_mode: lean # ค่าเริ่มต้นตั้งแต่ v0.20.6: 2.5% ของ context, พื้น 10K / เพดาน 25K
# gateway มี safety net ที่ 0.85
# งบ summary = content × 0.20, ต่ำสุด 2,000, สูงสุด min(context × 0.05, 12,000) → 10,000
ผลทางเงินสองข้อ ข้อแรก — ทุกครั้งที่บีบ คือการเรียก auxiliary compression model หนึ่งครั้งด้วยเนื้อหาเกือบทั้งหมดของ session และเอกสารกำหนดว่า model ตัวนั้น "must have a context window at least as large as the main agent model's" ดังนั้น "ใช้ model เล็กราคาถูกมาบีบ" ใช้ได้เฉพาะเมื่อ model เล็กนั้นมี context ใหญ่พอ ข้อสอง — หลังบีบ cache ส่วนประวัติแตก ดังนั้น turn ถัดไปแพงกว่าปกติเสมอ ถ้าเห็น /usage พุ่งเป็นช่วง ๆ โดยไม่มีสาเหตุ ให้ดูว่าตรงกับจังหวะ compression หรือไม่
/compress แบบ manual รายงาน timeout 120 วินาทีทั้งที่ worker เบื้องหลังทำสำเร็จในอีกหลายนาทีถัดมา ส่วน session ขนาดใหญ่ล้มด้วย lease lost / session_split_failed ถ้าคุณพึ่ง compression เพื่อคุมค่าใช้จ่ายของ session ยาว ให้ตรวจสถานะสองรายการนี้ก่อนอัปเกรด และดูใน /usage ว่า summary เกิดขึ้นจริง
เรื่อง cache ยังเป็นเหตุผลที่ prefix caching สำคัญมากเมื่อลงไปรัน model เอง — recipe ของนักพัฒนาอิสระบน dev.to (Qwen3.6-27B + vLLM บน RTX 3090 24 GB, พฤษภาคม 2026 — third-party ล้วน) วัด time-to-first-token ของ prompt ใหญ่ที่ 38.6 วินาทีแบบ cold เทียบกับ 1.59 วินาทีเมื่อ prefix cache ติด ต่างกันยี่สิบกว่าเท่า และ prompt ใหญ่ที่ว่านั้นก็คือ overhead คงที่ของ Hermes ที่เราเพิ่งวัดกันนั่นเอง
ตัวเลขต้นทุนวัดอะไรกันแน่
Hermes มีพื้นผิวรายงานต้นทุนอยู่หลายชั้น และแต่ละชั้นวัดคนละอย่าง ไล่จากใกล้ตัวที่สุด:
# ใน session
/usage # token ของ session นี้ + rate-limit headers และ cost detail (v0.9.0),
# account limits (v0.11.0), แยกตามหมวดของ context (v0.18.0)
/insights --days 7 # ภาพรวมย้อนหลัง
# นอก session
hermes insights --days 30 --source telegram
# งาน non-interactive (hermes -z) — เขียนรายงาน JSON "even when the run fails" (v0.19.0)
hermes -z <prompt> --usage-file ./nightly-usage.json
ไฟล์จาก --usage-file มี estimated_cost_usd, token แยกเป็น input/output/cache_read/cache_write/reasoning, api_calls, model, provider, session_id และสถานะ completed/failed — สำหรับ cron job ในองค์กร นี่คือสิ่งที่ควรส่งเข้า log pipeline ทุกคืน เพราะมันคือหลักฐานต่อการรันหนึ่งครั้งที่จับต้องได้ที่สุดที่ Hermes ให้
v0.21.0 เพิ่มสองอย่างที่ผมรอมานาน อย่างแรกคือ status bar ที่แสดง cache hit ratio, latency (ค่าเฉลี่ยเคลื่อนที่ของ 10 call ล่าสุด) และ tok/s ใน CLI แบบสด อย่างที่สองคือ model_overrides ที่ให้เราใส่ราคาต่อล้าน token ของ model ที่ Hermes ยังไม่รู้จัก — สำคัญมากสำหรับบันไดขั้นสามและสี่ เพราะ model หลัง gateway ขององค์กรหรือบน GPU ของเราไม่มีราคาอยู่ใน registry สาธารณะ และถ้าไม่ใส่ ทุกตัวเลขต้นทุนจะเป็นศูนย์ปลอม ๆ:
# config.yaml — v0.21.0
display:
show_cost: true
status_bar:
fields: [cache_hit, latency, tps] # cache_hit reset เมื่อสลับ model หรือ compression; แสดงผลอย่างเดียว
model_overrides:
<model-id>:
context_length: 128000
input_price_per_mtok: 0.50 # ราคาตามสัญญาของเรา (ตัวเลขสมมติ) — ไหลเข้า session cost summary
output_price_per_mtok: 1.50
capabilities: [tool_use, vision]
ทีนี้ถึงตัวเลขที่คนอ่านผิดมากที่สุด — หน้า Analytics ใน web dashboard (token 7/30/90 วัน, cache-hit %, ค่าใช้จ่าย, ตารางต่อ model) ซึ่ง ปิดเป็นค่าเริ่มต้น ด้วยเหตุผลที่เขียนไว้ใน PR ที่ปิดมัน — ไม่ใช่บนหน้าเอกสาร dashboard:
dashboard.show_token_analytics: false พร้อมเหตุผลว่าตัวเลขเป็น "a local lower-bound estimate (they exclude auxiliary calls, retries, fallbacks, and cache writes), so they can read far below the provider bill. Set true only if you understand they're not billing." ถ้อยคำนี้อยู่ใน PR ไม่ใช่บนหน้าเอกสาร dashboard ซึ่งไม่ได้กล่าวถึงค่านี้เลย — และ PR เดียวกันยกกรณีจริงที่ dashboard บอก 150K tokens ขณะที่บิลจริงคือ 27M — ต่างกันเกือบสองร้อยเท่า
ผมอยากให้ทุกคนที่ดูแลงบอ่านประโยคนั้นสองรอบ เพราะสิ่งที่มันตัดออก — auxiliary calls, retries, fallbacks, cache writes — คือทุกอย่างที่เราเพิ่งคุยกันมาสี่หัวข้อ พูดอีกแบบ ตัวเลขใน Hermes บอก รูปร่าง ของค่าใช้จ่าย ส่วน ขนาด ต้องอ่านจากบิลของ provider เท่านั้น เครื่องมือชุมชนอย่าง hermes-dashboard ของ Bichev (proxy + SQLite dashboard ที่อ่าน state.db เพื่อดู cache-hit และค่าใช้จ่ายรายวัน) เกิดขึ้นเพราะช่องว่างนี้ — ผมยกมาเป็นหลักฐานของช่องว่าง ไม่ใช่คำแนะนำให้ติดตั้ง และ hermes sessions cost ที่จะรายงาน cache-hit ต่อ session (PR #76596, 2 สิงหาคม 2026) ยัง ไม่ถูก merge ณ วันที่เขียน
เรื่องที่ต้องพูดให้ชัดที่สุด: Hermes ไม่มีเพดานค่าใช้จ่ายเป็นดอลลาร์ ไม่ว่ารายวันหรือรายเดือน issue #26382 (เปิด 15 พฤษภาคม 2026, ยังเปิดอยู่, P3) ขอ budget.daily_usd_cap/monthly_usd_cap และชี้ว่า /usage กับ /insights ดูได้แค่ย้อนหลัง ส่วน flag --max-budget-usd ที่คู่มือต้นทุนของบุคคลที่สามหลายฉบับยกมาเป็น "option ของ Hermes" นั้น เป็น flag ของ Claude Code — มันปรากฏใน repo แค่ใน SKILL.md ของ skill claude-code ที่แถมมา ไม่มีใน cli-commands.md, configuration.md หรือ FAQ สิ่งที่มีจริงคือรั้วเชิงพฤติกรรม:
# config.yaml — รั้วที่มีจริง (ไม่ใช่เพดานเงิน)
agent:
max_turns: 500 # งบจำนวนรอบต่องาน (ค่าเริ่มต้น 500)
run_budget_seconds: 1800 # แจ้งให้เริ่มสรุปงานเมื่อใช้เวลาไป 80%
# loop_caps: จำกัด web_search / การ spawn subagent ต่อ turn
# hard_stop_enabled: สำหรับ gateway ที่ไม่มีคนเฝ้า
สำหรับองค์กร คำตอบเรื่องเพดานเงินจึงอยู่ นอก Hermes — spend limit ฝั่ง OpenRouter/provider, quota บน gateway ขององค์กรในบันไดขั้นสาม หรือยอดเครดิตของ Portal เอง (Portal ที่เครดิตหมดก็หยุด ซึ่งเป็นเพดานแบบหนึ่ง)
GPU ของเราเอง — vLLM, llama.cpp, Ollama และพื้น 64K
บันไดขั้นที่สี่ตัดค่า token เป็นศูนย์ และเพิ่มปัญหาใหม่สองข้อ: context กับ tool calling ข้อแรกเป็นตัวเลขตายตัว — เอกสาร providers ระบุว่า Hermes "requires at least 64,000 tokens of context for agent use with tools" เพราะ overhead คงที่ที่เราวัดกันไปกินพื้นที่มากอยู่แล้ว ข้อสองคือ server ต้องแปลง tool call ของ model ให้ถูกรูป ซึ่งแต่ละ engine มี flag ของตัวเอง:
# vLLM — parser "hermes" คือรูปแบบ tool call ที่ Hermes Agent ใช้ (ยังมี llama3_json, mistral, deepseek_v3, xlam)
vllm serve <model> \
--enable-auto-tool-choice --tool-call-parser hermes \
--max-model-len 65536
# llama.cpp — "Without --jinja, the server ignores the tools parameter entirely"
llama-server -m <model.gguf> --jinja -fa -c 64000 -ngl 99
# ระวัง: -np แบ่ง context ตาม slot → -c 64000 -np 4 = 16K ต่อ slot ต่ำกว่าพื้น
# SGLang — เพิ่ม --tool-call-parser qwen ตอน launch; เพดาน output เริ่มต้นแค่ 128 tokens
# ยกด้วย --default-max-tokens หรือ model.max_tokens
# LM Studio — provider ชั้นหนึ่งตั้งแต่ v0.12.0 พร้อม doctor checks; endpoint :1234/v1
lms load <model> --context-length 64000
# config.yaml — ชี้ Hermes ไปที่ server ของเรา
model:
provider: custom
base_url: http://vllm.internal:8000/v1
api_key: none
context_length: 65536 # บอก Hermes ตรง ๆ ไม่ต้องรอมันเดา
เรื่อง context length Hermes มีลำดับการหาค่าที่ควรรู้: model.context_length ที่เราตั้ง → cache ถาวร → endpoint /models → Anthropic API → metadata ของ OpenRouter → registry models.dev → ค่าเริ่มต้น 128K สำหรับ model ที่ไม่รู้จัก สำหรับ server ของเราเอง ตั้งค่าแรกไว้ตรง ๆ จะประหยัดเวลาไล่ปัญหาได้มาก
คู่มือทางการ "Run Hermes Locally with Ollama — Zero API Cost" (เพิ่มใน v0.13.0, 7 พฤษภาคม 2026) คือหน้าที่ซื่อสัตย์ที่สุดในเอกสารทั้งชุด เพราะมันบอกตรง ๆ ว่า model เล็กส่วนใหญ่ chat ได้แต่ทำงานไม่ได้: ในตาราง model ของคู่มือ มี gemma4:31b (~20 GB, RAM 24 GB ขึ้นไป) ตัวเดียวที่ tool calling ได้ ส่วน gemma2:27b/9b และ llama3.2:3b "can only chat; they can't take actions" สเปกขั้นต่ำที่ระบุคือ RAM 8 GB / 4 cores แนะนำ 32 GB ขึ้นไป / 8 cores / GPU NVIDIA VRAM 8 GB ขึ้นไป และความเร็วบน CPU ล้วน ๆ อยู่ที่ ~2–5 tokens/วินาที สำหรับ 31B — ใช้ได้กับ cron job กลางคืน ไม่ใช่กับคนนั่งรอ
PARAMETER num_ctx 64000 ใน Modelfile ส่วนทางเลือกระดับ daemon OLLAMA_CONTEXT_LENGTH=64000 มาจาก FAQ ของ Ollama เอง ไม่ได้อยู่ในคู่มือ Hermes และถ้ารันบน CPU ให้ตั้ง HERMES_API_TIMEOUT=1800 ด้วย
# Modelfile — ยก context ให้ถึงพื้นที่ Hermes ต้องการ
FROM gemma4:31b
PARAMETER num_ctx 64000
# หรือระดับ daemon — ตัวแปรนี้มาจาก FAQ ของ Ollama เอง ไม่ได้อยู่ในคู่มือ Hermes
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
# CPU-only: ให้เวลา prefill นาน ๆ
export HERMES_API_TIMEOUT=1800
แล้ว Hermes 4 ของ Nous เอง อยู่ตรงไหนในภาพนี้? ตระกูล Hermes 4 (405B, 70B, 14B) ออกเมื่อ 26 สิงหาคม 2025 และ Hermes-4.3-36B (Apache-2.0, base จาก Seed-OSS-36B, ฝึกด้วย context ถึง 512K, มี GGUF ระดับ 4/5/6/8-bit) ออกเมื่อ 3 ธันวาคม 2025 — แต่ในคลัง Ollama เมื่อ 1 กันยายน 2026 entry ทางการล่าสุดคือ hermes3 ที่อัปเดตเมื่อหนึ่งปีก่อน ไม่มี Hermes 4 หรือ 4.3 แบบทางการ มีแต่ community upload บวกกับคำเตือนของ Nous เองที่ผมยกไว้ในหัวข้อ Portal ว่า 70B/405B ไม่ได้จูนสำหรับ tool call ข้อสรุปของผม: การรัน Hermes Agent บนน้ำหนัก Hermes 4 ในบ้าน ทำได้ ผ่าน GGUF กับ llama.cpp/LM Studio แต่ไม่ใช่เส้นทางที่ทางการรับรอง และคุณต้องทดสอบ tool calling เอง
ที่แย่กว่านั้นคือรายการ "Recommended Local Models" บนหน้า providers (Qwen 2.5 Coder, Llama 3.1, "Hermes 2/3", DeepSeek 67B, Mistral) ขัดกับคู่มือ Ollama บนเว็บเดียวกัน (gemma4:31b) และกับ recipe ของ Red Hat (Qwen2.5-7B-Instruct) — เอกสารทางการยังไม่เห็นตรงกันเอง ดังนั้นกฎเดียวที่ปลอดภัยคือ ทดสอบ tool calling ด้วย workload ของคุณเองก่อนตัดสินใจซื้อ GPU
สำหรับองค์กรที่มี Kubernetes อยู่แล้ว recipe ของ Red Hat Developer (2 มิถุนายน 2026, โดย Gerald Trotman) วาง Hermes บน OpenShift AI กับ KServe vLLM InferenceService, model เริ่มต้น Qwen/Qwen2.5-7B-Instruct, GPU 1 ตัวต่อ replica แล้วชี้ agent ไปที่ OPENAI_BASE_URL=http://hermes-llm-predictor.hermes.svc.cluster.local:8080/v1 — พร้อมประโยคที่ควรติดไว้หน้าห้อง server: "Inference on CPU is 10-100x slower" ฝั่ง AMD ก็มี technical article เรื่อง Hermes + vLLM บน Instinct MI300X ใน AMD Developer Cloud แต่หน้านั้น timeout ตอนผมดึง จึงยังไม่ยก flag ใด ๆ มาอ้าง
สองบันทึกปิดท้าย ข้อแรก — เรื่อง Nemotron ที่หลายคนถามถึง: Nous ไม่ได้ เป็นสมาชิกก่อตั้ง Nemotron Coalition (ประกาศของ NVIDIA เมื่อ 16 มีนาคม 2026 ระบุแปดห้องแล็บโดยไม่มี Nous) แต่เข้าร่วมราวต้นมิถุนายน และเสิร์ฟ nvidia/nemotron-3-ultra:free บน Portal ผ่าน Nebius เฉพาะช่วง 4–18 มิถุนายน 2026 คู่มือหน้านั้นยังค้างอยู่ทั้งที่หมดเขตแล้ว (issue #53729 เปิดอยู่) และ Nemotron SKU ไหนฟรีอยู่ตอนนี้บ้าง ผมยืนยันไม่ได้ ข้อสอง — Hermes ปรับตัวให้ local inference ดีขึ้นเรื่อย ๆ: Ollama ได้ max_tokens เริ่มต้น 65536 ใน v0.17.0, v0.19.0 ข้ามการ probe Ollama สำหรับ provider ที่รู้ว่าไม่ใช่ Ollama (ส่วนหนึ่งของการลด TTFT ~80% จาก ~4.3 วินาทีเหลือ ~0.9 — ตัวเลขนี้มาจาก release tracker ของบุคคลที่สาม ไม่ใช่ release notes ทางการ) และ configuration.md ระบุว่า local endpoint จะ "auto-raise socket and stale timeouts" เพื่อรองรับ prefill ที่ช้า
ข้อมูลไปอยู่ที่ไหน — Telemetry, Privacy Policy และ Air Gap ที่ยังไม่มี
มาถึงอีกครึ่งของคำถามเปิด ตัว runtime ของ Hermes ให้คำตอบที่สะอาด — FAQ ทางการ: "API calls go only to the LLM provider you configure… Hermes Agent does not collect telemetry, usage data, or analytics" และบทสนทนา ความจำ skills อยู่ใน ~/.hermes/ ทั้งหมด แต่ "ไม่มี telemetry" ไม่เท่ากับ "ไม่ออก network เลย" — จากเอกสารเดียวกันและ release notes มีการเรียกออกที่รู้จักสี่อย่าง: ตรวจ update, ค้น registry models.dev (cache ก่อน), ดึง plugin index (cache 24 ชั่วโมง, ถ้า offline ใช้ seed ที่แถมมา) และ header originator: hermes-agent ในคำขอไปยัง Codex ฝ่าย network ขององค์กรควรรู้ทั้งสี่ก่อนเขียน firewall rule
คำถามใหญ่กว่าจึงไม่ใช่ตัว agent แต่คือ Portal ซึ่งเป็นบันไดขั้นแรก และนี่คือหน้าที่ผมอยากให้ทุก DPO อ่านเอง — Privacy Policy ของ Nous Research (ปรับปรุงล่าสุด 11 มิถุนายน 2026) สำหรับ Portal/Cloud ระบุว่า inference payload ถูก เก็บเป็นค่าเริ่มต้น; prompt และ output "may be shared with such providers as necessary to provide the Services"; Nous อาจใช้ข้อมูลเพื่อ "training, fine-tuning" และสร้างหรือ license ข้อมูลรวมหรืออนุพันธ์ได้ ทางออกคือ Privacy Mode ที่เมื่อเปิดแล้ว Nous จะไม่เก็บ payload และไม่ใช้เพื่อ training/product improvement/support ยกเว้นเท่าที่จำเป็นด้านความปลอดภัยและกฎหมาย — และเป็นการ opt-out ไปข้างหน้า ไม่ย้อนหลัง
💡 สิ่งที่นโยบายฉบับนี้ ไม่ ระบุ สำคัญพอ ๆ กับสิ่งที่ระบุ: ไม่มีระยะเวลาเก็บรักษา, ไม่มีที่ตั้งของข้อมูล, ไม่มี zero-data-retention tier, ไม่มีรายชื่อ sub-processor และไม่มี DPA ให้ลงนาม ผมจะไม่อนุมานสิ่งใดในข้อเหล่านี้แทน Nous — และคุณก็ไม่ควร
ฝั่ง software มีของช่วยอยู่สองชิ้น ชิ้นแรกคือ data-training-tier guard ใน v0.21.0 — "a unified selection-guard registry warns you across every picker surface when a model trains on your data" และงานแบบ non-interactive จะ fail closed เว้นแต่ตั้ง security.allow_data_training_tiers_noninteractive: true (เรื่อง secrets และ vault ผมยกไว้ใน #4 Security) ชิ้นที่สองคือ knob ของ OpenRouter:
# config.yaml — กันไม่ให้ route ไปหา provider ที่เก็บข้อมูล (OpenRouter)
provider_routing:
data_collection: deny
sort: price # price | throughput | latency
# only: [...] / ignore: [...] เลือกหรือตัด provider เป็นรายชื่อ
# v0.21.0 — งาน non-interactive จะหยุดเอง ถ้า model ที่เลือกฝึกจากข้อมูลเรา
security:
allow_data_training_tiers_noninteractive: false
แต่ต้องซื่อสัตย์กับข้อจำกัด: issue #32757 "[Bug]: Data privacy" (26 พฤษภาคม 2026) รายงานว่า data_collection: deny ไม่ได้กัน model ที่เก็บ log เมื่อ route ผ่าน Nous Portal issue ถูกปิดแบบ not planned พร้อม label ว่า implemented on main แต่ไม่มีคำตอบจากผู้ดูแลให้เห็น — ผมจัดเป็น คำถามที่ยังเปิด ไม่ใช่บั๊กที่แก้แล้ว
แล้ว Hermes Cloud ล่ะ? หน้า portal.nousresearch.com/cloud บอกว่ายังเป็น preview ต้องมีเครดิต $10 หรือ subscription ที่ active และ "scales to zero when idle — you only pay while it works" ซึ่งเป็นทั้งหมดที่หน้านั้นพูดเรื่องการคิดเงิน (คำว่า "hourly" ไม่ปรากฏบนหน้านั้น) ส่วน FAQ บนหน้านั้นตั้งคำถามเอง "Where are the servers located?" โดยไม่มีคำตอบในเนื้อหาที่ผมดึงได้ — ไม่มีขนาดเครื่อง ราคา หรือ region เผยแพร่ ตัวเลขที่เห็นตามบล็อกอื่นล้วนเป็นของบุคคลที่สาม สำหรับองค์กรไทย ข้อนี้ตัดสินตัวเองได้เลย: ตราบใดที่ตอบไม่ได้ว่า server อยู่ประเทศไหน Cloud ยังไม่ผ่านด่านแรกของ PDPA
ส่วนบันไดขั้นที่ห้า — air-gapped — ยังเป็นคำขอ ไม่ใช่ฟีเจอร์ issue #17696 (เปิด 30 เมษายน 2026, P3, ยังเปิดอยู่) ระบุว่า installer ต้องการ GitHub, source และ public registry สิ่งที่พอทำได้วันนี้คือ wheel บน PyPI (PR #26593 merge 15 พฤษภาคม 2026 — แต่เวอร์ชันบน PyPI ล่าสุด 0.19.0 ตามหลัง GitHub v0.21.0 และตาราง platform ระบุว่า pip install ยังไม่ supported) กับ image ทางการ nousresearch/hermes-agent (ดึงเกิน 5 ล้านครั้ง ขนาดราว 900 MB) ที่ mirror เข้า registry ภายในได้ และอย่าสับสนกับ terminal.docker_network: false:
# config.yaml — sandbox ไม่มี network (--network=none)
terminal:
docker_network: false
# ตัด egress เฉพาะคำสั่งของ agent (terminal, execute_code, file tools)
# ไม่ได้ตัดการเรียก model ของตัว agent เอง — และการสลับค่านี้จะลบ container เดิมที่มี network ทิ้ง
PDPA และ Decision Tree สำหรับองค์กรไทย
ใน #4 ผมเล่าเหตุการณ์ Hermes ในโหมด YOLO ที่วิ่งอยู่ในเครือข่ายกระทรวงการคลังไว้แล้ว เหตุการณ์นั้นเป็นเรื่อง ใครควบคุม agent หัวข้อนี้เป็นอีกด้านของเหรียญเดียวกัน — เมื่อเราเป็นคนควบคุมเอง ข้อมูลที่ agent อ่านไปอยู่ที่ไหน และคำตอบต้องผ่าน PDPA
กรอบกฎหมายที่เกี่ยวโดยตรงคือประกาศลูกภายใต้มาตรา 28 (การส่งข้อมูลไปยังประเทศที่มีมาตรฐานคุ้มครองเพียงพอ) และมาตรา 29 (นโยบายคุ้มครองข้อมูลภายในเครือกิจการ) ซึ่งประกาศในราชกิจจานุเบกษาเมื่อ 25 ธันวาคม 2023 และมีผล 24 มีนาคม 2024 บทสรุปของสำนักกฎหมายหลายแห่งระบุตรงกันว่า ยังไม่มีการประกาศรายชื่อประเทศที่มีมาตรฐานเพียงพอ — ผมยังไม่ได้อ่านตัวประกาศฉบับเต็มจาก pdpc.or.th จึงขอวางไว้เป็นข้อสังเกตจากแหล่งรอง ผลในทางปฏิบัติคือ การส่ง prompt ที่มีข้อมูลส่วนบุคคลออกนอกประเทศต้องพึ่งข้อยกเว้นอื่นตามมาตรา 28 (ซึ่งควรให้ฝ่ายกฎหมายเป็นผู้ชี้) และการอธิบายเรื่องนี้กับ DPO จะง่ายกว่ามากถ้า inference อยู่ในประเทศ
ข่าวดีคือทางเลือกในประเทศมีจริงแล้ว: AWS Asia Pacific (Thailand) ap-southeast-7 เปิดเมื่อ 8 มกราคม 2025 (3 AZ), Google Cloud region กรุงเทพฯ เปิด 21 มกราคม 2026 (3 zones) ส่วน Microsoft ประกาศ region ประเทศไทยแล้วแต่ผมยังไม่พบวันที่ GA สองรายแรกให้เราวางบันไดขั้นสาม (LLM gateway ในประเทศ) หรือขั้นสี่ (vLLM บน GPU ใน region ไทย) ได้โดยที่ Hermes ไม่รู้ความต่างเลย
อีกทางที่น่าจับตาสำหรับงานภาษาไทยคือ Typhoon ของ SCB 10X — endpoint https://api.opentyphoon.ai/v1 เป็น OpenAI-compatible และเอกสารของ Typhoon มีตัวอย่าง tool calling (เช่น typhoon-v2.1-12b-instruct) ส่วน typhoon2.5-qwen3-4b อยู่บน Ollama พร้อม function calling และ context 256K บนกระดาษจึงเป็น custom provider ของ Hermes ได้:
# config.yaml — Typhoon เป็น named custom provider (ยังไม่ผ่านการทดสอบกับ Hermes)
providers:
typhoon:
api: https://api.opentyphoon.ai/v1
transport: chat_completions
key_env: TYPHOON_API_KEY
/model custom:typhoon:typhoon-v2.1-12b-instruct
ย้ำว่า ยังไม่มีใครทดสอบ Typhoon กับ Hermes อย่างเป็นทางการ ที่ตั้ง server และราคาก็ไม่อยู่ในหน้าเอกสารที่ผมดึงได้ และพื้น 64K ยังใช้บังคับ — นี่คือการทดลองที่ผมตั้งใจจะทำเอง ไม่ใช่คำแนะนำ
บริบทที่กำลังมาคือ ร่าง พ.ร.บ. AI ที่ ETDA เปิดรับฟังความเห็นเมื่อ 9 กรกฎาคม 2026 ถึง 14 สิงหาคม 2026 เป็นกรอบแบบ risk-based และบทวิเคราะห์ของ Baker McKenzie คาดว่าจะใช้เวลา 2–3 ปีกว่าจะบังคับใช้ — ยังไม่ผูกพัน แต่การเลือกบันไดขั้นที่ควบคุมข้อมูลได้ตั้งแต่วันนี้ คือการซื้อประกันราคาถูกสำหรับวันที่มันบังคับใช้
รวมทุกอย่างเป็นตารางตัดสินใจหนึ่งตาราง — ไม่ใช่คำตอบสำเร็จรูป แต่เป็นคำถามที่ควรถามเรียงลำดับ:
| ข้อมูลที่ agent จะเห็น | ขั้นบันไดที่ผมแนะนำ | เหตุผล |
|---|---|---|
| สาธารณะ / ไม่มีข้อมูลส่วนบุคคล (สรุปเอกสารเปิด, โค้ดบน repo สาธารณะ) | ขั้น 1 Portal (เปิด Privacy Mode) หรือขั้น 2 BYO | เร็ว ถูก และไม่มีอะไรให้ PDPA คุ้มครอง |
| ข้อมูลภายในองค์กร ไม่มีข้อมูลส่วนบุคคล | ขั้น 2 ที่ตั้ง data_collection: deny หรือขั้น 3 | ตัดทอดของ Nous ออก เหลือสัญญากับ provider เดียว |
| มีข้อมูลส่วนบุคคลของนักศึกษา / ลูกค้า / บุคลากร | ขั้น 3 gateway ใน region ไทย หรือขั้น 4 | inference ไม่ออกนอกประเทศ คำถามมาตรา 28 จึงไม่เกิดตั้งแต่ต้น |
| ข้อมูลอ่อนไหว / ข้อมูลราชการชั้นความลับ | ขั้น 4 บน hardware ของเราเอง — และยอมรับว่าขั้น 5 ยังไม่มี | ไม่มี provider ให้ไว้ใจ และ Hermes เองยังต้องออก network เพื่อติดตั้ง/อัปเดต |
จุดยืนของผม ณ 1 กันยายน 2026: สำหรับทีมวิจัยหรือทีม dev ที่ข้อมูลไม่มีชื่อคน Portal ยังเป็นบันไดที่คุ้มที่สุด — เปิด Privacy Mode, ตั้ง auxiliary เป็น model ฟรี, ดู /usage ให้เป็นนิสัย สำหรับหน่วยงานที่แตะข้อมูลส่วนบุคคล คำตอบที่ยั่งยืนคือขั้นสาม — LLM gateway ขององค์กรใน region ไทย ที่ Hermes มองเห็นเป็น provider ธรรมดาตัวหนึ่ง — เพราะมันย้ายคำถาม PDPA ทั้งหมดไปอยู่ในสัญญาฉบับเดียวที่ฝ่ายกฎหมายอ่านออก และปล่อยให้ Hermes ทำสิ่งที่มันเก่ง: สลับ model ได้ทุกเมื่อโดยไม่มี lock-in
🎯 สิ่งสำคัญที่ต้องจำ
- บันไดห้าขั้น = Portal → BYO keys → custom provider หลัง gateway องค์กร → self-hosted → air-gapped (ขั้นสุดท้ายยังเป็น issue #17696 ที่เปิดอยู่)
- Portal Free = "Free models only, $0 monthly credits" — Tool Gateway มาพร้อม tier จ่ายเงินและเปิดทีละ tool; "10%" มีสองตัวที่คนละความหมาย (โบนัสเครดิต $20 → $22 กับส่วนลด 10% สำหรับ token-billed providers) ส่วน 20% ของ Teknium ยืนยันไม่ได้นอก X
- Hermes 4 ≠ Hermes Agent = Nous เองบอกว่า 70B/405B ไม่ได้จูนสำหรับ tool call — ใช้ผ่าน subscription proxy กับแอป chat ไม่ใช่กับ agent
- Pool → Fallback → Auxiliary = สามชั้นที่ทุกชั้นเก็บค่าผ่านทางเป็นการอ่านประวัติใหม่เต็มราคา — session ยาวให้เปิด session ใหม่แทนการเด้งไปมา
- hermes prompt-size = วัด overhead คงที่ของตัวเองแบบ offline ก่อนเชื่อตัวเลข 73% จาก issue #4379 ที่ผูกกับ v0.6.0
- system_and_3 = breakpoint สี่จุด อ้าง ~75% — แตกเมื่อสลับ model, fallback, หมุน key และ compression; provider ที่ cache ได้มีมากกว่าที่ developer guide เขียน
- Dashboard Analytics = "local lower-bound estimate" ปิดเป็นค่าเริ่มต้นเพราะเคยบอก 150K เทียบบิลจริง 27M — ไม่มีเพดานดอลลาร์ใน Hermes และ
--max-budget-usdเป็นของ Claude Code - พื้น 64K = context ขั้นต่ำสำหรับงาน tool; Ollama เริ่มต้นที่ 2,048 ตามคู่มือ Hermes (4,096 ตาม FAQ ของ Ollama) — และ gemma4:31b คือ model เดียวในคู่มือทางการที่ tool calling ได้
- PDPA = Portal เก็บ payload เป็นค่าเริ่มต้นและไม่ระบุที่ตั้ง server; ข้อมูลส่วนบุคคลควรอยู่บันไดขั้นสามขึ้นไปใน region ไทย (AWS ap-southeast-7, Google Cloud Bangkok)
Picture the end of the month. The provider's invoice arrives, and the number on it is several times what the agent's own dashboard had been telling you. You open /insights and see one figure; you open the provider's billing page and see another. Which one is the truth? And the question underneath that one — while the agent was working through the night, whose servers did your organization's data sleep on?
Those two questions — what does it cost and where does the data go — are more entangled than they look, because every step up or down the model ladder, from a single Nous Portal bill to GPUs in your own server room, trades cost against control at the same time. I have touched money twice already in this series: the Portal tier table and the "MIT-free, pay only for the model" framing live in #1 Hermes 101, and a brief section on cost control and BYO keys versus Portal sits in #7 Production. This post is the full treatment of both.
I will walk the five rungs of the ladder, the three layers of fallback, where the tokens actually go, the economics of the prompt cache, the cost figures Hermes reports (and what they cannot see), running open-weight models on your own GPUs, and then close with the question no Thai organization gets to skip — PDPA. Every number here comes from the official docs, the release notes, the issue tracker or the Portal pages as I fetched them on September 1, 2026. Where the sources disagree with each other, I will tell you how they disagree rather than pick the nicer figure.
The Model Ladder: Five Rungs an Organization Climbs
When I talk to a department weighing Hermes, the first question is usually "Portal, or our own keys?" — which is too narrow a question. Hermes lets you stand on any of five rungs, and each rung answers three questions differently: who holds the key, where the data travels, and who sends the bill.
- Nous Portal — one OAuth, one bill. The README promises "300+ models — pick any of them with /model <name>", and a paid subscription unlocks the Tool Gateway (web, image, TTS, cloud browser). The fastest rung, and the only one that adds Nous as an extra hop in the data path.
- BYO keys — an OpenRouter key or a direct provider key (Anthropic, OpenAI and so on). Bills scatter across key owners, but the data path is one hop shorter.
- A named custom provider behind your corporate gateway — declare a provider of your own in config, pointing at an internal LLM gateway (LiteLLM, the IT department's proxy, or a Bedrock contract you already hold), authenticated with short-lived SSO tokens.
- Self-hosted inference — vLLM, llama.cpp, SGLang, LM Studio or Ollama on your own GPUs. Token cost zero; electricity and staff cost not zero.
- Air-gapped — the rung many government agencies want, and the one Hermes has not officially reached (more on that in the data section).
What makes the ladder usable is that every rung is switched with the same commands. The configuring-models doc draws the line clearly: hermes model is the out-of-session wizard that adds providers, while /model inside a session only switches among providers already configured — with three scopes:
# This session only — reverts when the session ends
/model <name> --provider <p>
# Persist to config.yaml
/model <name> --provider <p> --global
# One turn, then auto-restore
/model <name> --once
# A model_aliases entry lets a short name expand to provider + model
/model fav
One thing gateway operators must remember: gateway sessions pick up config changes only for new sessions unless the gateway is restarted. If you changed the model and the Telegram bot is still answering on the old one, that is the first thing to suspect.
The third rung is where mid-sized and larger organizations tend to settle, and the docs support it better than I expected. providers.<name> takes api, key_env and a transport of either chat_completions or anthropic_messages; key_cmd runs a CLI of yours to mint a short-lived token from enterprise SSO and refreshes it before expiry:
# config.yaml — a provider of your own, behind the corporate gateway
providers:
corp-gateway:
api: https://llm-gateway.internal.example.ac.th/v1
transport: chat_completions # or anthropic_messages
key_cmd: "my-auth-cli print-token --profile prod" # short-lived SSO token, auto-refreshed
# Switch to it in a session
/model custom:corp-gateway:<model>
And since v0.21.0 (August 31, 2026), third parties can ship providers as pip packages that Hermes discovers through entry points — which means a platform team can package the internal provider once and distribute it to every machine without editing configs one by one (a thread that continues into #10 Desktop & Fleet).
Portal as a Routing Tier, Not a Religion
In #7 I declined to say what the Free tier includes, because the docs page and the pricing page disagreed. The pricing page at portal.nousresearch.com, fetched September 1, 2026, now settles it: Free is "Free models only" with $0 in monthly credits. The four-tier table itself (Free $0 · Plus $20 → $22 credits · Super $100 → $110 · Ultra $200 → $220) is already printed in #1 Hermes 101, so I will not repeat it — what #1 does not carry is the following:
- Rollover caps — unused credits carry into the next month, but only up to $10 (Plus) / $50 (Super) / $100 (Ultra). An organization with bursty usage — heavy at the end of term, quiet over the break — quietly forfeits whatever sits above that cap in every light month.
- The paid cards add "Hosted tool usage" and "High rate limits" (Free gets "Standard rate limits").
- Top-ups — pay-as-you-go credit top-ups come in $10, $20, $50, $100 or $200.
- Copy that contradicts itself — the tier cards say "200+ Models" while the catalog copy on the same page says "over 300" and the README says "300+". I go with 300+, as the README does, and note that the cards have not caught up. The same page prices the catalog at "free to $24 per million input tokens".
The "10%" needs untangling, because there are two of them and they mean different things. The first is the credit bonus on the pricing page — "10% bonus" means you pay $20 and receive $22 in credits (and $100 → $110, $200 → $220). The second is the provider discount in the configuring-models doc — "Portal subscribers also get 10% off token-billed providers" — a lower per-token price when you route through the Portal, not extra credit. A third figure — Teknium (a Nous co-founder) wrote "20% Discount on all models" in his X post for the Herald release — exists only on X, appears on neither the Portal pages nor the docs, and I cannot verify whether it still applies or what it ever applied to. I will not print the 20% as settled — log into the Portal and read the live number before it goes into a budget.
What a paid subscription genuinely adds is the Tool Gateway. The doc says it is "included with every paid Nous Portal subscription", while Free-tier accounts "can use Portal for inference but don't include managed tools" (with the caveat that "some accounts are also entitled to a free tool pool"). Nous manages five backends: Firecrawl for web, FAL for image generation, OpenAI TTS, Browser Use for a cloud browser, and a Modal cloud terminal as an optional add-on — and, the part I like, it is "opt-in per tool, not all-or-nothing":
# config.yaml — adopt Nous-managed tools one at a time
web:
backend: nous # Firecrawl via the Portal
image_gen:
provider: nous # FAL
tts:
provider: nous # OpenAI TTS
stt:
provider: nous
browser:
cloud_provider: nous # Browser Use — only when you want it
Since v0.19.0 (July 20, 2026) the plan can be managed from inside the TUI/CLI with /subscription and /topup — neither is in the user guide's slash-command list yet, so I am citing the release notes — and hermes portal info shows login, subscription and routing status. Under the hood Hermes mints a short-lived JWT from your stored Portal refresh token on every inference call.
Now the question people get wrong most often: "Nous makes the Hermes 4 models — why not run Hermes Agent on Hermes 4 and get the cheapest possible loop?" The answer comes from Nous itself, in the run-hermes-with-nous-portal guide:
💡 "Hermes-4-70B and Hermes-4-405B are available on the Portal at deep discounts, but they're chat/reasoning models, not tool-call-tuned. They will struggle with multi-step agent loops… For Hermes Agent itself, stick to the frontier agentic models." Same name; not the same job.
Where Hermes 4 does fit is the subscription proxy. hermes proxy start exposes a local OpenAI-compatible server (/v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models) so that non-agent apps such as Open WebUI can spend the same subscription on chat and research work; only nous and xAI are shipped today:
# Start the proxy on this machine, then point other apps at localhost
hermes proxy start
A closing footnote for Claude users, which affects readers of both this series and the OpenClaw series: the Anthropic OAuth path in Hermes requires a "Claude Max subscription + purchased extra usage credits" and consumes only those overage credits, while "Claude Pro subscribers cannot use the OAuth path; the supported alternative is an ANTHROPIC_API_KEY", billed per token. If you set up OpenClaw on Claude Pro earlier this year, your second rung is an API key, not a subscription.
Three Layers of Resilience: Pools, Fallbacks, Auxiliaries
In #7 I summarized the fallback chain in one line as "the way out of a 429." The full picture has three stacked layers, and the fallback-providers doc keeps them distinct: credential pools rotate several keys for the same provider, fallback_providers switch to a different provider:model, and each auxiliary task resolves its provider independently. Pools are always tried before fallbacks.
The first layer — the fallback chain, shipped in v0.6.0 (March 30, 2026) — is a top-level YAML list, with a management command since v0.12.0:
# config.yaml — the order to try when the primary fails
fallback_providers:
- provider: openrouter
model: <provider/model>
- provider: nous
model: <model>
# Or manage it from the CLI without touching the file
hermes fallback list
hermes fallback add <provider> <model>
hermes fallback remove <provider> <model>
The legacy singular fallback_model still works, but the list wins when both exist, and there is deliberately no environment variable — the doc wants this controlled by YAML and CLI only. What you need to know precisely is what triggers it:
| Situation | Behaviour |
|---|---|
| HTTP 429 · 500 · 502 · 503 | retries are exhausted first, then fallback |
| HTTP 401 · 403 · 404 | fallback immediately |
| repeated malformed responses | fallback |
| a transient 429 with a known reset window | no fallback — Hermes waits |
| scope | per turn — every new user message starts on the primary again |
The second layer — credential pools, from v0.7.0 (April 3, 2026, which Nous called "the resilience release") — supports OpenRouter, Anthropic, Nous Portal and custom OpenAI-compatible endpoints:
# Add a second key for the same provider to its pool
hermes auth add <provider>
hermes auth list
hermes auth reset # clear cooldown state
# config.yaml — rotation strategy per provider
credential_pool_strategies:
openrouter: round_robin # fill_first (default) | round_robin | least_used | random
The rotation rules are finer-grained than you might expect: a 429 is retried once, and only a second consecutive 429 rotates the key (one-hour cooldown); a 402 rotates immediately (one-hour cooldown); a 401 tries a token refresh first, then rotates (five-minute cooldown); and only when every key in the pool is exhausted does the request fall through to the fallback chain. All of this state lives in ~/.hermes/auth.json — one more file that belongs in the backup routine from #7.
The third layer — auxiliary tasks — is the background work that is not the main agent. The configuring-models page enumerates eleven slots: title generation, vision, compression, approval scoring, web extraction, skills hub, MCP tool routing, triage specifier, Kanban decomposer, profile describer and curator (in the fallback-providers doc the same slots appear as YAML keys such as vision, compression, skills_hub, mcp, approval, title_generation, triage_specifier). The default provider: auto means the main model first, and since v0.15.0 (May 28, 2026) there is a layered ladder: main model → auxiliary.<task>.fallback_chain → fallback_providers → the built-in discovery chain (OpenRouter → Nous Portal → custom endpoint → Codex OAuth → API-key providers → give up).
# config.yaml — background work on a free/cheap model, separate from the main one (shape follows the fallback-providers doc's example)
auxiliary:
compression:
provider: auto # try the main model first
model: ""
fallback_chain:
- provider: openrouter
model: inclusionai/ring-2.6-1t:free # the doc's own :free example
vision:
provider: auto # use the main model if it can see
This is the documented "zero-cost auxiliary" pattern: v0.8.0 (April 8, 2026) added a free Xiaomi MiMo v2 Pro on the Portal for compression, vision and summarization, and the fallback-providers doc's own example puts inclusionai/ring-2.6-1t:free on OpenRouter in the compression fallback_chain (its other example sets google/gemini-3-flash-preview as the compression model — not a :free one). Two caveats. Compression that cannot find a model will "degrade to no-summary" — quietly, with no summary produced. And subagents inherit the parent's whole chain; the cost argument for delegation lives in #2 Agent Teams.
Every layer charges the same toll, and the doc says so in plain words: a fallback or a credential rotation resets the cache keys, so "the next request re-reads the entire history at full input-token price" — and again when you switch back. I will return to that in the cache section.
Where the Tokens Go: The Fixed Cost of Every Call
Before talking about price per million tokens, you need to know what each request actually sends, because an agent like Hermes never sends just the user's message — it attaches the system prompt, the schema of every enabled tool, the skills index, memory and profile, every single time.
The number that made this concrete comes from issue #4379 (opened April 1, 2026 by a user named Bichev; still open, P2, labelled area/usage-cost), which measured at v0.6.0 a fixed per-call overhead of roughly 13,935 tokens — about 73% of a typical request — broken down as system prompt ~5,176, 31 tool definitions ~8,759, skills index ~2,200. One evening of Telegram, WhatsApp and cron gateways consumed ~3.9 million tokens across 207 API calls.
hermes prompt-size, which runs offline and costs nothing.
# The fixed per-message budget — system prompt, skills index, memory, profile, tool schemas
# "Runs entirely offline — no API call" (v0.16.0, tagged 2026.6.5)
hermes prompt-size
hermes prompt-size --platform telegram --json # per platform, because toolsets differ
# Estimated token cost per toolset
hermes tools
The diet the official FAQ prescribes is simple: measure first, then disable the toolsets and skills you do not use. Since v0.11.0 each cron job can carry its own enabled_toolsets to "cap token overhead + cost per job" — a morning news digest does not need to carry the browser-automation schemas (the scheduling side is in #8 Automation).
The bigger instrument is Tool Search: an opt-in mode that replaces the schemas of MCP tools and non-core plugin tools with three bridge tools — tool_search / tool_describe / tool_call — and lets the model look up the tool it needs when it needs it, instead of carrying every schema on every turn (progressive disclosure). The key rule: "Built-in Hermes tools never defer" — the built-ins are always present, so the benefit is largest for organizations that attach many MCP servers, as described in #5 Integrations.
The next lever the configuration doc names explicitly as a cost lever is reasoning effort, settable at the agent level and per auxiliary task:
# config.yaml — thinking depth is money
agent:
reasoning_effort: medium # none | minimal | low | medium | high | max (v0.19.0 added the top tiers, max and ultra)
auxiliary:
compression:
reasoning_effort: low # summarizing does not need deep thought
vision:
reasoning_effort: none # reading an image needs no reasoning
Nous has also been trimming its own hidden costs release by release, which tells you where money used to leak: v0.15.0 rebuilt session_search without an LLM — the release notes say the old version "cost ~$0.30/call and took ~30 seconds to summarize three sessions, sometimes confabulating"; v0.17.0 made the skill Curator's consolidation pass opt-in (curator.consolidate: true) so routine runs incur "no aux-model spend"; and v0.18.0 routed the post-turn self-improvement fork to an auxiliary model that digests context instead of replaying the whole conversation. The rule I draw from all of it: every background job is a model call, and every model call is money — make sure you know which background jobs are on.
Prompt-Cache Economics
If the fixed overhead is the problem, the prompt cache is half of the answer — because the part that repeats every turn (system prompt, tool schemas) is exactly the part that caches best. The mechanics per the developer guide: caching is enabled automatically when the model name is an Anthropic Claude model and the provider supports cache_control; the strategy is system_and_3, which places four breakpoints — the system prompt plus the last three non-system messages — and the doc claims it "Reduces input token costs by ~75% on multi-turn conversations".
# config.yaml
prompt_caching:
cache_ttl: "5m" # only "5m" (default) or "1h" are accepted — configurable since v0.12.0
Three things break the cache, and two of them are things we just configured in the previous sections: mid-history edits, a change of model identity (/model, an automatic fallback, a credential rotation), and compression — where the system-prompt cache survives and the rest re-establishes within one or two turns. The official tips page condenses it into a rule I want every team to memorize:
💡 "an explicit /model switch, an automatic provider fallback, or a credential-pool rotation all force the next turn to re-read the entire conversation at full input price" — and in long sessions "it's often cheaper to start a fresh session on the other model than to bounce back and forth".
The natural follow-up is "which providers actually cache?", and this is where the docs lag the release notes. The developer guide still says Anthropic native or OpenRouter only, but configuration.md already adds Nous Portal, and the release notes add several more: xAI via the x-grok-conv-id header (v0.8.0), the Bedrock Converse API via cachePoint and DeepSeek on OpenCode gateways (v0.20.0), Claude through LiteLLM on the OpenAI wire (v0.20.2/v0.21.0), and api.meta.ai via the Responses API (v0.21.0). v0.20.0 also made tool schemas cacheable on native Anthropic "without history loss". Put differently: if you stand on rung three, behind a corporate gateway, the odds that caching engages are better than the main doc page suggests — but confirm it yourself by watching cache_read in /usage.
The other half of the equation is compression, whose mechanics I covered in #3 Memory; today I add only the cost dimension. The defaults per the developer guide:
# Compression defaults (developer guide) — worked example for a 200K model
compression:
threshold: 0.50 # compress at 50% of the context window → 100,000 tokens
target_ratio: 0.20 # share of the tail that is kept → tail budget 20,000
tail_mode: lean # default since v0.20.6: 2.5% of context, 10K floor / 25K cap
# the gateway keeps a safety net at 0.85
# summary budget = content × 0.20, min 2,000, max min(context × 0.05, 12,000) → 10,000
Two financial consequences. First, every compression is one call to the auxiliary compression model with nearly the whole session as input, and the doc requires that model to have "a context window at least as large as the main agent model's" — so "use a small cheap model to compress" only works when the small model also has a large enough window. Second, compression breaks the history part of the cache, so the turn after it is always more expensive than usual. If /usage spikes periodically for no visible reason, check whether the spikes line up with compression events.
/compress reports a 120 s timeout while the background worker succeeds minutes later; large-session compression fails with lease lost / session_split_failed). If you rely on compression to hold down the cost of long sessions, check both before upgrading, and confirm in /usage that a summary was in fact produced.
Caching is also why prefix caching matters so much once you run models yourself. An independent developer's recipe on dev.to (Qwen3.6-27B + vLLM on a 24 GB RTX 3090, May 2026 — third-party throughout) measured time-to-first-token on a large prompt at 38.6 seconds cold versus 1.59 seconds with the prefix cache warm, a difference of more than twenty times — and that "large prompt" is precisely the fixed Hermes overhead we just measured.
What the Cost Numbers Actually Measure
Hermes has several cost-reporting surfaces, and each measures something different. From the nearest outward:
# In a session
/usage # this session's tokens + rate-limit headers and cost detail (v0.9.0),
# account limits (v0.11.0), per-category context breakdown (v0.18.0)
/insights --days 7 # the retrospective view
# Outside a session
hermes insights --days 30 --source telegram
# Non-interactive runs (hermes -z) — a JSON report written "even when the run fails" (v0.19.0)
hermes -z <prompt> --usage-file ./nightly-usage.json
The --usage-file report carries estimated_cost_usd; input, output, cache_read, cache_write and reasoning tokens; api_calls; model; provider; session_id; and a completed/failed status. For an organization's cron jobs, this is what should flow into the log pipeline every night — it is the most tangible per-run evidence Hermes produces.
v0.21.0 added two things I had been waiting for. The first is a status bar that shows the cache-hit ratio, latency (a rolling mean of the last ten calls) and tokens per second live in the CLI. The second is model_overrides, which lets you enter per-million-token prices for models Hermes does not know — essential on rungs three and four, because a model behind your corporate gateway or on your own GPU has no price in any public registry, and without an override every cost figure is a false zero:
# config.yaml — v0.21.0
display:
show_cost: true
status_bar:
fields: [cache_hit, latency, tps] # cache_hit resets on model switch and compression; display only
model_overrides:
<model-id>:
context_length: 128000
input_price_per_mtok: 0.50 # your contracted rate (illustrative) — flows into session cost summaries
output_price_per_mtok: 1.50
capabilities: [tool_use, vision]
Now the figure people misread most: the Analytics page in the web dashboard (7/30/90-day tokens, cache-hit %, cost, per-model table), which is off by default for a reason spelled out in the PR that turned it off — not on the dashboard docs page:
dashboard.show_token_analytics: false on the grounds that the figures are "a local lower-bound estimate (they exclude auxiliary calls, retries, fallbacks, and cache writes), so they can read far below the provider bill. Set true only if you understand they're not billing." That wording lives in the PR, not on the dashboard docs page, which does not mention the setting at all — and the same PR cites a real case where the dashboard reported 150K tokens against an actual bill of 27M, nearly two hundred times apart.
I want everyone who owns a budget to read that sentence twice, because what it excludes — auxiliary calls, retries, fallbacks, cache writes — is everything the last four sections were about. Put differently, Hermes's numbers tell you the shape of your spend; the size can only be read off the provider's bill. Community tooling such as Bichev's hermes-dashboard (a proxy plus a SQLite dashboard that reads state.db for cache-hit rates and daily cost) exists precisely because of this gap — I cite it as evidence of the gap, not as a recommendation to install it. And hermes sessions cost, which would report cache hits per session (PR #76596, August 2, 2026), is not merged as I write.
The thing that must be said most plainly: Hermes has no spend cap in dollars, daily or monthly. Issue #26382 (opened May 15, 2026; still open, P3) asks for budget.daily_usd_cap/monthly_usd_cap and points out that /usage and /insights only look backwards. The --max-budget-usd flag that several third-party cost guides list as a "Hermes option" is a Claude Code flag — it appears in the repo only inside the bundled claude-code skill's SKILL.md, and not in cli-commands.md, configuration.md or the FAQ. What does exist are behavioural guardrails:
# config.yaml — the guardrails that actually exist (not a dollar cap)
agent:
max_turns: 500 # iteration budget per task (default 500)
run_budget_seconds: 1800 # wrap-up notice at 80% of the budget
# loop_caps: cap web_search calls / subagent spawns per turn
# hard_stop_enabled: for unattended gateways
For an organization, the dollar ceiling therefore lives outside Hermes — a spend limit on the OpenRouter or provider side, a quota on the corporate gateway of rung three, or the Portal's own credit balance (a Portal that runs out of credit stops, which is a ceiling of a kind).
Bring Your Own GPUs: vLLM, llama.cpp, Ollama and the 64K Floor
The fourth rung takes the token cost to zero and adds two new problems: context and tool calling. The first is a hard number — the providers doc says Hermes "requires at least 64,000 tokens of context for agent use with tools", because the fixed overhead we measured already eats a large share. The second is that the server must parse the model's tool calls correctly, and every engine has its own flags:
# vLLM — the "hermes" parser is the tool-call format Hermes Agent speaks (others: llama3_json, mistral, deepseek_v3, xlam)
vllm serve <model> \
--enable-auto-tool-choice --tool-call-parser hermes \
--max-model-len 65536
# llama.cpp — "Without --jinja, the server ignores the tools parameter entirely"
llama-server -m <model.gguf> --jinja -fa -c 64000 -ngl 99
# careful: -np divides the context across slots → -c 64000 -np 4 = 16K per slot, below the floor
# SGLang — add --tool-call-parser qwen at launch; the default output cap is only 128 tokens,
# raise it with --default-max-tokens or model.max_tokens
# LM Studio — a first-class provider with doctor checks since v0.12.0; endpoint :1234/v1
lms load <model> --context-length 64000
# config.yaml — point Hermes at your server
model:
provider: custom
base_url: http://vllm.internal:8000/v1
api_key: none
context_length: 65536 # tell Hermes directly instead of letting it guess
On context length, Hermes resolves the value in a documented order: your model.context_length → the persistent cache → the endpoint's /models → the Anthropic API → OpenRouter metadata → the models.dev registry → a 128K default for unknown models. For your own server, setting the first one explicitly saves a lot of debugging.
The official guide "Run Hermes Locally with Ollama — Zero API Cost" (added in v0.13.0, May 7, 2026) is the most honest page in the whole documentation set, because it says outright that most small models can chat but cannot act. In its model table, gemma4:31b (~20 GB, 24+ GB RAM) is the only entry with tool calling; gemma2:27b/9b and llama3.2:3b "can only chat; they can't take actions". The stated minimum is 8 GB RAM / 4 cores, the recommendation 32+ GB / 8+ cores / an NVIDIA GPU with 8+ GB VRAM, and CPU-only speed for the 31B is ~2–5 tokens per second — fine for an overnight cron job, not for a person waiting.
PARAMETER num_ctx 64000 in a Modelfile; the daemon-level alternative, OLLAMA_CONTEXT_LENGTH=64000, comes from Ollama's own FAQ, not from the Hermes guide. On CPU-only hardware set HERMES_API_TIMEOUT=1800 as well.
# Modelfile — raise the context to the floor Hermes needs
FROM gemma4:31b
PARAMETER num_ctx 64000
# or at the daemon level — this variable comes from Ollama's own FAQ, not the Hermes guide
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
# CPU-only: allow long prefill
export HERMES_API_TIMEOUT=1800
So where does Nous's own Hermes 4 fit in this picture? The Hermes 4 family (405B, 70B, 14B) was released on August 26, 2025, and Hermes-4.3-36B (Apache-2.0, based on Seed-OSS-36B, trained with context up to 512K, with GGUFs at 4/5/6/8-bit) on December 3, 2025 — but in the Ollama library on September 1, 2026 the newest official entry is hermes3, updated a year ago. There is no official Hermes 4 or 4.3 entry, only community uploads, on top of Nous's own warning, quoted in the Portal section, that the 70B/405B are not tool-call-tuned. My conclusion: running Hermes Agent on Hermes 4 weights at home is possible via a GGUF with llama.cpp or LM Studio, but it is not an officially endorsed path, and you must test tool calling yourself.
Worse, the "Recommended Local Models" list on the providers page (Qwen 2.5 Coder, Llama 3.1, "Hermes 2/3", DeepSeek 67B, Mistral) contradicts the Ollama guide on the same site (gemma4:31b) and Red Hat's recipe (Qwen2.5-7B-Instruct) — the official docs do not agree with themselves, so the only safe rule is test tool calling on your own workload before buying a GPU.
For organizations already running Kubernetes, the Red Hat Developer recipe (June 2, 2026, by Gerald Trotman) places Hermes on OpenShift AI with a KServe vLLM InferenceService, default model Qwen/Qwen2.5-7B-Instruct, one GPU per replica, and points the agent at OPENAI_BASE_URL=http://hermes-llm-predictor.hermes.svc.cluster.local:8080/v1 — with a sentence worth taping to the server-room door: "Inference on CPU is 10-100x slower". AMD has a technical article on Hermes plus vLLM on an Instinct MI300X in the AMD Developer Cloud as well, but the page timed out when I fetched it, so I quote no flags from it.
Two closing notes. First, on the Nemotron question people keep asking: Nous was not a founding member of the Nemotron Coalition (NVIDIA's March 16, 2026 announcement names eight labs, Nous not among them) but joined around early June and served nvidia/nemotron-3-ultra:free on the Portal through Nebius only from June 4 to June 18, 2026. That guide is still live past its expiry (issue #53729 open), and I cannot verify whether any Nemotron SKU is free today. Second, Hermes keeps adapting to local inference: Ollama got a default max_tokens of 65536 in v0.17.0, v0.19.0 skips the Ollama probe for providers known not to be Ollama (part of the ~80% TTFT cut, from ~4.3 s to ~0.9 s — a figure from third-party release trackers, not the official release notes), and configuration.md says local endpoints "auto-raise socket and stale timeouts" to accommodate slow prefill.
Where the Data Goes: Telemetry, the Privacy Policy and the Air Gap That Isn't
Now the other half of the opening question. The Hermes runtime gives a clean answer — the official FAQ: "API calls go only to the LLM provider you configure… Hermes Agent does not collect telemetry, usage data, or analytics", with conversations, memory and skills all under ~/.hermes/. But "no telemetry" is not "no network egress at all": from the same doc and the release notes, four known outbound calls remain — the update check, models.dev registry lookups (cache-first), the plugin index fetch (24-hour cache, falling back to a bundled seed offline), and an originator: hermes-agent header on Codex requests. Your network team should know all four before writing firewall rules.
The bigger question, then, is not the agent but the Portal, which is rung one — and this is the page I want every DPO to read for themselves. The Nous Research Privacy Policy (last updated June 11, 2026) for Portal and Cloud states that inference payloads are stored by default; that prompts and outputs "may be shared with such providers as necessary to provide the Services"; that Nous may use data for "training, fine-tuning" and may generate or license aggregated or derivative data. The way out is Privacy Mode: once enabled, Nous will not store inference payloads and will not use them for training, product improvement or support, except to the limited extent necessary for security and legal purposes — and it is a prospective opt-out, not a retroactive one.
💡 What this policy does not state matters as much as what it does: no retention period, no data location, no zero-data-retention tier, no sub-processor list, and no DPA to sign. I will not infer any of those on Nous's behalf — and neither should you.
On the software side there are two helpers. The first is the data-training-tier guard in v0.21.0 — "a unified selection-guard registry warns you across every picker surface when a model trains on your data", and non-interactive workloads fail closed unless you set security.allow_data_training_tiers_noninteractive: true (secrets and vault integration are covered in #4 Security). The second is OpenRouter's routing knob:
# config.yaml — refuse routes to data-collecting providers (OpenRouter)
provider_routing:
data_collection: deny
sort: price # price | throughput | latency
# only: [...] / ignore: [...] to allow-list or block providers by name
# v0.21.0 — non-interactive runs stop themselves if the chosen model trains on your data
security:
allow_data_training_tiers_noninteractive: false
But honesty about the limits: issue #32757 "[Bug]: Data privacy" (May 26, 2026) reported that data_collection: deny did not prevent a data-logging model when routed via Nous Portal. It was closed as not planned with a label saying implemented on main, and no maintainer reply is visible — I file it as an open question, not a fixed bug.
And Hermes Cloud? The portal.nousresearch.com/cloud page says it is in preview, requires $10 in credits or an active subscription, and "scales to zero when idle — you only pay while it works" — which is the whole of what the page says about billing (the word "hourly" does not appear on it); its own FAQ asks "Where are the servers located?" and the content I could fetch gives no answer. No machine sizes, prices or regions are published — the figures on other blogs are all third-party. For a Thai organization this one decides itself: until someone can say which country the servers are in, Cloud does not clear the first PDPA gate.
As for the fifth rung — air-gapped — it is still a request, not a feature. Issue #17696 (opened April 30, 2026, P3, still open) notes that the installer needs GitHub, source and public registries. What can be done today: the PyPI wheel (PR #26593, merged May 15, 2026 — though the latest PyPI release, 0.19.0, lags GitHub's v0.21.0, and the platform matrix lists pip installs as unsupported), and the official nousresearch/hermes-agent image (5M+ pulls, roughly 900 MB), which can be mirrored into an internal registry. And do not confuse any of this with terminal.docker_network: false:
# config.yaml — a sandbox with no network (--network=none)
terminal:
docker_network: false
# cuts egress only for agent commands (terminal, execute_code, file tools)
# NOT for the agent process's own model calls — and flipping it removes the existing networked container
PDPA and a Decision Tree for Thai Organizations
In #4 I told the story of a Hermes instance in YOLO mode running inside the Ministry of Finance network. That incident was about who controls the agent. This section is the other side of the same coin — when we are the ones in control, where does the data the agent reads end up, and the answer has to pass through PDPA.
The directly relevant law is the pair of sub-regulations under section 28 (transfers to countries with adequate protection) and section 29 (binding corporate rules), gazetted on December 25, 2023 and in force since March 24, 2024. Several law-firm summaries agree that no adequacy list has been published — I have not read the full notification text on pdpc.or.th myself, so I flag that as a secondary-source observation. The practical effect is that sending prompts containing personal data abroad must rest on one of the other section 28 exceptions (which your legal team, not this post, should identify), and that conversation with your DPO is far easier when inference stays in the country.
The good news is that in-country options now exist: AWS Asia Pacific (Thailand) ap-southeast-7 launched on January 8, 2025 (3 AZs); Google Cloud's Bangkok region launched on January 21, 2026 (3 zones); Microsoft has announced a Thailand region, but I found no GA date. The first two let you place rung three (an in-country LLM gateway) or rung four (vLLM on GPUs in a Thai region) without Hermes noticing any difference.
Another option worth watching for Thai-language work is Typhoon from SCB 10X: the endpoint https://api.opentyphoon.ai/v1 is OpenAI-compatible, Typhoon's docs include tool-calling examples (e.g. typhoon-v2.1-12b-instruct), and typhoon2.5-qwen3-4b is on Ollama with function calling and a 256K context. On paper, that makes it a Hermes custom provider:
# config.yaml — Typhoon as a named custom provider (untested with Hermes)
providers:
typhoon:
api: https://api.opentyphoon.ai/v1
transport: chat_completions
key_env: TYPHOON_API_KEY
/model custom:typhoon:typhoon-v2.1-12b-instruct
To be clear: nobody has tested Typhoon with Hermes officially, its hosting location and pricing were not on the pages I fetched, and the 64K floor still applies — this is an experiment I intend to run myself, not a recommendation.
The context on the horizon is Thailand's draft AI Act, which ETDA opened for consultation on July 9, 2026 through August 14, 2026 — a risk-based framework that Baker McKenzie's analysis expects to take two to three years to enact. Not binding yet; but choosing a rung that keeps the data under your control today is cheap insurance for the day it is.
Everything above folds into one decision table — not a ready-made answer, but the questions to ask, in order:
| What the agent will see | The rung I recommend | Why |
|---|---|---|
| Public / no personal data (summarizing open documents, code in public repos) | Rung 1 Portal (Privacy Mode on) or rung 2 BYO | fast, cheap, and nothing for PDPA to protect |
| Internal data, no personal data | Rung 2 with data_collection: deny, or rung 3 | removes the Nous hop; one contract with one provider |
| Personal data of students / customers / staff | Rung 3 gateway in a Thai region, or rung 4 | inference never leaves the country, so the section 28 question never arises |
| Sensitive data / classified government data | Rung 4 on your own hardware — accepting that rung 5 does not exist yet | no provider to trust, and Hermes itself still needs the network to install and update |
My position as of September 1, 2026: for a research or development team whose data carries no names, the Portal is still the best-value rung — turn on Privacy Mode, put the auxiliaries on a free model, and make /usage a habit. For any unit that touches personal data, the durable answer is rung three — a corporate LLM gateway in a Thai region that Hermes sees as just another provider — because it moves the entire PDPA question into a single contract your legal team can read, and leaves Hermes doing what it is good at: switching models at will, with no lock-in.
🎯 Key Takeaways
- The five-rung ladder = Portal → BYO keys → a custom provider behind the corporate gateway → self-hosted → air-gapped (the last rung is still open issue #17696)
- Portal Free = "Free models only, $0 monthly credits" — the Tool Gateway comes with paid tiers and is enabled per tool; there are two different "10%"s (the $20 → $22 credit bonus and the 10% off token-billed providers), and Teknium's 20% is unverifiable outside X
- Hermes 4 ≠ Hermes Agent = Nous itself says the 70B/405B are not tool-call-tuned — use them through the subscription proxy from chat apps, not under the agent
- Pool → Fallback → Auxiliary = three layers, each charging the same toll of re-reading the history at full price — on long sessions start a fresh session instead of bouncing
- hermes prompt-size = measure your own fixed overhead offline before trusting the 73% figure from issue #4379, which is pinned to v0.6.0
- system_and_3 = four breakpoints, ~75% claimed — broken by a model switch, a fallback, a key rotation and compression; more providers cache than the developer guide lists
- Dashboard Analytics = a "local lower-bound estimate", off by default after it read 150K against a 27M bill — Hermes has no dollar cap, and
--max-budget-usdbelongs to Claude Code - The 64K floor = the minimum context for tool work; Ollama defaults to 2,048 per the Hermes guide (4,096 per Ollama's own FAQ) — and gemma4:31b is the only model in the official guide that can call tools
- PDPA = the Portal stores payloads by default and names no server location; personal data belongs on rung three or higher in a Thai region (AWS ap-southeast-7, Google Cloud Bangkok)