Supplier Governance FinOps Sustainable AI

Suppliers, Cost and Footprint — outsource ความรับผิดชอบไม่ได้Suppliers, Cost and Footprint — Outsourcing Technology Does Not Outsource Accountability

Outsourcing เทคโนโลยีไม่ได้ outsource ความรับผิดชอบ — 11 ข้อสัญญากับ supplier รวมทางออก ต้นทุนรวมต่อผลลัพธ์ที่สำเร็จ และ footprint ที่ต้องวัดหลังการใช้งานเพิ่มขึ้น ไม่ใช่ต่อ tokenOutsourcing technology does not outsource accountability — eleven supplier contract clauses including exit, total cost per successful outcome, and a footprint measured after usage growth rather than per token.

By Anirach Mingkhwan AI Transformation for Organizations 2026 • Post #17 28 min read
Suppliers, Cost and Footprint — outsource ความรับผิดชอบไม่ได้
ในบทความนี้
  1. ทำไม outsource เทคโนโลยีแล้วความรับผิดชอบยังอยู่ที่เรา และหลักฐาน release หมดอายุได้เงียบ ๆ
  2. เวิร์กช็อป 11 ข้อสัญญากับ supplier — สิ่งที่ต้องขอ หลักฐานที่ได้ และอะไรพังถ้าไม่มี
  3. ช่อง Economics ของกระดานคะแนน — ต้นทุนรวมต่อ Successful Outcome ไม่ใช่ต่อ token
  4. อ่านตัวเลข IEA ให้ถูก 485 → 950 TWh ห้าสิ่งที่ต้องวัด และกับดัก Rebound
  5. กฎ routing โมเดลเล็กที่สุดที่ผ่านเกณฑ์ และตารางออกแบบการวัดที่มีเจ้าของทุกช่อง
  6. ตัวชี้วัดสำคัญพร้อมช่อง Scorecard และรูปแบบความล้มเหลวเจ็ดแบบ
  7. ก้าวต่อไป — จากสัญญาและมิเตอร์ สู่ 90 วันแรกที่องค์กรทำได้จริง
In this post
  1. Why outsourcing technology leaves accountability with us, and how release evidence expires in silence
  2. Worksheet: eleven supplier clauses — what to ask for, the evidence it feeds, and what breaks if it is missing
  3. The board scorecard's Economics column — total cost per successful outcome, not per token
  4. Reading the IEA numbers correctly, 485 → 950 TWh, the five things to measure, and the rebound trap
  5. The smallest-adequate-model routing rule, and a measurement design table with an owner in every row
  6. The metrics that matter with their Scorecard column, and seven failure patterns
  7. The road ahead — from contracts and meters to a first 90 days an organisation can actually walk

🤔 ถ้า vendor เปลี่ยนโมเดลคืนนี้โดยไม่บอก หลักฐาน release ของคุณเมื่อวานยังใช้ได้ไหม?

ตอนที่แล้ว One Evidence System จบลงที่ข้อเสนอว่าองค์กรควรผูกภาระผูกพันจากหลายกรอบเข้ากับหลักฐานชุดเดียว แทนที่จะสร้างแฟ้มใหม่ทุกครั้งที่มีกฎใหม่ ตอนนี้คือคำถามที่ตามมาทันทีและมักไม่มีใครถามในห้องประชุมเดียวกัน — ถ้าหลักฐานส่วนใหญ่ในแฟ้มนั้นมาจากระบบที่เราไม่ได้เป็นเจ้าของ เราจะกำกับมันด้วยอะไร และเมื่อระบบนั้นเปลี่ยนตัวเองในคืนวันอังคาร ใครเป็นคนรู้

คำตอบสั้น ๆ ของทั้งบทความคือประโยคเดียวจากบทที่ 10 ของหนังสือ AI Transformation as an Organizational Core: "Outsourcing technology does not outsource accountability." — การจ้างเทคโนโลยีไม่ใช่การโอนความรับผิด สิ่งที่ย้ายไปกับสัญญาคืองาน ไม่ใช่ ความรับผิดรับชอบ (accountability) และการทำให้ข้อเท็จจริงนี้กำกับได้จริงต้องอาศัยสามอย่างพร้อมกัน คือ สัญญาที่ครอบคลุม 11 หัวข้อรวมทางออก ตัวหารที่ถูกต้องของต้นทุน และมิเตอร์วัด footprint ที่วัดผลรวมหลังการใช้งานโต ไม่ใช่วัดต่อ token

1. คืนที่ vendor เปลี่ยนโมเดล

ลองนึกภาพตามนี้ วันจันทร์ทีมของคุณผ่าน ด่านอนุมัติการนำระบบออกใช้ (release gate) เรียบร้อย มีชุดทดสอบครบ มีผลบนกรณีขอบ มีสถิติของกรณีรุนแรง มีลายเซ็นเจ้าของงาน แฟ้มหลักฐานหนาพอที่จะวางบนโต๊ะผู้ตรวจได้ทันที คืนวันอังคารผู้ให้บริการโมเดลปรับรุ่นเบื้องหลัง endpoint เดิม ชื่อรุ่นใน API ไม่เปลี่ยน ราคาไม่เปลี่ยน เอกสารไม่เปลี่ยน วันพุธระบบของคุณยังทำงาน ยังตอบ ยังผ่าน dashboard ทุกเส้น สิ่งเดียวที่เปลี่ยนคือ แฟ้มหลักฐานที่คุณเซ็นไปเมื่อวันจันทร์ อธิบายระบบที่ไม่มีอยู่แล้ว

ผมย้ำก่อนไปต่อว่านี่ไม่ใช่รายงานเหตุการณ์จริงที่ผมมีชื่อบริษัทหรือวันที่ในมือ มันคือการทำให้ประโยคหนึ่งของหนังสือเห็นภาพขึ้น และประโยคนั้นตรงไปตรงมามาก บทที่ 9 เขียนไว้ว่า "Production is the final evaluation environment, not the end of evaluation. Supplier model changes, new user behavior, or policy updates can invalidate previous evidence." — คู่มือภาษาไทยของหนังสือเล่มเดียวกันมีประโยคนี้ว่า "Production คือสภาพประเมินสุดท้าย ไม่ใช่จุดจบ เพราะ Supplier, User Behavior หรือ Policy เปลี่ยนแล้ว Evidence เดิมอาจใช้ไม่ได้"

ความน่ากลัวของกลไกนี้ไม่ได้อยู่ที่ความรุนแรง แต่อยู่ที่ความเงียบ ระบบไม่ล่ม ไม่มี alert ไม่มีใครโทรมาตอนตีสาม สิ่งที่หมดอายุคือ คำกล่าวอ้าง ว่าเราเคยพิสูจน์อะไรไว้ และคำกล่าวอ้างไม่เคยส่งเสียงตอนมันตาย มันจะเงียบไปเรื่อย ๆ จนถึงวันที่มีคนถาม — ผู้ตรวจ ผู้กำกับ ลูกค้าที่ถูกปฏิเสธคำขอ หรือทนายของอีกฝ่าย — แล้วเราถึงจะพบว่าเอกสารที่เราถืออยู่ในมือพูดถึงสิ่งที่ต่างจากสิ่งที่รันอยู่จริง

บทที่ 10 ของหนังสือจัดพฤติกรรมที่นำไปสู่สภาพนี้ไว้ในรายการ รูปแบบความล้มเหลว ด้วยถ้อยคำสั้นที่สุดเท่าที่จะสั้นได้ คือ "assuming the vendor owns risk" หรือในคู่มือภาษาไทยว่า "คิดว่า Vendor รับความเสี่ยงแทน" และบทที่ 9 เติมอีกข้อที่คู่กันเสมอคือ "treating vendor testing as sufficient" — "เชื่อ Vendor Test" สองข้อนี้ไม่ใช่ความประมาทของคนขี้เกียจ มันคือข้อสรุปที่ สมเหตุสมผล ถ้าเราคิดว่าสิ่งที่เราซื้อคือ "บริการ" แบบเดียวกับที่เราซื้อไฟฟ้าหรือพื้นที่เก็บข้อมูล ปัญหาคือสิ่งที่เราซื้อไม่ใช่แบบนั้น เราซื้อองค์ประกอบหนึ่งที่จะเข้าไปมีส่วนในการตัดสินใจที่มีคนรับผลจริง และหน้าที่ตอบคำถามว่าการตัดสินใจนั้นถูกต้องหรือไม่ ไม่เคยย้ายไปกับใบแจ้งหนี้

หนังสือให้นิยาม ความรับผิดรับชอบ ไว้ในภาคผนวกคำศัพท์แบบที่ปิดทางออกทุกทาง: "การมีเจ้าของที่ชัดเจนต่อการตัดสินใจและผลกระทบ พร้อมการชี้แจง หลักฐาน การยกระดับปัญหา การเยียวยา และผลตามความรับผิด โดยไม่โยนความรับผิดให้โมเดล" — ประโยคสุดท้ายมักถูกอ่านผ่าน แต่มันคือหัวใจ ถ้าโยนให้โมเดลไม่ได้ ก็โยนให้เจ้าของโมเดลไม่ได้เช่นกัน

💡 มุมมองของผม: ในบรรดาหลักปฏิบัติห้าประการของบทที่ 10 ข้อที่ผมคิดว่าเปลี่ยนวิธีทำงานได้ทันทีที่สุดคือข้อ 4 — "Make supplier change and exit governable · Vendor dependence is part of the risk perimeter" หรือในคู่มือภาษาไทย "ทำให้ Supplier Change และ Exit กำกับได้ Vendor Dependence อยู่ในขอบเขตความเสี่ยง" คำที่ทำงานหนักที่สุดคือ risk perimeter องค์กรส่วนใหญ่วาดขอบเขตความเสี่ยงไว้ที่ขอบระบบของตัวเอง แล้วเขียนกำกับไว้ตรงขอบว่า "ส่วนนี้เป็นของ vendor" ประโยคนี้บอกว่าเส้นนั้นวาดผิดที่ ความพึ่งพา vendor อยู่ข้างใน ขอบเขต ไม่ใช่ข้างนอก

อีกสี่ข้อที่เหลือของบทนี้ทำงานรอบข้อ 4 เสมอ ข้อ 1 Maintain a live inventory and named owner — Governance เริ่มจากรู้ว่ามีอะไร ซึ่งแปลว่า Inventory ต้องบันทึกด้วยว่าแต่ละระบบแขวนอยู่บนผู้ให้บริการรายใด ข้อ 2 Map obligations to evidence — ใช้ Artifact ร่วมอย่างรับผิดชอบโดยไม่อ้างว่ากรอบเหมือนกัน ข้อ 3 Use law as the floor and ethics for the gap — ระบุ Prohibited Use กลุ่มผลกระทบ และ Remedy และข้อ 5 Measure lifecycle sustainability and rebound — ประสิทธิภาพต่อหน่วยอาจดีขึ้นขณะผลรวมสูงขึ้น ซึ่งเป็นหัวข้อของ ส่วนที่ 4 ของบทความนี้ทั้งส่วน

คำถามเดียวที่ผมใช้ทดสอบทุกสัญญา: ถ้าผู้ให้บริการเปลี่ยนรุ่นโมเดลคืนนี้ องค์กรของเราจะรู้เรื่องนี้จาก อะไร — จากหนังสือแจ้งตามสัญญา จากการเฝ้าดูของเราเอง หรือจากการที่ตัวเลขในรายงานเดือนหน้าดูแปลก ๆ คำตอบข้อที่สามคือคำตอบที่องค์กรส่วนใหญ่ให้จริง และมันหมายความว่าเราไม่ได้กำกับอะไรเลย เราแค่กำลังเฝ้ารอ

2. สิบเอ็ดข้อสัญญา — และข้อที่สิบเอ็ดคือทางออก

บทที่ 10 ไม่ได้ปล่อยเรื่องนี้ไว้เป็นหลักการลอย ๆ มันไล่หัวข้อสัญญาไว้เป็นรายการปิด 11 ข้อ[1] ด้วยประโยคเดียว: "Govern suppliers through contracts covering data use, customer-data training, model and subprocessor changes, evaluation evidence, logs, security tests, incident notice, audit support, continuity, intellectual property, and exit." คู่มือภาษาไทยเขียนไว้ว่า "กำกับ Supplier ผ่านสัญญาที่ครอบคลุม Data Use การฝึกด้วยข้อมูลลูกค้า การเปลี่ยน Model หรือ Subprocessor, Evaluation Evidence, Log, Security Test, Incident Notice, Audit Support, Continuity, IP และ Exit"

ก่อนจะลงตาราง ขอวางเงื่อนไขให้ชัดหนึ่งครั้ง เพราะเนื้อหาส่วนนี้อ่านคล้ายการร่างสัญญาและมันไม่ใช่ — หนังสือระบุชัดว่านี่ไม่ใช่คำแนะนำทางกฎหมาย การใช้งานจริงต้องผ่านการทบทวนตามเขตอำนาจและอุตสาหกรรม รายการ 11 ข้อนี้เป็นกรอบตั้งต้นสำหรับตั้งคำถาม ไม่ใช่ถ้อยคำที่เอาไปวางในสัญญาได้ทันที และหนังสือย้ำเองว่า "Role and system classification require qualified legal analysis" — การจำแนกบทบาทและระบบต้องใช้คำปรึกษากฎหมาย

ตารางข้างล่างคือเวิร์กช็อปที่ผมใช้เดินกับทีมจัดซื้อและทีมกฎหมายพร้อมกันในห้องเดียว ครึ่งวันต่อผู้ให้บริการหนึ่งราย คอลัมน์ที่สองคือสิ่งที่ต้องขอ คอลัมน์ที่สามคือหลักฐานที่คำขอนั้นจะกลายไปเป็น และคอลัมน์ที่สี่ — คอลัมน์ที่ทำให้การประชุมจบเร็วขึ้นมาก — คือสิ่งที่พังถ้าช่องนั้นว่าง

Clause สิ่งที่ต้องขอ หลักฐานที่ได้ อะไรพังถ้าไม่มี
1 · Data use ขอบเขตการใช้ข้อมูลของเรา ระยะเวลาเก็บ ภูมิภาคที่ประมวลผล และรายชื่อผู้ที่เข้าถึงได้ บันทึกฐานทางกฎหมายและ Lineage ของข้อมูลใน Inventory ไม่มีใครตอบผู้กำกับได้ว่าข้อมูลลูกค้าไปอยู่ที่ใดและถูกใช้ทำอะไร
2 · Customer-data training ข้อห้ามหรือเงื่อนไขการนำข้อมูลลูกค้าไปฝึกโมเดล พร้อมกลไกปฏิเสธที่ตรวจสอบได้ ไม่ใช่แค่ช่องติ๊กในหน้าตั้งค่า หนังสือยืนยันจากผู้ให้บริการ คู่กับผลตรวจการตั้งค่าที่เราทำเอง ข้อมูลลูกค้าไหลเข้าโมเดลรุ่นถัดไปของผู้ให้บริการ โดยไม่มีใครรู้และย้อนกลับไม่ได้
3 · Model and subprocessor changes การแจ้งล่วงหน้าเมื่อเปลี่ยนรุ่นโมเดล เปลี่ยนพฤติกรรมเบื้องหลัง endpoint เดิม หรือเพิ่ม Subprocessor พร้อมสิทธิ์คงรุ่นเดิมชั่วคราว บันทึกการเปลี่ยนรุ่นที่ผูกกับ Manifest ของ Release แต่ละรอบ หลักฐานที่ผ่าน gate เมื่อวาน อธิบายระบบที่ไม่มีอยู่แล้ววันนี้ — คือฉากใน ส่วนที่ 1
4 · Evaluation evidence ผล การประเมินระบบ (evaluation) ที่ระบุชุดทดสอบ ข้อจำกัด และเงื่อนไขที่ผลนั้นใช้ไม่ได้ แฟ้มหลักฐานที่เข้าไปเป็นส่วนหนึ่งของ Release Dossier ของเรา เรามีแต่ผลทดสอบของผู้ให้บริการในบริบทของผู้ให้บริการ ซึ่งไม่ใช่บริบทของเรา
5 · Logs สิทธิ์เข้าถึง Log ระดับ request ที่มีเวลา รหัสรุ่น และข้อมูลพอจะสร้างเหตุการณ์ย้อนกลับได้ Trace ที่ใช้สืบสวนเหตุการณ์และตอบข้อร้องเรียนรายกรณี เมื่อเกิดเรื่อง เราสืบได้แค่ถึงขอบ API แล้วมืดสนิท
6 · Security tests รายงานผลทดสอบความมั่นคง ขอบเขตที่ทดสอบ และวันที่ พร้อมสิทธิ์ให้เราหรือบุคคลที่สามทดสอบส่วนที่เราเปิดใช้ หลักฐานสำหรับด่านความมั่นคงของ Release Gate คำว่า "ปลอดภัย" ในสไลด์ขายกลายเป็นหลักฐานชิ้นเดียวที่เรามี
7 · Incident notice นิยามของเหตุการณ์ กรอบเวลาแจ้ง ช่องทาง และข้อมูลขั้นต่ำที่ต้องได้รับในการแจ้งครั้งแรก จุดเริ่มนับเวลาของวงจรเรียนรู้จากเหตุการณ์ผิดปกติ เรารู้เรื่องจากข่าวหรือจากลูกค้า ซึ่งช้ากว่ากรอบเวลาที่กฎหมายกำหนดไว้เสมอ
8 · Audit support ความร่วมมือในการตรวจ ขอบเขต ความถี่ และรูปแบบหลักฐานที่ยอมรับ รวมถึงรายงานของบุคคลที่สาม ชุดหลักฐานที่ผู้ตรวจภายในและผู้กำกับใช้ได้โดยไม่ต้องขอเป็นรายครั้ง การตรวจกลายเป็นการขอความกรุณา และตารางเวลาขึ้นกับผู้ให้บริการ
9 · Continuity ระดับบริการ ความพร้อมใช้ แผนสำรอง และพฤติกรรมของระบบเมื่อบริการไม่ตอบ แผน Fallback ที่ผ่านการซ้อมจริง ไม่ใช่ย่อหน้าในเอกสาร บริการล่มแล้วธุรกิจล่มตาม เพราะไม่เคยมีใครเดินเส้นทางที่ไม่ผ่านโมเดล
10 · Intellectual property สิทธิ์ในผลลัพธ์ ข้อมูลนำเข้า Prompt และ Artifact ที่เราสร้าง รวมถึงความรับผิดเมื่อถูกกล่าวหาว่าละเมิด เอกสารสิทธิ์ที่ฝ่ายกฎหมายหยิบไปใช้ได้จริงเมื่อมีข้อพิพาท ข้อพิพาทเรื่องผลงานเกิดขึ้นหลังจากผลลัพธ์ถูกส่งออกไปแล้วหลายหมื่นชิ้น
11 · Exit สิทธิ์นำข้อมูลออกในรูปแบบที่ใช้ต่อได้ กรอบเวลา การลบข้อมูลที่พิสูจน์ได้ และเงื่อนไขที่เราบอกเลิกได้ แผนทางออกที่มีเจ้าของ มีวันที่ และมีตัวเลขต้นทุนที่ประเมินแล้ว เราไม่ได้เลือกผู้ให้บริการอีกต่อไป ผู้ให้บริการเป็นฝ่ายเลือกเรา

ข้อ 11 คือข้อที่หายไปบ่อยที่สุด และหายไปด้วยเหตุผลที่เข้าใจได้ ตอนเซ็นสัญญาไม่มีใครอยากคุยเรื่องเลิกสัญญา มันดูเหมือนการไม่ไว้ใจคู่ค้าตั้งแต่วันแรก บทที่ 3 ของหนังสือแก้ปัญหาการเมืองข้อนี้ด้วยการเปลี่ยนกรอบให้เป็นเรื่องงบประมาณแทน หลักปฏิบัติข้อ 5 ของบทนั้นเขียนว่า "Fund scale and exit together — every investment needs fallback rollback and retirement conditions" หรือ "ให้งบขยายมาพร้อมงบทางออก ทุกข้อเสนอต้องมี Fallback, Rollback และเงื่อนไขยุติ" เมื่อทางออกเป็นบรรทัดหนึ่งในงบ มันก็ไม่ใช่การไม่ไว้ใจใครอีกต่อไป มันคือรูปแบบมาตรฐานของเอกสารขออนุมัติ

ความกระจุกตัวที่ Provider เป็นความเสี่ยงระดับพอร์ต

สัญญาที่ดีที่สุด 11 ข้อครบทุกข้อ ยังตอบคำถามหนึ่งไม่ได้ เพราะคำถามนั้นไม่ได้อยู่ในระดับสัญญา บทที่ 3 ใส่ไว้ในรายการตัวชี้วัดของ พอร์ตโฟลิโอการตัดสินใจ (decision portfolio) ด้วยคำสองคำคือ provider concentration — คู่มือภาษาไทยเรียกว่า "ความกระจุกตัวที่ Provider" และจับคู่ไว้กับรูปแบบความล้มเหลวข้อหนึ่งของบทเดียวกันคือ "มองข้ามความเสี่ยงจากโมเดลเดียวทั้งพอร์ต"

กลไกที่ทำให้ความเสี่ยงนี้รอดสายตาทุกกระบวนการอนุมัติ น่าสนใจกว่าตัวความเสี่ยงเสียอีก ทุกทีมประเมินระบบของตัวเองแยกกัน แต่ละทีมสรุปได้อย่างซื่อสัตย์ว่า "ความเสี่ยงของระบบเรายอมรับได้" และทุกข้อสรุปนั้นถูกต้องในระดับของมัน แต่ไม่มีเวทีไหนถามคำถามที่เห็นได้จากมุมพอร์ตเท่านั้น — ถ้างานที่องค์กรขาดไม่ได้ทุกงานอยู่บนผู้ให้บริการรายเดียวกัน ความเสี่ยงระดับองค์กรคือเท่าไร ความกระจุกตัวเป็นคุณสมบัติของ ผลรวม ไม่ใช่ของชิ้นส่วน จึงมองไม่เห็นจากรายงานของชิ้นส่วนใด ๆ ไม่ว่าจะเขียนดีแค่ไหน

ผู้ให้บริการโมเดลมีหน้าที่ของตัวเองตามกฎหมายด้วย

สัญญาไม่ใช่เครื่องมือเดียว ในสหภาพยุโรป กฎหมายกำหนดหน้าที่บางอย่างไว้ที่ผู้ให้บริการโมเดลโดยตรง ตรวจสอบข้อความทางการเมื่อ 5 กันยายน 2026: EU AI Act มาตรา 53 กำหนดให้ผู้ให้บริการ general-purpose AI model จัดทำและปรับปรุงเอกสารทางเทคนิคของโมเดล รวมถึงกระบวนการฝึกและทดสอบกับผลการประเมิน และต้องส่งข้อมูลกับเอกสารให้ผู้พัฒนาระบบที่จะนำโมเดลไปประกอบ เพื่อให้เข้าใจขีดความสามารถและข้อจำกัดของโมเดลได้จริง[2] — และมาตรานี้ไม่ถูกแก้โดยการแก้ไขปี 2026

สองข้อควรระวังที่ต้องเดินทางไปกับย่อหน้าข้างบนเสมอ หนึ่ง กฎหมายฉบับนี้ไม่ได้ผูกหน้าที่ไว้กับคำว่า "supplier" แต่ผูกกับบทบาทที่นิยามไว้ — provider, deployer, importer, distributor และผู้ให้บริการ general-purpose AI model — องค์กรของคุณอาจเป็นบทบาทใดบทบาทหนึ่งในนั้นเองด้วย และการจำแนกบทบาทต้องใช้คำปรึกษากฎหมาย สอง กำหนดเวลาเพิ่งขยับ ตรวจสอบเมื่อ 5 กันยายน 2026: EU AI Act มีผลใช้บังคับทั่วไปตั้งแต่ 2 สิงหาคม 2026 ส่วนกำหนดของระบบ high-risk ถูกเลื่อนโดย Regulation (EU) 2026/1744 ลงวันที่ 8 กรกฎาคม 2026 มีผล 27 กรกฎาคม 2026 ไปเป็น 2 ธันวาคม 2027 สำหรับระบบตาม Annex III และ 2 สิงหาคม 2028 สำหรับระบบตาม Annex I[3] การเลื่อนกำหนดไม่ได้ลบหน้าที่ มันแค่ย้ายวันที่ และองค์กรที่รอวันที่นั้นค่อยเริ่มขอเอกสารจากผู้ให้บริการ จะเริ่มช้ากว่ารอบสัญญาของตัวเองหลายรอบ

ในระดับมาตรฐานและกรอบ หนังสือใช้ถ้อยคำที่ผมคิดว่าควรลอกไปทั้งประโยค: "Several frameworks can share evidence without being treated as interchangeable" — หลายกรอบใช้หลักฐานร่วมกันได้แต่ไม่ใช่สิ่งเดียวกัน NIST AI RMF ให้ Lifecycle Function, ISO/IEC 42001 กำหนดระบบการจัดการ AI ซึ่ง ณ 5 กันยายน 2026 ยังเป็นมาตรฐานที่เผยแพร่แล้วตั้งแต่ธันวาคม 2023[4] ส่วน ISO/IEC 42005 เน้น การประเมินผลกระทบ (impact assessment) ตลอดวงจร และ OECD AI Principles ซึ่งรับรองครั้งแรกพฤษภาคม 2019 ปรับปรุงพฤษภาคม 2024 ขอให้ผู้เกี่ยวข้อง "apply a systematic risk management approach to each phase of the AI system lifecycle on an ongoing basis" พร้อมกับดูแล traceability ตลอดวงจรชีวิตของระบบ[5]

กับดักที่หนังสือเรียกชื่อตรง ๆ: รูปแบบความล้มเหลวข้อสุดท้ายของบทที่ 10 คือ "describing guidance or certification as proof of legal compliance" — "อ้าง Guideline หรือ Certificate ว่าเท่ากับ Legal Compliance" ในบริบทของบทความนี้มันแปลตรงตัวว่า ใบรับรองของผู้ให้บริการเป็นหลักฐานว่าเขามีระบบการจัดการแบบหนึ่ง ไม่ใช่หลักฐานว่าระบบของ เรา ถูกกฎหมายในเขตอำนาจของเรา และไม่ใช่หลักฐานว่าผลลัพธ์ที่ระบบของเราส่งถึงลูกค้าคนหนึ่งเมื่อวานนี้ถูกต้อง สองสิ่งนี้อยู่คนละชั้นของข้อโต้แย้ง

3. ช่อง Economics — ตัวหารที่ถูกต้องของต้นทุน

คู่มือผู้บริหารฉบับย่อของหนังสือวางกระดานคะแนนสำหรับคณะกรรมการไว้หกช่อง คือ Value, Quality, Risk, People, Learning และ Economics โดยช่อง Economics เขียนไว้สั้นมากเพียงบรรทัดเดียว: "Total cost per successful outcome and reusable capability share" — ต้นทุนรวมต่อผลลัพธ์ที่สำเร็จ และสัดส่วนความสามารถที่นำกลับมาใช้ซ้ำได้ คู่มือภาษาไทยย่อไว้ว่า "Economics จากต้นทุนต่อ Successful Outcome"

และประโยคที่ต้องเดินทางไปกับกระดานนี้ทุกครั้งคือ "No single composite score should replace this view. A faster process with rising severe errors is not progress. A safe system that produces no outcome value is not transformation. A productive workflow that exhausts reviewers is not sustainable." — ห้ามยุบหกช่องเป็นคะแนนเดียว กระบวนการที่เร็วขึ้นพร้อมความผิดพลาดรุนแรงที่มากขึ้นไม่ใช่ความก้าวหน้า ระบบที่ปลอดภัยแต่ไม่สร้างคุณค่าไม่ใช่การเปลี่ยนผ่าน และกระบวนงานที่ผลิตได้มากแต่ทำให้ผู้ตรวจหมดแรงไม่ยั่งยืน

คำสำคัญของช่องนี้อยู่ที่คำว่า ต่อ ไม่ใช่คำว่า ต้นทุน ตัวเศษเป็นเรื่องที่ทุกคนถกเถียงกันสนุก แต่สิ่งที่ตัดสินว่ารายงานฉบับนั้นบอกความจริงหรือไม่คือตัวหาร และตัวหารที่อุตสาหกรรมใช้โดยปริยายคือ token FinOps Foundation ซึ่งเป็นแหล่งอ้างอิงกลางของวิชาชีพนี้ ระบุ KPI ด้านต้นทุน AI ไว้เป็นต้นทุนต่อ token ต่อ inference และต่อ API call[6] ตัวเลขเหล่านี้ไม่ได้ผิด มันเป็นตัวเลขที่จำเป็นสำหรับวิศวกรที่ต้องปรับระบบ แต่มันตอบคำถามของคณะกรรมการไม่ได้ เพราะต้นทุนต่อ token ที่ลดลง ขณะที่จำนวนงานซึ่งจบได้จริงลดลงด้วย คือรายงานที่ตัวเลขดีขึ้นและธุรกิจแย่ลงพร้อมกัน และไม่มีอะไรในรายงานนั้นที่ผิด

ที่ต้องพูดให้ตรง: หน่วย "ต้นทุนรวมต่อ Successful Outcome" เป็นของหนังสือ ไม่ใช่ของ FinOps Foundation ผมยกแหล่งนั้นมาเพื่อแสดงว่ามิเตอร์เริ่มต้นของอุตสาหกรรมคือ token จริง ๆ ไม่ใช่ภาพจำที่ผมสร้างขึ้นเอง และเพราะมันเป็นมิเตอร์เริ่มต้น องค์กรจึงต้องออกแรงเปลี่ยนมันอย่างตั้งใจ ไม่มีทางที่ตัวหารที่ถูกต้องจะโผล่มาเองในรายงาน

คำว่า "รวม" มีห้าบรรทัด

ตัวเศษก็ถูกตัดทอนได้เหมือนกัน บทที่ 4 ขยายคำว่า "ต้นทุนรวม" ไว้ครบทุกบรรทัดในรายการตัวชี้วัดของมัน คือ "value after model, control, review, incident, and rework cost" — คุณค่าหลังหักค่าโมเดล Control การตรวจ Incident และ Rework ห้าบรรทัดนี้แทบไม่เคยอยู่ในเอกสารเดียวกัน เพราะแต่ละบรรทัดถูกบันทึกโดยระบบคนละระบบและตกอยู่กับงบคนละก้อน

Cost line สิ่งที่อยู่ในนั้น จุดที่มันมักถูกซ่อน
Model ค่าเรียกโมเดลทั้ง input และ output รวมการเรียกซ้ำเมื่อผลรอบแรกใช้ไม่ได้ อยู่ในบิลรวมของทั้งองค์กร แยกไม่ออกว่างานใดเป็นคนกิน
Control Guard ก่อนเกิดผล การตรวจสิทธิ์ การบันทึกและจัดเก็บ Trace ถูกนับเป็นค่าแพลตฟอร์มกลาง จึงไม่เคยผูกกลับไปหางานที่ก่อมัน
Review เวลาผู้ตรวจต่อชิ้นงาน คูณด้วยค่าแรงจริงของคนระดับที่ตรวจได้ ถูกมองว่าเป็น "งานเดิมที่เขาทำอยู่แล้ว" จึงไม่เข้าบัญชีของระบบใหม่
Incident เวลาสืบสวน การเยียวยาผู้ได้รับผล การแจ้ง และช่วงที่บริการหยุด อยู่ในระบบจัดการเหตุการณ์ที่ไม่เคยคุยกับระบบต้นทุน
Rework งานที่ต้องทำซ้ำเพราะผลลัพธ์ใช้ไม่ได้ ทั้งที่ต้นทางและที่ปลายทาง ตกไปอยู่กับทีมปลายน้ำที่ไม่ได้เป็นเจ้าของงบ AI จึงไม่มีใครรายงาน

บทที่ 3 เติมบรรทัดที่หกซึ่งผมคิดว่าสำคัญที่สุดสำหรับการตัดสินใจว่าจะเลิกอะไร คือ evidence cost หรือ "ต้นทุนหลักฐาน" — เงินและเวลาที่ใช้ไปกับการพิสูจน์ว่าระบบทำงานถูก เมื่อไรที่ต้นทุนหลักฐานของงานหนึ่งสูงกว่าคุณค่าที่งานนั้นสร้าง คำตอบไม่ใช่การลดหลักฐาน คำตอบคือการยุติงานนั้น การลดหลักฐานเพื่อให้ตัวเลขสวยขึ้นคือการย้ายต้นทุนไปไว้ในช่อง Risk แล้วเลิกวัดมัน

อีกครึ่งของช่อง Economics คือความสามารถที่ใช้ซ้ำได้

ครึ่งหลังของบรรทัด Economics คือ reusable capability share ซึ่งเชื่อมตรงเข้ากับ โรงงาน AI และข้อมูล (AI and data factory) ของบทที่ 5 หนังสือระบุว่าโรงงานมีบริการที่ใช้ซ้ำได้หกอย่าง คือ Data products ให้ข้อเท็จจริงที่กำกับแล้ว, Context services ประกอบคำสั่ง หลักฐานที่ค้นมา และ Memory, Model services ส่งงานไปยังโมเดลที่เล็กที่สุดที่เพียงพอและจัดการการเปลี่ยน Provider, Evaluation services ดูแลกรณีที่ติดฉลาก การสอบเทียบผู้ตัดสิน และการรันชุดทดสอบ, Tool registry services กำหนด Scope, Schema, ชั้นของผลกระทบ และเงื่อนไขอนุมัติ และ Observability services เก็บ Trace, Outcome, Cost, Drift และ Incident

แล้วประโยคปิดของย่อหน้านั้นในหนังสือคือสิ่งที่ทำให้บทความนี้เป็นบทความเดียว ไม่ใช่สองเรื่องที่ถูกเย็บติดกัน: "Shared security, privacy, FinOps, and sustainability controls span the six." — Security, Privacy, FinOps และ Sustainability พาดผ่านทั้งหกบริการ ในรูปที่ 6 ของหนังสือ แถบสีเข้มด้านล่างของโรงงานเขียนไว้ว่า SHARED ASSURANCE • SECURITY • PRIVACY • FINOPS • GREENOPS พร้อมคำแปลไทยว่า "การประกันความเชื่อมั่น ความมั่นคง ความเป็นส่วนตัว ต้นทุน และพลังงาน" นั่นคือคำตอบเชิงโครงสร้างว่าทำไมต้นทุนกับ footprint ถึงอยู่ในบทความเดียวกัน — ในสถาปัตยกรรมของหนังสือ ทั้งสองเรื่องเป็นแถบเดียวกันที่พาดผ่านทุกบริการ ไม่ใช่โครงการแยกของสองฝ่าย

ผลที่ตามมาในทางปฏิบัติคือ ตัวชี้วัดของบทที่ 5 จับสองเรื่องนี้ไว้ในหน่วยเดียวกันตั้งแต่ต้น: "energy and cost per successful task" — พลังงานและต้นทุนต่องานที่สำเร็จ ถ้าองค์กรวัดสองอย่างนี้ด้วยตัวหารเดียวกัน การถกเถียงเรื่อง "ประหยัดเงินแต่เปลืองพลังงาน" หรือกลับกัน จะจบลงด้วยตัวเลขแทนที่จะจบด้วยความเห็น

4. Footprint — อ่านตัวเลขให้ถูก แล้ววัดของตัวเอง

หนังสือเปิดหัวข้อนี้ด้วยประโยคที่กำหนดตำแหน่งของทั้งเรื่อง: "Sustainability belongs inside the business case." — ความยั่งยืนอยู่ใน Business Case ไม่ใช่ในรายงานประจำปีที่ออกมาหลังจากทุกอย่างตัดสินไปแล้ว และภาคผนวกคำศัพท์นิยาม AI ที่ยั่งยืน (sustainable AI) ไว้ว่า "การออกแบบและใช้ AI โดยวัดและพิจารณาพลังงาน คาร์บอน น้ำ ฮาร์ดแวร์ ที่ตั้ง การใช้ทรัพยากร ประสิทธิภาพผู้ให้บริการ และคุณค่าต่อสังคม" — สังเกตคำว่า ประสิทธิภาพผู้ให้บริการ ในรายการนั้น มันคือสะพานที่หนังสือทอดกลับไปหา ส่วนที่ 2 ด้วยคำของหนังสือเอง เรื่องสัญญากับเรื่องพลังงานไม่ได้แยกจากกัน

ทีนี้ถึงตัวเลข IEA ประเมินว่า data centre ทั่วโลกใช้ไฟฟ้าราว 485 TWh ในปี 2025 — ตัวเลขของ data centre ทั้งหมด ไม่ใช่ของงาน AI อย่างเดียว — และอาจเพิ่มเป็นราว 950 TWh ในปี 2030 หรือราว 3 เปอร์เซ็นต์ของความต้องการไฟฟ้าทั้งโลกในปีนั้น[7] ตรวจสอบเมื่อ 5 กันยายน 2026: รายงาน Key Questions on Energy and AI ของ IEA เผยแพร่วันที่ 16 เมษายน 2026 และยังคงตัวเลขชุดนี้ตรงกับที่หนังสืออ้างไว้ รายงานฉบับนี้ต่อยอดจากรายงานเรือธง Energy and AI ของ IEA เมื่อปี 2025[8]

และประโยคที่ต้องเดินทางไปกับตัวเลขทั้งสามเสมอคือประโยคของหนังสือเอง: "These are uncertain sector scenarios, not the meter for one use case." — เป็น Scenario ของภาคที่มีความไม่แน่นอนสูง ไม่ใช่มิเตอร์ของ use case เดียว รายงานของ IEA เองก็ระบุว่าค่าประมาณกลางของมันยังใกล้เคียงกับเส้นทางที่รายงานปี 2025 วางไว้ และคอขวดตลอดห่วงโซ่คุณค่ากำลังลดความเป็นไปได้ของ scenario ที่ก้าวร้าวกว่านี้ในระยะสั้น นั่นคือภาษาของการสร้างแบบจำลอง ไม่ใช่ภาษาของการพยากรณ์

สิ่งที่ห้ามทำกับตัวเลขชุดนี้: ห้ามเขียนว่า "AI จะใช้ไฟ 950 TWh" เพราะตัวเลขเป็นของ data centre ทั้งหมด ห้ามหารมันด้วยจำนวนผู้ใช้ จำนวน query หรือจำนวน token เพื่อให้ได้ค่าพลังงานต่อครั้ง และห้ามยกไปเป็นค่าอ้างอิงของผู้ให้บริการรายใดรายหนึ่ง IEA แยกประสิทธิภาพ การนำไปใช้ และขีดความสามารถของโมเดล ออกเป็นสามแนวโน้มที่เปลี่ยนเร็วและไม่แน่นอน การหารตัวเลขระดับภาคด้วยอะไรก็ตามเพื่อให้ได้ตัวเลขระดับ use case คือการสร้างข้อมูลขึ้นมาเอง แล้วอ้างชื่อ IEA กำกับ

ถ้าตัวเลขระดับโลกใช้เป็นมิเตอร์ไม่ได้ แล้วองค์กรควรวัดอะไร หนังสือให้รายการห้าอย่างไว้ในประโยคเดียว: พลังงานรวม, Emission เท่าที่ประเมินได้, น้ำเมื่อสำคัญ, Hardware Lifecycle และ Rebound เมื่อราคาต่อหน่วยลดแต่ Usage เพิ่ม สี่ข้อแรกเป็นการบัญชี ข้อที่ห้าเป็นเรื่องพฤติกรรม และเป็นข้อที่ทำให้อีกสี่ข้อมีความหมาย

Rebound คือกลไกที่ทำให้รายงานความยั่งยืนส่วนใหญ่อ่านแล้วสบายใจโดยไม่มีเหตุผล เรื่องเล่ามาตรฐานเป็นแบบนี้ ทีมย้ายไปใช้โมเดลที่เล็กลง ต้นทุนต่อครั้งลดลงชัดเจน พลังงานต่อครั้งลดลงตาม สไลด์ถูกนำเสนอต่อคณะกรรมการพร้อมกราฟที่ลาดลงอย่างสวยงาม สิ่งที่สไลด์นั้นไม่ได้แสดงคือสิ่งที่เกิดขึ้นหลังจากนั้น — เมื่อของถูกลง ทีมอื่นก็เปิดใช้บ้าง งานที่เคยไม่คุ้มค่าที่จะทำก็กลายเป็นคุ้ม ปริมาณการเรียกทั้งองค์กรจึงโตเร็วกว่าที่ต้นทุนต่อครั้งลดลง และผลรวมสัมบูรณ์ทั้งเงินและพลังงานลงเอยสูงกว่าเดิม ทั้งที่ทุกตัวเลข ต่อหน่วย ในสไลด์นั้นถูกต้องทุกตัว

หลักปฏิบัติข้อ 5 ของบทที่ 10 บีบเรื่องนี้ให้เหลือประโยคเดียวที่ผมอยากให้ติดไว้บนหน้าปกของทุกรายงานประเภทนี้: "Unit efficiency can coexist with rising total impact" — "ประสิทธิภาพต่อหน่วยอาจดีขึ้นขณะผลรวมสูงขึ้น" คำว่า coexist เลือกมาอย่างระมัดระวัง หนังสือไม่ได้บอกว่าประสิทธิภาพต่อหน่วยเป็นเรื่องหลอกลวง มันบอกว่าตัวเลขสองตัวนี้เป็นจริงพร้อมกันได้ และรายงานที่แสดงตัวเดียวจึงไม่ผิด แต่ไม่ครบ วิธีแก้ในทางปฏิบัติมีข้อเดียว คือทุกครั้งที่รายงานค่าต่อหน่วย ต้องรายงานผลรวมสัมบูรณ์ในกรอบเวลาเดียวกันไว้ข้าง ๆ กันเสมอ

ข้อควรระวังเรื่องการอ้างอิงหนึ่งข้อ รายการห้าอย่างที่ต้องวัดข้างต้นเป็นของหนังสือ ไม่ใช่ของ OECD สิ่งที่ตรวจสอบได้จากเอกสาร OECD คือหลักการเชิงคุณค่าข้อแรกชื่อ "inclusive growth, sustainable development and well-being" และหลักการข้อ 1.5 Accountability ที่เรียกร้องการบริหารความเสี่ยงอย่างเป็นระบบตลอดวงจรชีวิตของระบบ AI[5] ไม่มีข้อความใดที่ตรวจพบว่ากำหนดให้ต้องวัดพลังงาน คาร์บอน หรือน้ำ ผมจึงอ้าง OECD สำหรับกรอบการบริหารความเสี่ยงตลอดวงจร และอ้างหนังสือกับ IEA สำหรับรายการที่ต้องวัด

5. ใช้โมเดลเล็กที่สุดที่ผ่านเกณฑ์ — และวัดมันให้ได้

กฎที่หนังสือให้ไว้เป็นประโยคเดียวและครบในตัว: "Route tasks to the smallest model that meets quality and risk requirements and compare footprint per successful outcome rather than token alone." คู่มือภาษาไทยเขียนว่า "ใช้โมเดลเล็กที่สุดที่ผ่านเกณฑ์และเปรียบเทียบ Footprint ต่อ Successful Outcome"

คำที่คนอ่านข้ามบ่อยที่สุดคือ ที่ผ่านเกณฑ์ กฎนี้ไม่ใช่ "ใช้โมเดลเล็กที่สุด" และไม่ใช่ "ใช้โมเดลถูกที่สุด" มันคือกฎสองชั้น ชั้นแรกคือเกณฑ์คุณภาพและเกณฑ์ความเสี่ยงซึ่งประกาศไว้ก่อน ชั้นที่สองคือเลือกตัวที่เล็กที่สุดในบรรดาตัวที่ผ่านชั้นแรก ลำดับนี้ห้ามสลับ องค์กรที่สลับลำดับจะได้ระบบที่ถูกลงจริงและผิดบ่อยขึ้นจริง แล้วจ่ายส่วนต่างคืนในบรรทัด Review, Incident และ Rework ของ ส่วนที่ 3 ซึ่งไม่มีใครนับ

สถาปัตยกรรมของหนังสือวางความรับผิดชอบนี้ไว้ที่จุดเดียวอย่างจงใจ Model services ของโรงงาน AI และข้อมูล ทำสองอย่างในบริการเดียว คือ "ส่งงานไปยังโมเดลขนาดเหมาะสม" และ "จัดการการเปลี่ยน Provider" นี่ไม่ใช่ความบังเอิญของการจัดกลุ่ม การเลือกขนาดโมเดลกับการรับมือการเปลี่ยนรุ่นของผู้ให้บริการ เป็นคำถามเดียวกันที่ถามคนละเวลา — ทั้งคู่ถามว่า "งานนี้ต้องการความสามารถระดับใด และเรารู้ได้อย่างไรว่ายังได้ระดับนั้นอยู่"

คำแนะนำนี้ยังมีแหล่งจากนอกหนังสือมายืนยันโดยอิสระ FinOps Foundation เขียนไว้ตรง ๆ ว่าให้ "avoid using the most complex and expensive models for every task, as this often leads to unnecessary costs"[6] — เลี่ยงการใช้โมเดลที่ซับซ้อนและแพงที่สุดกับทุกงาน เพราะมักนำไปสู่ต้นทุนที่ไม่จำเป็น ต่างกันตรงที่หนังสือเติมชั้นเกณฑ์ความเสี่ยงเข้าไปข้างหน้า ซึ่งเป็นชั้นที่เอกสารด้านต้นทุนไม่มีหน้าที่ต้องมี

ข้อควรระวังสำหรับคนที่กำลังจะตั้งเป้า: ทั้ง "สัดส่วนความสามารถที่ใช้ซ้ำได้" และ "สัดส่วนงานที่ถูกส่งไปโมเดลเล็กที่สุดที่ผ่านเกณฑ์" เป็น นิยามของตัวชี้วัด ในหนังสือ ไม่ใช่ค่าที่วัดได้จากที่ใด หนังสือไม่ได้ให้ตัวเลขเป้าหมายไว้ และผมก็จะไม่แต่งขึ้นมา ค่าที่ถูกต้องขึ้นกับส่วนผสมของงานในองค์กรนั้น ๆ สิ่งที่วัดได้จริงคือแนวโน้มของตัวเองเทียบกับตัวเองในไตรมาสก่อน

ตารางออกแบบการวัด

ตารางข้างล่างคือเวิร์กช็อปชุดที่สองของบทความนี้ และมันควรถูกกรอกให้เสร็จ ก่อน เริ่มวัด ไม่ใช่หลังจากมีข้อมูลกองแรกแล้ว คอลัมน์ที่สำคัญที่สุดคือคอลัมน์ที่สี่ ตัวชี้วัดที่ไม่มีชื่อคนกำกับจะกลายเป็นสไลด์ที่ไม่มีใครอัปเดตภายในสองไตรมาส และคอลัมน์ที่สองบังคับให้ทุกแถวใช้ตัวหารเดียวกัน คือ Successful Outcome

What we measure Unit (per successful outcome) Source Owner Review
ต้นทุนรวมของงานที่สำเร็จ บาทต่อ Successful Outcome รวมทั้งห้าบรรทัดของต้นทุน Observability (cost) + ระบบบัญชี + ใบแจ้งหนี้ผู้ให้บริการ เจ้าของผลิตภัณฑ์ ร่วมกับ FinOps รายเดือน
ต้นทุนหลักฐาน บาทต่อ Successful Outcome แยกออกจากต้นทุนโมเดลเสมอ เวลาผู้ตรวจ + ค่าทดสอบ + ค่าจัดเก็บ Trace เจ้าของ Domain รายไตรมาส
สัดส่วนงานที่ถูกส่งไปโมเดลเล็กที่สุดที่ผ่านเกณฑ์ เปอร์เซ็นต์ของ request ที่สำเร็จ แยกตามชั้นความเสี่ยง Log ของ Model Service routing เจ้าของแพลตฟอร์ม รายเดือน
พลังงาน kWh ต่อ Successful Outcome หรือค่าที่ผู้ให้บริการรายงานพร้อมวิธีคิด รายงานผู้ให้บริการ + มิเตอร์ของเราเองสำหรับส่วนที่รันเอง เจ้าของแพลตฟอร์ม รายไตรมาส
Emission เท่าที่ประเมินได้ หน่วยคาร์บอนต่อ Successful Outcome พร้อมวิธีประเมินและช่วงความไม่แน่นอน ค่าพลังงาน × ความเข้มคาร์บอนของพื้นที่ที่ประมวลผล ฝ่ายความยั่งยืน รายปี
น้ำ เมื่อสำคัญ ปริมาณต่อ Successful Outcome เฉพาะที่ตั้งที่เรื่องนี้มีนัยจริง รายงานของผู้ให้บริการหรือศูนย์ข้อมูลที่เราใช้ ฝ่ายความยั่งยืน รายปี
Hardware Lifecycle อายุการใช้งานและปลายทางของอุปกรณ์ที่เราเป็นเจ้าของ ทะเบียนทรัพย์สิน + เงื่อนไขในสัญญาผู้ให้บริการ ฝ่ายโครงสร้างพื้นฐาน รายปี
Total Footprint หลัง Usage Growth ผลรวมสัมบูรณ์ เทียบกับช่วงเวลาเดียวกันของปีก่อน ค่าต่อหน่วย × ปริมาณการใช้จริงทั้งองค์กร คณะกรรมการที่ดูแลพอร์ต รายไตรมาส

แถวสุดท้ายมีอยู่เพื่อกันข้อผิดพลาดข้อเดียว คือการรายงานเจ็ดแถวแรกแล้วหยุด เจ็ดแถวแรกทั้งหมดเป็นค่า ต่อหน่วย และตามหลักปฏิบัติข้อ 5 ของบทที่ 10 ค่าต่อหน่วยที่ดีขึ้นกับผลรวมที่สูงขึ้นเป็นจริงพร้อมกันได้ แถวที่แปดจึงไม่ใช่ส่วนขยาย มันคือแถวที่ทำให้เจ็ดแถวแรกอ่านได้อย่างซื่อสัตย์

เรื่องนี้ไม่ได้เกิดในสุญญากาศ — มันอยู่ในเวิร์กช็อปของตอนที่แล้ว

สองหัวข้อของบทความนี้ไม่ได้ต้องการเวทีใหม่ ทั้งคู่เป็นขั้นตอนในเวิร์กช็อป Impact and obligation mapping ซึ่งเป็นเวิร์กช็อปเดียวกับที่ตอน #16 One Evidence System เดินไปแล้ว เวิร์กช็อปนั้นมีเจ็ดขั้น และขั้นที่ห้ากับหกคือของเรา — "Review vendor evidence and exit" ตรวจหลักฐานของผู้ให้บริการและทางออก และ "Estimate energy and model-sizing options" ประเมินพลังงานกับทางเลือกด้านขนาดโมเดล ก่อนจะไปถึงขั้นที่เจ็ดคือตัดสิน Proceed, Modify, Pause หรือ Prohibit พร้อมระบุช่องว่างของหลักฐาน

ผมชอบการจัดวางแบบนี้มาก เพราะมันทำให้คำถามเรื่องผู้ให้บริการและคำถามเรื่องพลังงานเกิดขึ้น ก่อน คำตัดสิน ไม่ใช่หลังจากระบบขึ้นแล้วในรูปของแบบสอบถามที่ส่งเวียนกลับมาให้กรอก และมันหมายความว่าองค์กรไม่ต้องตั้งคณะทำงานใหม่เพื่อเริ่มเรื่องนี้ ต้องเพิ่มสองบรรทัดในวาระที่มีอยู่แล้วเท่านั้น

หมายเหตุเรื่องขอบเขต: ค่าพลังงานต่อ query หรือต่อ token ที่วนอยู่ในบทความทั่วไปบนอินเทอร์เน็ต ไม่มีอยู่ในหนังสือเล่มนี้ และผมไม่ยกมาใช้ สิ่งที่หนังสือสั่งคือให้ วัดของตัวเอง ในหน่วยต่อ Successful Outcome ซึ่งเป็นการออกแบบการวัด ไม่ใช่ค่าคงที่ที่หยิบจากที่อื่นมาใส่ได้ ถ้าผู้ให้บริการของคุณรายงานค่าพลังงานมาให้ ให้บันทึกวิธีคิดของเขาไว้ในคอลัมน์ Source ด้วยเสมอ เพราะสองผู้ให้บริการที่คิดคนละวิธีเทียบกันไม่ได้

6. ตัวชี้วัดสำคัญ

รายการตัวชี้วัดของบทที่ 10 มีสิบสามตัว ตารางข้างล่างหยิบมาเฉพาะตัวที่เป็นเรื่องของผู้ให้บริการ ต้นทุน และ footprint แล้วเติมสองตัวจากบทที่ 3 กับกระดานคะแนนคณะกรรมการเข้าไป พร้อมช่อง Scorecard ตามธรรมเนียมของซีรีส์นี้ เพื่อให้เห็นว่าแต่ละตัวไปลงช่องไหนของกระดาน

ตารางนี้เอียงไปทาง Economics และ Risk อย่างเห็นได้ชัด และนั่นถูกต้องตามบทบาทของบทความนี้ ตอนอื่นในซีรีส์เป็นเจ้าของช่อง Value, Quality, People และ Learning — ข้อควรระวังคืออย่ายกตารางนี้ไปใช้เป็นกระดานคะแนนทั้งใบ เพราะหนังสือห้ามยุบหกช่องเป็นมุมมองเดียวไว้ตั้งแต่หน้าแรก

Metric สิ่งที่บอกเรา สัญญาณเตือน Scorecard
Vendor Evidence สัดส่วนผู้ให้บริการที่ส่งหลักฐานครบตามสัญญา และหลักฐานนั้นอ้างถึงรุ่นที่รันอยู่จริง หลักฐานล่าสุดลงวันที่ก่อนการเปลี่ยนรุ่นครั้งล่าสุด — คือหลักฐานที่หมดอายุแล้ว Risk
ความกระจุกตัวที่ Provider สัดส่วนของพอร์ตที่แขวนอยู่กับผู้ให้บริการหรือตระกูลโมเดลรายเดียว งานที่องค์กรขาดไม่ได้ทุกงานอยู่บนผู้ให้บริการรายเดียวกัน และไม่มีใครเคยรวมภาพนี้ Risk
ต้นทุนรวมต่อ Successful Outcome ต้นทุนจริงของผลลัพธ์ที่ใช้งานได้หนึ่งชิ้น หลังหักครบทั้งห้าบรรทัด รายงานเป็นต้นทุนต่อ token หรือต่อ request ซึ่งดีขึ้นได้แม้ผลสำเร็จลดลง Economics
ต้นทุนหลักฐาน เงินและเวลาที่ใช้พิสูจน์ว่าระบบทำงานถูก แยกจากต้นทุนการรันระบบ สูงกว่าคุณค่าที่ระบบสร้าง — สัญญาณให้ยุติ ไม่ใช่สัญญาณให้ลดการตรวจ Economics
สัดส่วนความสามารถที่ใช้ซ้ำได้ ส่วนของงานใหม่ที่สร้างจากบริการเดิมหกอย่างของโรงงาน แทนที่จะสร้างใหม่ทั้งชิ้น ทุกโครงการใหม่มี Pipeline, ชุดประเมิน และ Guard ของตัวเอง Economics
Smallest-adequate-model Routing สัดส่วนงานที่วิ่งบนโมเดลเล็กที่สุดที่ยังผ่านเกณฑ์คุณภาพและความเสี่ยง ตัวเลขดีขึ้นพร้อมกับอัตราการตรวจซ้ำและ Escalation ที่สูงขึ้น — คือการลดเกณฑ์ ไม่ใช่การ route Economics
พลังงานและ Emission ต่อ Successful Outcome ต้นทุนทางกายภาพของผลลัพธ์หนึ่งชิ้น ในหน่วยเดียวกับต้นทุนทางการเงิน รายงานโดยไม่ระบุวิธีประเมินและช่วงความไม่แน่นอน จึงเทียบข้ามปีไม่ได้ Economics
Total Footprint หลัง Usage Growth ผลรวมสัมบูรณ์ของพลังงานและต้นทุน หลังจากปริมาณการใช้เติบโตแล้ว ไม่เคยถูกรายงานคู่กับค่าต่อหน่วย จึงไม่มีใครเห็น Rebound จนสายเกินไป Economics

รูปแบบความล้มเหลว

  • คิดว่า Vendor รับความเสี่ยงแทน — ถ้อยคำของหนังสือเอง เป็นข้อสรุปที่สมเหตุสมผลถ้าเราเข้าใจผิดว่าสิ่งที่ซื้อคือบริการสาธารณูปโภค แต่หน้าที่ตอบว่าการตัดสินใจถูกหรือไม่ ไม่เคยย้ายไปกับใบแจ้งหนี้
  • เชื่อ Vendor Test — รูปแบบความล้มเหลวของบทที่ 9 ผลทดสอบของผู้ให้บริการอยู่ในบริบทของผู้ให้บริการ ไม่ใช่ในกระบวนงาน ภาษา ข้อมูล และผู้ใช้ของเรา
  • ไม่มีข้อ Exit — ข้อนี้ผมสรุปจากหลักปฏิบัติข้อ 4 ของบทที่ 10 คู่กับหลักปฏิบัติข้อ 5 ของบทที่ 3 ไม่ใช่ถ้อยคำที่หนังสือพิมพ์ไว้เป็นรายการความล้มเหลว แต่ผลลัพธ์ตรงกัน คือความสัมพันธ์ที่เราออกไม่ได้ ไม่ใช่ความสัมพันธ์ที่เราเลือก
  • ใช้ต้นทุนต่อ token เป็น KPI — สรุปจากคำสั่งของหนังสือที่ให้เทียบ footprint ต่อ Successful Outcome แทนที่จะเทียบต่อ token อย่างเดียว ตัวเลขต่อ token มีประโยชน์สำหรับวิศวกร แต่ดีขึ้นได้พร้อมกับที่งานสำเร็จน้อยลง
  • อ้างประสิทธิภาพโดยไม่แสดงผลรวม — หลักปฏิบัติข้อ 5 ของบทที่ 10 ตรง ๆ ประสิทธิภาพต่อหน่วยที่ดีขึ้นกับผลกระทบรวมที่สูงขึ้น อยู่ด้วยกันได้เสมอ
  • วางความยั่งยืนไว้นอก Business Case — กลับด้านของประโยค "Sustainability belongs inside the business case" เมื่อความยั่งยืนเป็นรายงานประจำปีแยกเล่ม มันจะไม่มีวันเปลี่ยนคำตัดสินใด ๆ ได้ทัน
  • อ้าง Guideline หรือ Certificate ว่าเท่ากับ Legal Compliance — ถ้อยคำของหนังสือเอง และเป็นข้อที่กัดแรงที่สุดในบริบทของผู้ให้บริการ เพราะใบรับรองคือสิ่งที่ฝ่ายขายมีติดตัวเสมอ

หกในเจ็ดข้อนี้แก้ได้ด้วยเอกสาร คือสัญญาใน ส่วนที่ 2 กับตารางการวัดใน ส่วนที่ 5 แต่ข้อสุดท้ายแก้ด้วยเอกสารไม่ได้ เพราะมันเป็นความผิดพลาดเรื่องประเภทของข้อโต้แย้ง ใบรับรองบอกว่าองค์กรหนึ่งมีระบบการจัดการแบบหนึ่ง หลักฐานทางกฎหมายบอกว่าการกระทำหนึ่งในเขตอำนาจหนึ่งชอบด้วยกฎหมาย ไม่มีจำนวนใบรับรองเท่าใดที่แปลงเป็นข้อที่สองได้

7. ก้าวต่อไป

ถ้าจะสรุปทั้งบทเป็นการเปลี่ยนวิธีทำงานอย่างเดียว ผมจะเลือกการเปลี่ยนตัวหาร เปลี่ยนคำถามในรายงานประจำเดือนจาก "เดือนนี้เราจ่ายค่าโมเดลไปเท่าไร" เป็น "เดือนนี้ผลลัพธ์ที่ใช้งานได้หนึ่งชิ้น มีต้นทุนรวมและใช้พลังงานเท่าไร และผลรวมทั้งองค์กรเป็นเท่าไรเทียบกับปีก่อน" คำถามแรกผลิตการต่อรองราคา คำถามที่สองผลิตการตัดสินใจ

สามอย่างที่ทำได้ทันทีสัปดาห์หน้า หนึ่ง หยิบสัญญาผู้ให้บริการ AI ฉบับที่ใหญ่ที่สุดที่องค์กรมีอยู่ แล้วเดินตารางใน ส่วนที่ 2 ทีละข้อ ทำเครื่องหมายว่าข้อใดมี ข้อใดไม่มี — ประสบการณ์ของผมคือข้อ 3, 5 และ 11 จะว่างพร้อมกันบ่อยที่สุด สอง เปิดสเปรดชีตหนึ่งใบแล้วกรอกตารางใน ส่วนที่ 5 ให้ครบทุกช่อง โดยเฉพาะคอลัมน์ Owner ช่องที่กรอกไม่ได้คือคำตอบที่มีประโยชน์ที่สุดของแบบฝึกหัดนี้ และสาม เอารายงานต้นทุน AI ฉบับล่าสุดมาดูว่าตัวหารคืออะไร ถ้าเป็น token หรือ request ให้เพิ่มอีกหนึ่งบรรทัดที่หารด้วยจำนวนผลลัพธ์ที่ใช้งานได้จริง แล้ววางสองบรรทัดนั้นไว้ข้างกันในรายงานเดือนถัดไป

ทั้งหมดนี้ — พอร์ตที่จำแนกแล้ว กระบวนงานที่ออกแบบใหม่ โรงงานที่ประกอบขึ้น สัญญาที่ครบ และมิเตอร์ที่หารถูก — ยังไม่ใช่ลำดับที่องค์กรลงมือทำได้ในเช้าวันจันทร์ มันเป็นรายการสิ่งที่ต้องมี ไม่ใช่แผน สิ่งที่ขาดคือลำดับเวลาและขอบเขตที่เล็กพอจะเสร็จจริง ซึ่งเป็นเรื่องของตอนหน้าทั้งตอน

🧭 ชั้นที่บทความนี้ขยับ: ชั้น AI and data factory (โรงงาน AI และข้อมูล) และชั้น Operating model (รูปแบบการดำเนินงาน) — บทความนี้ตอบคำถามผู้นำข้อ Q7 ("ใครเป็นเจ้าของคุณค่า ความเสี่ยง ผลกระทบ ข้อยกเว้น และการเรียนรู้") ซึ่งเป็นคำถามที่การจ้างผู้ให้บริการภายนอกทำให้ตอบยากที่สุด เพราะคำตอบดูเหมือนย้ายออกไปนอกองค์กรแล้วทั้งที่ไม่ได้ย้าย บนกระดานคะแนนองค์กร บทนี้ขยับช่อง Economics เป็นหลัก (ต้นทุนรวมต่อ Successful Outcome, ต้นทุนหลักฐาน, สัดส่วนความสามารถที่ใช้ซ้ำได้, พลังงานต่อผลลัพธ์ และผลรวมหลัง Usage Growth) และช่อง Risk เป็นรอง (Vendor Evidence และความกระจุกตัวที่ Provider) ตอนถัดไป #18 The First 90 Days — Align, Design, Build ก่อนขยายอะไรทั้งสิ้น เปลี่ยนรายการสิ่งที่ต้องมีทั้งหมดนี้ให้เป็นลำดับเวลาที่ทีมจริงเดินตามได้

🎯 สิ่งสำคัญที่ต้องจำ

  • Accountability stays = outsource เทคโนโลยีได้ แต่ outsource ความรับผิดรับชอบ ไม่ได้ — งานย้ายไปกับสัญญา หน้าที่ตอบคำถามไม่ย้าย
  • Eleven clauses + exit = Data Use, การฝึกด้วยข้อมูลลูกค้า, การเปลี่ยน Model หรือ Subprocessor, Evaluation Evidence, Log, Security Test, Incident Notice, Audit Support, Continuity, IP และ Exit — ข้อสุดท้ายคือข้อที่หายบ่อยที่สุด
  • Provider concentration = ความกระจุกตัวที่ Provider เป็นความเสี่ยงระดับพอร์ต มองเห็นได้จากมุมพอร์ตเท่านั้น ไม่ใช่จากรายงานของระบบใดระบบหนึ่ง
  • Economics column = ต้นทุนรวมต่อ Successful Outcome และสัดส่วนความสามารถที่ใช้ซ้ำได้ โดยคำว่า "รวม" หมายถึงห้าบรรทัดคือ Model, Control, Review, Incident และ Rework
  • 485 → 950 TWh = สถานการณ์ระดับภาคของ data centre ทั้งหมด ไม่ใช่มิเตอร์ของ use case ใด ห้ามหารเพื่อให้ได้ค่าต่อ query
  • Per successful outcome = ตัวหารเดียวกันสำหรับทั้งต้นทุนและ footprint ไม่ใช่ต่อ token ซึ่งดีขึ้นได้ขณะที่งานสำเร็จน้อยลง
  • Rebound = ราคาต่อหน่วยที่ถูกลงทำให้การใช้งานโตจนผลรวมสูงขึ้น ทุกค่าต่อหน่วยจึงต้องรายงานคู่กับผลรวมสัมบูรณ์เสมอ

อ้างอิง

ตรวจสอบทุกแหล่งเมื่อ 5 กันยายน 2026 (เวลาประเทศไทย) · ป้ายหลักฐานสี่แบบ: Law ตัวบทกฎหมายหรือประกาศทางการ · Standard มาตรฐานหรือกรอบทางการที่เผยแพร่แล้ว · Study งานวิจัยหรือสัญญาณภาคสนาม · Synthesis การสังเคราะห์ของผู้เขียนหรือแหล่งที่ไม่ใช่งานวิจัย

  1. Synthesis Anirach Mingkhwan. AI Transformation as an Organizational Core — Bilingual Companion Playbook — บทที่ 3, 4, 5, 9, 10 คู่มือผู้บริหารฉบับย่อ และภาคผนวก C, D. ต้นฉบับของผู้เขียน ไม่มี URL สาธารณะ; ข้อมูลหลักฐาน ณ 5 กันยายน 2026. รองรับ: รายการ 11 ข้อสัญญากับ Supplier, หลักปฏิบัติห้าประการของบทที่ 10, ช่อง Economics ของกระดานคะแนน, ความกระจุกตัวที่ Provider, บริการหกอย่างของโรงงาน AI และข้อมูล, เวิร์กช็อป Impact and obligation mapping, รายการตัวชี้วัดและรูปแบบความล้มเหลว
  2. Law European Union. Regulation (EU) 2024/1689 (Artificial Intelligence Act) — มาตรา 53 และมาตรา 113 — ตรวจกับตัวบทฉบับรวมแก้ไข ณ 27 กรกฎาคม 2026. eur-lex.europa.eu — เข้าถึง 2026-09-05. รองรับ: หน้าที่ของผู้ให้บริการ general-purpose AI model ในการจัดทำเอกสารทางเทคนิคและส่งข้อมูลให้ผู้พัฒนาระบบปลายทาง และการที่มาตรา 53 ไม่ถูกแก้โดยการแก้ไขปี 2026
  3. Law European Union. Regulation (EU) 2026/1744 of 8 July 2026 amending Regulation (EU) 2024/1689 — ประกาศใน OJ 24 กรกฎาคม 2026 มีผล 27 กรกฎาคม 2026. eur-lex.europa.eu — เข้าถึง 2026-09-05. รองรับ: การบังคับใช้ทั่วไปตั้งแต่ 2 สิงหาคม 2026 และการเลื่อนกำหนดของระบบ high-risk เป็น 2 ธันวาคม 2027 (Annex III) และ 2 สิงหาคม 2028 (Annex I)
  4. Standard ISO/IEC. ISO/IEC 42001:2023 — Information technology, Artificial intelligence, Management system — เผยแพร่ธันวาคม 2023. iso.org — เข้าถึง 2026-09-05. รองรับ: การมีอยู่และสถานะเผยแพร่แล้วของมาตรฐานระบบการจัดการ AI ที่หนังสืออ้างถึง — คำอธิบายเนื้อหาของมาตรฐานในบทความนี้เป็นถ้อยคำของหนังสือ ไม่ได้ดึงจากตัวมาตรฐาน และหน้านี้ปิดกั้นการเรียกดูอัตโนมัติ
  5. Standard OECD. OECD AI Principles — รับรองพฤษภาคม 2019 ปรับปรุงพฤษภาคม 2024 (OECD.AI Policy Observatory). oecd.ai — เข้าถึง 2026-09-05. รองรับ: หลักการเชิงคุณค่าห้าข้อ ข้อ 1.1 inclusive growth, sustainable development and well-being และข้อ 1.5 Accountability ที่กำหนดการบริหารความเสี่ยงอย่างเป็นระบบตลอดวงจรชีวิตระบบ AI พร้อม traceability
  6. Standard FinOps Foundation (Linux Foundation). FinOps for AI — หมวดเทคโนโลยีของ FinOps Framework และเอกสารภาพรวมของคณะทำงาน ปรับปรุง 17 กุมภาพันธ์ 2026. finops.org — เข้าถึง 2026-09-05. รองรับ: KPI ต้นทุน AI ที่ใช้กันในอุตสาหกรรมคือต้นทุนต่อ token, ต่อ inference และต่อ API call และคำแนะนำให้เลี่ยงการใช้โมเดลที่ซับซ้อนและแพงที่สุดกับทุกงาน
  7. Study International Energy Agency. Key Questions on Energy and AI — Executive summary — เผยแพร่ 16 เมษายน 2026. iea.org — เข้าถึง 2026-09-05. รองรับ: ตัวเลขไฟฟ้าของ data centre ราว 485 TWh ในปี 2025 เพิ่มเป็นราว 950 TWh ในปี 2030 หรือราว 3 เปอร์เซ็นต์ของความต้องการไฟฟ้าโลก พร้อมกรอบว่าเป็น scenario ที่มีความไม่แน่นอนสูง
  8. Study International Energy Agency. Energy and AI — รายงานเรือธง เผยแพร่ 10 เมษายน 2025. iea.org — เข้าถึง 2026-09-05. รองรับ: บริบทและลำดับที่มาของรายงานฉบับปี 2026 ซึ่งต่อยอดจากฉบับนี้ — ไม่มีตัวเลขใดในบทความนี้ที่นำมาจากหน้านี้

🤔 If your vendor swapped the model tonight without telling you, would yesterday's release evidence still hold?

The previous post, One Evidence System, ended on the proposal that an organisation should tie the obligations of many frameworks to a single body of evidence instead of opening a new folder every time a new rule arrives. This post asks the question that follows immediately and is almost never asked in the same room — if most of the evidence in that folder comes from systems we do not own, what are we governing it with, and when one of those systems changes itself on a Tuesday night, who finds out?

The whole post's short answer is one sentence from Chapter 10 of AI Transformation as an Organizational Core: "Outsourcing technology does not outsource accountability." What travels with the contract is the work, not accountability. Making that fact governable in practice takes three things at once: a contract covering eleven clauses including exit, the correct denominator for cost, and a footprint meter that measures the total after usage grows rather than measuring per token.

1. The Night the Vendor Changed the Model

Picture it this way. On Monday your team clears the release gate cleanly: the test suite is complete, there are results on the edge cases, there are severe-case statistics, the owner has signed. The evidence file is thick enough to put in front of an auditor on the spot. On Tuesday night the model provider upgrades the version sitting behind the same endpoint. The model name in the API does not change. The price does not change. The documentation does not change. On Wednesday your system still runs, still answers, still passes every line on the dashboard. The only thing that has changed is that the evidence file you signed on Monday now describes a system that no longer exists.

Let me be clear before going further: this is not an incident report with a company name and a date behind it. It is one sentence of the book made visible, and the sentence is very direct. Chapter 9 puts it this way: "Production is the final evaluation environment, not the end of evaluation. Supplier model changes, new user behavior, or policy updates can invalidate previous evidence." — and the Thai companion of the same book renders it as "Production is the final evaluation environment, not the end, because when the supplier, user behaviour or policy changes, the old evidence may no longer hold."

What makes this mechanism frightening is not its violence but its silence. Nothing crashes. No alert fires. Nobody calls at three in the morning. What expires is the claim that we once proved something, and a claim never makes a sound when it dies. It stays quiet until the day someone asks — an auditor, a regulator, a customer whose application was refused, or the other side's lawyer — and only then do we discover that the document in our hands describes something different from what has actually been running.

Chapter 10 files the behaviour that leads to this state under its list of failure patterns, in the shortest possible words: "assuming the vendor owns risk", or in the Thai companion, "believing the vendor carries the risk for us". Chapter 9 adds the one that always travels beside it: "treating vendor testing as sufficient" — "trusting the vendor's tests". Neither of these is the carelessness of a lazy team. Both are the reasonable conclusion if we think what we bought is a "service" in the sense that electricity or storage is a service. The trouble is that what we bought is not that. We bought a component that will take part in decisions with real people on the receiving end, and the duty to answer for whether those decisions were right has never moved with an invoice.

The book defines accountability in its glossary in a way that closes every exit: "Clear ownership for decisions and effects, supported by answerability, evidence, escalation, remedy, and consequences — not responsibility assigned to the model." That last clause is usually skimmed, and it is the heart of it. If responsibility cannot be assigned to the model, it cannot be assigned to the model's owner either.

💡 My view: of Chapter 10's five operating principles, the one I think changes how an organisation works most immediately is the fourth — "Make supplier change and exit governable · Vendor dependence is part of the risk perimeter", or in the Thai companion, "make supplier change and exit governable; vendor dependence sits inside the risk perimeter". The phrase doing the heavy lifting is risk perimeter. Most organisations draw the perimeter at the edge of their own systems and write along that edge "this part belongs to the vendor". This principle says the line has been drawn in the wrong place. Vendor dependence is inside the perimeter, not outside it.

The other four principles of the chapter always work around the fourth. Principle 1, Maintain a live inventory and named owner — governance begins with knowing what exists, which means the inventory must also record which provider each system hangs from. Principle 2, Map obligations to evidence — reuse artifacts responsibly without claiming the frameworks are identical. Principle 3, Use law as the floor and ethics for the gap — define prohibited uses, affected groups and remedy. And principle 5, Measure lifecycle sustainability and rebound — unit efficiency can coexist with rising total impact, which is the subject of the whole of section 4 of this post.

The one question I use to test every contract: if the provider changes the model version tonight, what will tell our organisation about it — a contractual notice, our own monitoring, or the fact that next month's report looks a bit odd? The third answer is the one most organisations actually give, and it means we are not governing anything. We are simply waiting.

2. Eleven Clauses — and the Eleventh Is Exit

Chapter 10 does not leave this as a floating principle. It lists the contract topics as a closed list of eleven[1] in a single sentence: "Govern suppliers through contracts covering data use, customer-data training, model and subprocessor changes, evaluation evidence, logs, security tests, incident notice, audit support, continuity, intellectual property, and exit." The Thai companion renders it as "govern suppliers through contracts covering data use, training on customer data, model or subprocessor change, evaluation evidence, logs, security tests, incident notice, audit support, continuity, IP and exit."

Before the table, one condition stated clearly, once, because this section reads like contract drafting and it is not — the book states plainly that this is not legal advice, and that real use must go through review for the jurisdiction and the industry. The eleven items are a starting frame for asking questions, not wording to drop into an agreement. The book itself insists that "Role and system classification require qualified legal analysis."

The table below is the worksheet I walk through with the procurement team and the legal team together, in one room, half a day per provider. The second column is what to ask for, the third is the evidence that request turns into, and the fourth — the column that shortens the meeting most — is what breaks if the box is empty.

Clause What to ask for Evidence it feeds What breaks if it is missing
1 · Data use The scope of use of our data, the retention period, the regions where processing happens, and the list of who can access it A record of legal basis and data lineage in the inventory Nobody can tell a regulator where customer data went or what it was used for
2 · Customer-data training A prohibition or conditions on training models with customer data, with an opt-out mechanism that can be verified, not merely a checkbox on a settings page Written confirmation from the provider, paired with our own configuration check Customer data flows into the provider's next model generation, unnoticed and irreversible
3 · Model and subprocessor changes Advance notice when the model version changes, when behaviour behind the same endpoint changes, or when a subprocessor is added, plus the right to pin the previous version temporarily A version-change log tied to the manifest of each release Evidence that passed the gate yesterday describes a system that no longer exists today — the scene in section 1
4 · Evaluation evidence The results of the provider's evaluation, naming the test set, the limitations, and the conditions under which those results do not apply An evidence file that becomes part of our own release dossier All we hold is the provider's test results in the provider's context, which is not our context
5 · Logs Access rights to request-level logs carrying timestamps, version identifiers, and enough detail to reconstruct an event Traces used to investigate incidents and answer individual complaints When something happens, we can trace to the edge of the API and no further
6 · Security tests Security test reports, the scope tested, and the dates, plus the right for us or a third party to test the surface we have enabled Evidence for the security stage of the release gate The word "secure" on a sales slide becomes the only evidence we hold
7 · Incident notice The definition of an incident, the notification window, the channel, and the minimum information the first notice must carry The starting point of the clock for the incident learning loop We hear about it from the news or from a customer, which is always later than the window the law sets
8 · Audit support Cooperation with audits, the scope, the frequency, and the forms of evidence accepted, including third-party reports An evidence set internal audit and regulators can use without asking case by case Audit becomes a favour to be requested, and the calendar belongs to the provider
9 · Continuity Service levels, availability, contingency plans, and how the system behaves when the service does not respond A fallback plan that has been rehearsed for real, not a paragraph in a document The service goes down and the business goes with it, because nobody has ever walked the path that does not pass through the model
10 · Intellectual property Rights in outputs, inputs, prompts and the artifacts we create, plus liability if infringement is alleged Rights documentation the legal team can actually pick up when a dispute arrives The dispute over the work arrives after tens of thousands of outputs have already gone out of the door
11 · Exit The right to take data out in a usable form, the timeframe, provable deletion, and the conditions under which we may terminate An exit plan with an owner, a date, and a costed number We no longer choose the provider; the provider chooses us

Clause 11 is the one that goes missing most often, and it goes missing for an understandable reason. At signing time nobody wants to talk about ending the relationship; it looks like distrust on day one. Chapter 3 solves that political problem by reframing it as a budgeting question. Its fifth principle reads "Fund scale and exit together — every investment needs fallback rollback and retirement conditions", or "fund the scale-up together with the way out; every proposal needs fallback, rollback and retirement conditions." Once exit is a line in the budget, it stops being an accusation and becomes the standard shape of an approval paper.

Provider concentration is a portfolio-level risk

Even the best contract, with all eleven clauses in place, cannot answer one question, because the question does not live at contract level. Chapter 3 puts it in the metrics list of the decision portfolio in two words — provider concentration — which the Thai companion calls "concentration at the provider", and pairs it with one of the same chapter's failure patterns: "ignoring shared-model concentration".

The mechanism that lets this risk slip past every approval process is more interesting than the risk itself. Every team assesses its own system separately. Each team concludes, honestly, that "the risk of our system is acceptable", and every one of those conclusions is correct at its own level. But no forum asks the question that is only visible from the portfolio: if every job the organisation cannot do without runs on the same provider, what is the risk at organisational level? Concentration is a property of the sum, not of the parts, and so it is invisible in any component's report, however well written.

Model providers carry legal duties of their own

Contracts are not the only instrument. In the European Union, the law places some duties directly on model providers. Verified against the official text on 5 September 2026: EU AI Act Article 53 requires providers of a general-purpose AI model to draw up and keep up to date the technical documentation of the model, including its training and testing process and evaluation results, and to make information and documentation available to developers who intend to integrate the model, so that they can genuinely understand its capabilities and limitations[2] — and that article was not amended by the 2026 amendment.

Two cautions must always travel with the paragraph above. First, this law does not attach duties to the word "supplier"; it attaches them to defined roles — provider, deployer, importer, distributor, and provider of a general-purpose AI model — and your organisation may itself be one of them. Classifying roles requires qualified legal advice. Second, the deadlines have just moved. Verified on 5 September 2026: the EU AI Act has been generally applicable since 2 August 2026, while the high-risk deadlines were postponed by Regulation (EU) 2026/1744 of 8 July 2026, in force from 27 July 2026, to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems[3]. A postponement does not delete a duty; it moves a date. An organisation that waits for that date before starting to ask providers for documentation will be starting several of its own contract cycles too late.

At the level of standards and frameworks, the book uses wording I think is worth copying whole: "Several frameworks can share evidence without being treated as interchangeable". NIST AI RMF supplies lifecycle functions; ISO/IEC 42001 specifies an AI management system, which as of 5 September 2026 remains a standard published in December 2023[4]; ISO/IEC 42005 addresses lifecycle impact assessment; and the OECD AI Principles, adopted in May 2019 and updated in May 2024, ask actors to "apply a systematic risk management approach to each phase of the AI system lifecycle on an ongoing basis" together with traceability across the system's life[5].

The trap the book names outright: the last failure pattern in Chapter 10 is "describing guidance or certification as proof of legal compliance" — "claiming a guideline or a certificate equals legal compliance". In the context of this post it translates directly: a provider's certificate is evidence that they operate a management system of a certain kind. It is not evidence that our system is lawful in our jurisdiction, and it is not evidence that the output our system delivered to one customer yesterday was correct. Those two sit on different levels of the argument.

3. The Economics Column — Getting the Denominator Right

The book's executive brief lays out a board scorecard with six columns — Value, Quality, Risk, People, Learning and Economics — and the Economics cell is written in a single very short line: "Total cost per successful outcome and reusable capability share". The Thai companion compresses it to "Economics, from cost per successful outcome."

And the sentence that must travel with this scorecard every time is: "No single composite score should replace this view. A faster process with rising severe errors is not progress. A safe system that produces no outcome value is not transformation. A productive workflow that exhausts reviewers is not sustainable." The six columns may never be collapsed into one score. A faster process with more severe errors is not progress. A safe system that creates no value is not transformation. And a workflow that produces a great deal while exhausting its reviewers is not sustainable.

The important word in this column is per, not cost. The numerator is what everyone enjoys arguing about, but what decides whether the report tells the truth is the denominator — and the denominator the industry uses by default is the token. The FinOps Foundation, the profession's central reference, lists AI cost KPIs as cost per token, per inference and per API call[6]. These figures are not wrong; they are exactly what an engineer tuning a system needs. But they cannot answer the board's question, because a falling cost per token while the number of jobs that actually complete also falls is a report in which the numbers improve and the business declines at the same time — and nothing in that report is inaccurate.

To be straight about attribution: the unit "total cost per successful outcome" belongs to the book, not to the FinOps Foundation. I cite that source to show that the industry's default meter really is the token, rather than a caricature I invented. And because it is the default, an organisation has to spend deliberate effort to change it. The correct denominator will never appear in a report on its own.

The word "total" has five lines in it

The numerator can be truncated too. Chapter 4 expands "total cost" to all of its lines in its metrics list: "value after model, control, review, incident, and rework cost" — value after the cost of the model, controls, review, incidents and rework. These five lines almost never appear in the same document, because each is recorded by a different system and lands on a different budget.

Cost line What is in it Where it usually hides
Model Model calls, input and output alike, including the retries when the first result is unusable Inside one organisation-wide bill, with no way to tell which job consumed what
Control Guards before an effect occurs, permission checks, and the recording and storage of traces Counted as central platform cost, so it is never tied back to the work that caused it
Review Reviewer minutes per item, multiplied by the real cost of people senior enough to review Treated as "work they were doing anyway", so it never enters the new system's accounts
Incident Investigation time, remedy for those affected, notification, and the period the service was down Inside an incident management system that never speaks to the cost system
Rework Work redone because the output was unusable, both upstream and downstream Landing on a downstream team that owns no AI budget, so nobody reports it

Chapter 3 adds a sixth line that I think matters most for deciding what to stop: evidence cost — the money and time spent proving the system works correctly. When the evidence cost of a job exceeds the value that job creates, the answer is not to reduce the evidence. The answer is to retire the job. Cutting evidence to make the number look better simply moves the cost into the Risk column and then stops measuring it.

The other half of the Economics column is reusable capability

The second half of the Economics line is reusable capability share, which connects straight into Chapter 5's AI and data factory. The book says the factory has six reusable services: data products provide governed facts; context services assemble instructions, retrieved evidence and memory; model services route tasks to the smallest adequate model and manage provider change; evaluation services hold labeled cases, judge calibration and test execution; tool registry services define scopes, schemas, effect classes and approval requirements; and observability services capture traces, outcomes, cost, drift and incidents.

And the closing sentence of that paragraph in the book is what makes this a single post rather than two topics stitched together: "Shared security, privacy, FinOps, and sustainability controls span the six." Security, privacy, FinOps and sustainability run across all six services. In Figure 6 of the book, the dark band beneath the factory reads SHARED ASSURANCE • SECURITY • PRIVACY • FINOPS • GREENOPS, drawn bilingually with the Thai gloss beside it. That is the structural answer to why cost and footprint belong in one post — in the book's architecture they are the same band running under every service, not two departments' separate projects.

The practical consequence is that Chapter 5's metric holds the two together in one unit from the start: "energy and cost per successful task" — energy and cost per successful task. If an organisation measures both with the same denominator, the argument about "saving money but burning energy", or the reverse, ends in numbers instead of ending in opinions.

4. Footprint — Read the Numbers Correctly, Then Measure Your Own

The book opens this topic with a sentence that fixes the position of the whole subject: "Sustainability belongs inside the business case." — sustainability lives in the business case, not in an annual report published after every decision has already been taken. And the glossary defines sustainable AI as "designing and operating AI with measured attention to energy, carbon, water, hardware, location, utilization, vendor efficiency, and societal value." Notice vendor efficiency in that list. It is the bridge the book itself throws back to section 2: in the book's own words, the contract question and the energy question are not separate.

Now the numbers. The IEA estimates that data centres worldwide used about 485 TWh of electricity in 2025 — a figure for all data centres, not for AI workloads alone — and projects roughly 950 TWh in 2030, about 3 percent of global electricity demand in that year[7]. Verified on 5 September 2026: the IEA's Key Questions on Energy and AI was published on 16 April 2026 and still carries this set of figures exactly as the book cites them. That report builds on the IEA's 2025 flagship, Energy and AI[8].

And the sentence that must always travel with all three figures is the book's own: "These are uncertain sector scenarios, not the meter for one use case." — uncertain sector scenarios, not the meter for a single use case. The IEA's own report notes that its central estimate remains close to the trajectory the 2025 report laid out, and that bottlenecks along the value chain are making the more aggressive scenarios less likely in the near term. That is the language of modelling, not the language of forecasting.

What you must never do with these numbers: never write "AI will consume 950 TWh", because the figure is for all data centres. Never divide it by a number of users, queries or tokens to obtain an energy figure per interaction. And never cite it as a reference value for any one provider. The IEA separates efficiency, adoption and model capability into three fast-moving, uncertain trends. Dividing a sector-level figure by anything to obtain a use-case-level figure is manufacturing data and then putting the IEA's name on it.

If a global figure cannot be a meter, what should an organisation measure? The book gives a list of five in one sentence: total energy, emissions where estimable, water where material, hardware lifecycle, and rebound when lower unit cost drives far more use. The first four are accounting. The fifth is behaviour, and it is the one that gives the other four their meaning.

Rebound is the mechanism that makes most sustainability reports comfortable reading for no good reason. The standard story runs like this. A team moves to a smaller model. Cost per call falls visibly. Energy per call falls with it. The slide goes to the board with a beautifully descending line. What the slide does not show is what happened next — once it is cheap, other teams switch it on too; work that was not worth doing becomes worth doing; call volume across the organisation grows faster than unit cost fell; and the absolute total, in both money and energy, ends up higher than before. Every per-unit number on that slide is correct.

Chapter 10's fifth principle compresses this into one sentence I would put on the cover of every report of this kind: "Unit efficiency can coexist with rising total impact". The word coexist was chosen carefully. The book does not say unit efficiency is a deception. It says the two numbers can be true at the same time, so a report that shows only one is not wrong but is incomplete. There is exactly one practical remedy: every time a per-unit figure is reported, the absolute total for the same period must be reported next to it.

One caution about attribution. The list of five things to measure belongs to the book, not to the OECD. What is verifiable in the OECD documents is the first value-based principle, "inclusive growth, sustainable development and well-being", and Principle 1.5 Accountability, which calls for a systematic risk management approach across the AI system lifecycle[5]. No text was found there requiring energy, carbon or water to be measured. So I cite the OECD for the lifecycle risk-management frame, and the book and the IEA for the list of what to measure.

5. Route to the Smallest Adequate Model — and Measure It

The rule the book gives is one self-contained sentence: "Route tasks to the smallest model that meets quality and risk requirements and compare footprint per successful outcome rather than token alone." The Thai companion renders it as "use the smallest model that meets the criteria, and compare footprint per successful outcome."

The words readers skip most often are that meets the requirements. This rule is not "use the smallest model", and it is not "use the cheapest model". It is a two-layer rule: the first layer is the quality criteria and the risk criteria, declared in advance; the second layer is choosing the smallest of those that clear the first. The order may not be swapped. An organisation that swaps it will get a system that is genuinely cheaper and genuinely wrong more often, and will pay the difference back in the review, incident and rework lines of section 3, where nobody counts it.

The book's architecture places this responsibility at a single point on purpose. The AI and data factory's model services do two things in one service: "route tasks to the smallest adequate model" and "manage provider change". That grouping is not an accident. Choosing model size and coping with a provider's version change are the same question asked at different times — both ask "what level of capability does this job need, and how do we know we are still getting it?"

The advice also has independent confirmation from outside the book. The FinOps Foundation states directly that organisations should "avoid using the most complex and expensive models for every task, as this often leads to unnecessary costs"[6]. The difference is that the book puts a risk-criteria layer in front, which a cost-focused document has no duty to carry.

A caution for anyone about to set a target: both "reusable capability share" and "share of tasks routed to the smallest adequate model" are metric definitions in the book, not values measured anywhere. The book gives no target numbers, and I will not invent any. The right value depends on the mix of work in that particular organisation. What can genuinely be measured is your own trend against your own previous quarter.

The measurement design table

The table below is this post's second worksheet, and it should be filled in before measurement starts, not after the first pile of data has arrived. The most important column is the fourth: a metric with no named owner becomes a slide nobody updates within two quarters. And the second column forces every row onto the same denominator — the successful outcome.

What we measure Unit (per successful outcome) Source Owner Review
Total cost of a successful outcome THB per successful outcome, all five cost lines included Observability (cost) + the accounting system + provider invoices Product owner, with FinOps Monthly
Evidence cost THB per successful outcome, always kept separate from model cost Reviewer time + testing cost + trace storage cost Domain owner Quarterly
Share of tasks routed to the smallest adequate model Percentage of successful requests, split by risk tier Model service routing logs Platform owner Monthly
Energy kWh per successful outcome, or the provider's reported figure with its method Provider reports + our own meters for what we run ourselves Platform owner Quarterly
Emissions where estimable Carbon units per successful outcome, with the estimation method and uncertainty range Energy figure × the carbon intensity of the region processing it Sustainability function Annually
Water where material Volume per successful outcome, only at locations where this genuinely matters Reports from the provider or the data centre we use Sustainability function Annually
Hardware lifecycle Service life and end destination of the equipment we own Asset register + the terms in the provider contract Infrastructure function Annually
Total footprint after usage growth Absolute total, against the same period last year Per-unit figure × actual organisation-wide usage The committee that oversees the portfolio Quarterly

The last row exists to prevent one mistake: reporting the first seven rows and stopping. All seven of those are per-unit figures, and by Chapter 10's fifth principle, an improving per-unit figure and a rising total can be true together. The eighth row is therefore not an extension. It is the row that makes the first seven readable honestly.

None of this happens in a vacuum — it belongs in the previous post's working session

Neither of this post's two topics needs a new forum. Both are steps in the Impact and obligation mapping working session, the same session post #16 One Evidence System already walked. That session has seven steps, and the fifth and sixth are ours — "Review vendor evidence and exit", and "Estimate energy and model-sizing options" — before reaching the seventh, deciding proceed, modify, pause or prohibit and naming the evidence gaps.

I like that placement a great deal, because it makes the supplier question and the energy question happen before the decision, rather than after the system is live in the form of a questionnaire circulated back for completion. And it means an organisation does not need a new working group to start on this. It needs two extra lines on an agenda that already exists.

A note on scope: the energy-per-query or energy-per-token figures circulating in general articles on the internet do not appear in this book, and I do not use them. What the book instructs is to measure your own, in units per successful outcome. That is a measurement design, not a constant that can be lifted from elsewhere and dropped in. If your provider reports an energy figure to you, always record their method in the Source column, because two providers using different methods cannot be compared.

6. Metrics That Matter

Chapter 10's metrics list has thirteen entries. The table below takes only those that concern suppliers, cost and footprint, then adds two from Chapter 3 and the board scorecard, with a Scorecard column as this series does, so that each metric's column on the board is visible.

This table leans visibly towards Economics and Risk, and that is correct for this post's role. Other posts in the series own the Value, Quality, People and Learning columns — the caution is not to lift this table and use it as a whole scorecard, because the book forbids collapsing the six columns into a single view from its first page.

Metric What it tells us Warning sign Scorecard
Vendor evidence The share of providers delivering the evidence the contract requires, and whether that evidence refers to the version actually running The most recent evidence is dated before the most recent version change — which makes it expired evidence Risk
Provider concentration The share of the portfolio hanging on a single provider or model family Every job the organisation cannot do without runs on the same provider, and nobody has ever assembled that picture Risk
Total cost per successful outcome The real cost of one usable outcome, after all five cost lines Reported as cost per token or per request, which can improve even as successful outcomes fall Economics
Evidence cost The money and time spent proving the system works correctly, separated from the cost of running it Higher than the value the system creates — a signal to retire it, not a signal to review less Economics
Reusable capability share The proportion of new work built from the factory's six existing services rather than built from scratch Every new project has its own pipeline, its own evaluation set and its own guards Economics
Smallest-adequate-model routing The share of tasks running on the smallest model that still meets the quality and risk criteria The number improves alongside rising re-review and escalation rates — that is lowering the bar, not routing Economics
Energy and estimated emissions per successful outcome The physical cost of one outcome, in the same unit as its financial cost Reported without the estimation method and uncertainty range, so it cannot be compared year on year Economics
Total footprint after usage growth The absolute total of energy and cost after usage volume has grown Never reported beside the per-unit figure, so nobody sees the rebound until it is too late Economics

Failure patterns

  • Assuming the vendor owns risk — the book's own words. It is a reasonable conclusion if we mistake what we bought for a utility service, but the duty to answer for whether a decision was right has never moved with an invoice.
  • Treating vendor testing as sufficient — Chapter 9's failure pattern. The provider's test results sit in the provider's context, not in our workflow, our language, our data and our users.
  • No exit clause — this one I derive from Chapter 10's fourth principle together with Chapter 3's fifth principle. It is not wording the book prints in its failure list, but the result is the same: a relationship we cannot leave is not a relationship we chose.
  • Cost per token as the KPI — derived from the book's instruction to compare footprint per successful outcome rather than per token alone. Per-token figures are useful to engineers, but they can improve while fewer jobs actually complete.
  • Efficiency claims without the total — Chapter 10's fifth principle directly. Improving unit efficiency and rising total impact can always live together.
  • Sustainability outside the business case — the inverse of "Sustainability belongs inside the business case". Once sustainability is a separate annual report, it can never change a decision in time.
  • Describing guidance or certification as proof of legal compliance — the book's own words, and the one that bites hardest in a supplier context, because a certificate is the thing a salesperson always has to hand.

Six of these seven can be fixed with documents — the contract in section 2 and the measurement table in section 5. The last one cannot, because it is a mistake about the type of the argument. A certificate says that one organisation operates a management system of a certain kind. Legal evidence says that a particular act in a particular jurisdiction is lawful. No quantity of the first converts into the second.

7. The Road Ahead

If I had to compress this whole post into one change of practice, I would choose changing the denominator. Change the question in the monthly report from "how much did we pay for models this month?" to "this month, what did one usable outcome cost in total and how much energy did it use, and what is the organisation-wide total against last year?" The first question produces price negotiations. The second produces decisions.

Three things you can do next week. One: take the largest AI supplier contract the organisation holds and walk the table in section 2 clause by clause, marking which are present and which are absent — in my experience clauses 3, 5 and 11 are most often empty together. Two: open one spreadsheet and fill in every cell of the table in section 5, especially the Owner column; the cells you cannot fill are the most useful answer this exercise produces. Three: take the latest AI cost report and look at what the denominator is. If it is tokens or requests, add one more line divided by the number of genuinely usable outcomes, and put the two lines side by side in next month's report.

All of this — a classified portfolio, redesigned workflows, an assembled factory, complete contracts and a meter that divides correctly — is still not a sequence an organisation can start on Monday morning. It is a list of what must exist, not a plan. What is missing is the timeline and a scope small enough to actually finish, and that is the whole subject of the next post.

🧭 Layer this post advances: the AI and data factory layer and the Operating model layer — this post answers leadership question Q7 ("who owns value, risk, impact, exceptions and learning"), which outsourcing makes hardest to answer, because the answer appears to have moved outside the organisation when it has not. On the organisational scorecard this post moves the Economics column chiefly (total cost per successful outcome, evidence cost, reusable capability share, energy per outcome, and the total after usage growth) and the Risk column secondarily (vendor evidence and provider concentration). The next post, #18 The First 90 Days — Align, Design, Build Before Scaling Anything, turns this whole list of what must exist into a timeline a real team can walk.

🎯 Key Takeaways

  • Accountability stays = you can outsource technology, but you cannot outsource accountability — the work travels with the contract, the duty to answer does not.
  • Eleven clauses + exit = data use, customer-data training, model or subprocessor change, evaluation evidence, logs, security tests, incident notice, audit support, continuity, IP and exit — the last one is the one most often missing.
  • Provider concentration = concentration at the provider is a portfolio-level risk, visible only from the portfolio, never from any single system's report.
  • Economics column = total cost per successful outcome and reusable capability share, where "total" means five lines: model, control, review, incident and rework.
  • 485 → 950 TWh = a sector scenario for all data centres, not the meter for any use case. Never divide it to obtain a per-query figure.
  • Per successful outcome = the same denominator for both cost and footprint, not per token, which can improve while fewer jobs succeed.
  • Rebound = a lower unit price drives usage up until the total rises, so every per-unit figure must be reported beside its absolute total.

References

Every source was verified on 5 September 2026 (Thailand time) · four evidence labels: Law statutory text or an official notice · Standard a published standard or official framework · Study research or field evidence · Synthesis the author's own synthesis or a non-research source.

  1. Synthesis Mingkhwan, A. AI Transformation as an Organizational Core — Bilingual Companion Playbook — Chapters 3, 4, 5, 9, 10, the executive brief, and Appendices C and D. The author's own manuscript, with no public URL; evidence snapshot 5 September 2026. Supports: the eleven-clause supplier list, Chapter 10's five operating principles, the scorecard's Economics column, provider concentration, the six services of the AI and data factory, the Impact and obligation mapping working session, and the metrics and failure-pattern lists
  2. Law European Union. Regulation (EU) 2024/1689 (Artificial Intelligence Act) — Articles 53 and 113 — checked against the consolidated text as at 27 July 2026. eur-lex.europa.eu — accessed 2026-09-05. Supports: the duties of providers of general-purpose AI models to draw up technical documentation and to supply information to downstream developers, and the fact that Article 53 was not amended by the 2026 amendment
  3. Law European Union. Regulation (EU) 2026/1744 of 8 July 2026 amending Regulation (EU) 2024/1689 — published in the OJ on 24 July 2026, in force 27 July 2026. eur-lex.europa.eu — accessed 2026-09-05. Supports: general application from 2 August 2026 and the postponement of the high-risk deadlines to 2 December 2027 (Annex III) and 2 August 2028 (Annex I)
  4. Standard ISO/IEC. ISO/IEC 42001:2023 — Information technology, Artificial intelligence, Management system — published December 2023. iso.org — accessed 2026-09-05. Supports: the existence and published status of the AI management system standard the book refers to — the description of the standard's content in this post is the book's wording, not drawn from the standard itself, and this page blocks automated retrieval
  5. Standard OECD. OECD AI Principles — adopted May 2019, updated May 2024 (OECD.AI Policy Observatory). oecd.ai — accessed 2026-09-05. Supports: the five value-based principles, Principle 1.1 inclusive growth, sustainable development and well-being, and Principle 1.5 Accountability, which requires a systematic risk management approach across the AI system lifecycle together with traceability
  6. Standard FinOps Foundation (Linux Foundation). FinOps for AI — the AI technology category of the FinOps Framework and the working group's overview material, updated 17 February 2026. finops.org — accessed 2026-09-05. Supports: that the AI cost KPIs in industry use are cost per token, per inference and per API call, and the recommendation to avoid using the most complex and expensive models for every task
  7. Study International Energy Agency. Key Questions on Energy and AI — Executive summary — published 16 April 2026. iea.org — accessed 2026-09-05. Supports: data-centre electricity of about 485 TWh in 2025 rising to about 950 TWh in 2030, or roughly 3 percent of global electricity demand, together with the framing that these are scenarios carrying substantial uncertainty
  8. Study International Energy Agency. Energy and AI — flagship report, published 10 April 2025. iea.org — accessed 2026-09-05. Supports: the context and lineage of the 2026 report, which builds on this one — no figure in this post is taken from this page
บทความจากซีรีส์ AI Transformation for Organizations 2026From the AI Transformation for Organizations 2026 series