The demo is over, but no one can approve deployment
The AI classifies customer inquiries fairly well. The results looked good in the presentation, and the business team responded positively. But once production deployment comes up, the questions pour in. How wrong can it be? Can it access customer data? Who turns it off if something goes wrong? There are no predetermined answers.
At this stage, tuning the model further will not solve the problem. The criteria for moving into production were not written before the pilot. This is also the practical shift proposed by Vellum’s AI transformation playbook: proceed through strategy, team, tools, pilots, and organizational rollout, while designing KPIs, accountability, data access, and approval criteria together from the very first stage.
NIST AI RMF likewise places GOVERN not as a final review step, but as a cross-cutting function spanning MAP, MEASURE, and MANAGE. Governance is less like a stamp received just before deployment and more like an operating approach that continuously decides what to build, who is responsible, and under what conditions to stop.
First, let go of the claim that “95% die”
Vellum’s original article cites MIT NANDA and states that 95% of enterprise generative AI pilots fail to produce measurable ROI. However, there is no verified basis for changing that statement into “95% failed at production deployment” or “the remaining 5% survived because of early governance”.
The currently verifiable official MIT NANDA page introduces the research group and its research directions, and directs readers to a separate access route for the report. Since the sample composition, definition of ROI, and method used to derive the 95% figure have not been directly verified in the publicly available official source, this number should not be used as a universal failure rate.
What to take away is not fear of 95%, but the sequence for making deployment decisions. Vellum’s five stages are a vendor-proposed playbook, not a controlled comparison of successful and unsuccessful companies. So rather than promising a particular probability of success, use it to establish promotion criteria for your own pilot.
Seven things to decide on one page before the pilot
Before a tool comparison chart, the document you need is a one-page operating agreement that connects “what must improve, by how much, and what risks require us to stop”. At minimum, the following seven items should be in one place.
| Item | What to document | What happens when it is missing |
|---|---|---|
| KPI baseline | Current processing time, cost, and error rate | Cannot compare whether improvement occurred |
| Target and timeline | What level to reach by when | The pilot continues indefinitely |
| KPI owner | The person who makes the final judgment on business outcomes | A good demo is treated as success |
| Data agreement | Source, owner, access, and retention conditions | The effort stops when it is time to connect real data |
| Separation of duties | Builder, approver, and deployer | The person who built it approves their own risk |
| Error budget | Acceptable errors and incidents that require an immediate stop | Cannot decide whether to continue when problems arise |
| Rollback path | How to return to the existing workflow and who has authority | Cannot safely turn it off after deployment |
This document does not need to be a long policy manual. NIST AI RMF is also a voluntary, outcome-focused framework that does not mandate a particular organizational chart or product. It is designed for organizations to select GOVERN, MAP, MEASURE, and MANAGE actions that fit their context of use and risk tolerance. If you are working with generative AI, you can use NIST’s Generative AI Profile as supplementary material for unique or amplified risks.
Start the first pilot with real actions turned off
You do not need to let AI approve refunds or send messages to customers from day one. Vellum proposes starting with shadow runs that send suggested results to a review queue, comparing SOPs and golden sets, then expanding to assist mode for a small user group and limited automated actions.
Organize the state with the current and target profile concepts in AI RMF 1.0, choose the actions you need from the NIST AI RMF Playbook, and prepare one workflow in the following order.
- Choose one workflow whose baseline you already know. You should already be tracking at least one of processing time, cost, or error rate. If you do not have a baseline, start recording it consistently this week before connecting AI.
- Put the value hypothesis and stop conditions in one sentence. For example: “Over 90 days, use agent-approved refund classification to reduce median handling time from 12 minutes to 9 minutes or less, while keeping the overall misclassification rate at or below the current 4%.” The 90 days and numbers are examples; set them to fit your organization’s work cycle.
- Separate authority over people and data. Document the data sources, owners, and permitted access methods, and designate the builder, business approver, privacy reviewer, and deployer. Even if one person holds multiple roles, the authority under which each decision was made must be distinguished.
- Run shadow mode with actions turned off. Build a golden set from normal cases and representative failure cases, and place AI results in a review queue rather than sending them to customers. Track the pass rate against the SOP and the types of errors.
- Open assist mode only after every gate is passed. Provide suggestion features only to a small user group after meeting the evaluation threshold, error budget, data controls, and owner approvals. Recheck the same gates whenever you expand the scope.
Target workflow: ______
Baseline / target / timeline: ______
KPI owner: ______
Data source, owner, and access conditions: ______
Builder / approver / deployer: ______
Pass thresholds and error budget: ______
Immediate stop conditions: ______
Rollback method and authorized person: ______
The first success is not the automation rate. The first operational achievement is reaching a state where the accountable person can review this document and the shadow-run results and approve opening a small-scale assist mode: “Under these conditions, we can proceed.” If even one of the baseline, data permissions, failure cases, logs, or rollback capability has not been verified, address that gap before improving the model.
If you want to go deeper
Complete 2026 AI Business Transformation Playbook You can review the five stages from strategy through organizational rollout, as well as examples of shadow runs and promotion gates. It is best read as a practical checklist while recognizing its limits as vendor guidance. vellum.ai
Artificial Intelligence Risk Management Framework (AI RMF 1.0) You can read the original source to understand why GOVERN spans the full lifecycle and how it relates to MAP, MEASURE, and MANAGE. nist.gov
NIST AI RMF Playbook You can explore specific actions to choose based on your organization’s context of use. nist.gov



