There is a number every board should have pinned above the AI budget line. MIT’s NANDA initiative, in The GenAI Divide: State of AI in Business 2025, examined more than 300 enterprise deployments and found 95% delivered no measurable P&L impact (fortune.com via Yahoo), against an estimated $30–40 billion of enterprise spend (forbes.com). Gartner, from a different direction, predicts over 40% of agentic initiatives will be cancelled by the end of 2027 (gartner.com) on grounds of cost, unclear value, and weak risk controls.

Read plainly, the dominant enterprise AI posture of the last three years, perpetual trialling, is a failure mode with a receipt attached. A deployment that has run for nine months without touching the P&L is a cost centre with good PR.

The trial trap has a structure

MIT’s research is specific about why the 95% stall. The initiatives that fail are disconnected from core workflow, cannot retain feedback or improve with use, and are measured on activity (usage, satisfaction, demos) rather than on financial outcomes. Meanwhile, the same research finds externally built solutions succeed roughly twice as often as internal builds, and budgets skew towards visible sales and marketing use cases while the strongest returns sit in unglamorous back-office operations (legal.io). The pattern is consistent: organisations fund visibility and call it strategy.

The 5% that generate value do the opposite. One workflow, deployed end to end, wired into systems of record, owned by an operator with a financial target, measured in the currency the CFO already uses.

The production standard

Moving to production-grade agentic operations means passing five gates, and refusing to fund anything that cannot.

  1. A named P&L owner. A business operator whose numbers move if the deployment works, and whose name is attached if it does not.
  2. A baseline taken before day one. Cost per case, cycle time, error rate, revenue per head. No baseline, no funding.
  3. Workflow integration, never workflow adjacency. The agent works inside the system of record, or the deployment does not proceed.
  4. Governance built in from the start. Decision rights, escalation paths and audit trails specified before go-live, closing off the “inadequate risk controls” failure Gartner identifies.
  5. A scale-or-stop date. A calendar date on which the deployment either standardises across the function or is switched off. Indefinite middle states are where the 95% lives.

Where AgentCraft sits

This is where AgentCraft comes in. Developed through Chesamel’s AI Innovation Lab, AgentCraft is our growing library of purpose-built AI agents designed to solve real operational and marketing challenges at scale.

We build agents around specific workflows, integrate them into how teams operate, and continually review and improve their performance. We can also build alongside clients, developing tailored agents around their business objectives, governance requirements and measurable outcomes.

The goal isn’t more AI activity. It’s a deployed capability that saves time, removes friction, and delivers measurable operational value.