Enterprise AI Programs Remain Stuck in Pilots

Regeneron’s clinical-development organization expanded from roughly five trials to more than 50 in six years, including more than 20 Phase 3 studies, creating operational pressure that could not be solved simply by adding staff.
In clinical document generation, Regeneron’s Kaniel Cassady said the main bottleneck was not writing documents such as protocols but the review cycles: “The biggest bottleneck in document generation is the review cycles.”
A practical AIoT architecture requires four linked layers—sensors, connectivity, a data platform and AI—and must account for operational contingencies such as what happens when the connection is interrupted.
The automated coding workflow described in the article uses a Kanban board to keep tasks, prompts and execution connected; moving a task into progress automatically starts sequential execution, with each task receiving the prior task’s results and a GitHub pull request produced for review.
The AI-agent framework presented in the planning chapter treats planning and feedback as distinct from merely giving an LLM tools: an agent must formulate a multistep plan, execute it toward a goal and learn from whether the plan worked, rather than depend on a human to advance each step.
Most enterprise AI projects never leave the pilot stage, stuck in endless testing loops because companies build technology before defining what problem it solves. McKinsey and industry leaders across clinical development, software engineering and IoT systems report the same pattern: success depends on identifying a concrete business outcome first, then selecting the right tools. Without this clarity, organizations waste months on proof-of-concept projects that never reach production.
Even well-designed AI systems fail without three critical ingredients: high-quality data, ongoing model evaluation after launch, and human review built into the workflow. Gartner research shows that 70% of AI initiatives stall because companies underestimate the operational changes needed to actually deploy the technology. The gap between a working demo and reliable production value is larger than most organizations expect.
Regeneron's clinical-development group faced acute operational pressure: trial volume jumped from roughly 5 trials to more than 50 in six years, including 20+ Phase 3 studies. Adding staff could not keep pace. The team deployed AI to generate clinical documents like protocols and regulatory submissions. But the real bottleneck was not document creation—it was the review cycle. Kaniel Cassady, a leader on the project, stated bluntly: "The biggest bottleneck in document generation is the review cycles." This insight forced a redesign: instead of automating writing, the team optimized the feedback loop and human sign-off process.
A functional IoT system powered by AI requires four tightly integrated components: sensors that collect data, connectivity to transmit it, a data platform to store and process it, and AI models that generate insights. IoT practitioners warn that skipping any layer or treating them as separate dooms the project. Real-world deployments fail not on the AI side but on basics: interrupted connections, sensor calibration drift, and missing fallback logic when the network goes down. A production system must specify what happens when connectivity fails—does the sensor buffer data, or does it shut down? These operational details determine success.
The difference between a chatbot and an autonomous agent is planning. A chatbot responds to a user prompt. An agent formulates a multistep plan, executes it toward a goal, evaluates whether it worked, and adjusts course. AI researchers emphasize that giving an LLM a list of tools is not enough; the agent must sequence dependent tasks, feed prior results into the next step, and learn from feedback. Without planning and feedback loops, automation fails on any problem requiring more than one action.
In automated software development, a concrete example illustrates this principle. A Kanban board tracks tasks and prompts. Moving a task into "in progress" automatically triggers sequential execution. Each task receives the output of the prior task. At the end, a GitHub pull request is generated for human review. This workflow converts an agent from a tool that executes isolated commands into a system that can plan and deliver production-ready code—but only because a human still decides whether to merge.
Pilot projects stall when organizations treat them as experiments rather than stepping stones to production. HBR research shows that 75% of AI pilots fail to scale because success was never defined. Teams built dashboards, ran models, and demonstrated value in a closed environment—then faced the hard question: "What operational change does this enable?" If the answer is vague, the project dies. The path forward requires ruthless clarity: identify the decision, process or bottleneck the AI solves, measure the impact in business terms, and build governance and human review into the production system from day one.
Publishers
18
Articles
3
Reach
21