Demonstrating Safety: Our Evidence Standards

A Provisional, Stage-Gated Framework for Evaluating Clinical AI in Utah’s Sandbox

Our Commitment to Clinical Integrity

Clinical AI holds potential to address critical care gaps and decrease administrative burdens in Utah, but responsible integration demands ongoing clinical evaluation. The Office of Artificial Intelligence Policy (OAIP) prioritizes a precautionary, safety-first oversight model that is designed to adapt as our pool of empirical clinical evidence grows.

Utah’s AI sandbox is organized to establish a structured, data-responsive environment to systematically evaluate technologies through real-world clinical practice and continuous monitoring. Static pre-deployment assessments of AI are useful, but can quickly become obsolete due to model, population, and behavioral drift. Thus, models of continuous post-deployment quality assurance may be better suited for measuring AI safety and efficacy.

Our Objectives: Reasonably Ensuring Safety

Our current evidentiary expectations are built around three primary hypotheses that sandbox participants are asked to evaluate and support through empirical data:

  1. Non-inferiority of care

    Evidence suggesting that the quality of care delivered via the AI-assisted pathway is not inferior to the care patients would traditionally receive from existing healthcare providers and systems.

  2. Bodily and psychological safety

    Continuous tracking to demonstrate, to the best of our current measuring capabilities, the absence of substantial bodily or psychological injury attributable to the AI tool.

  3. Potential for public-health benefits in Utah

    Measurable data suggesting how the tool might improve clinical quality, geographic access, or financial affordability for Utah residents relative to standard care paths.

To support these goals, we emphasize that patient participation in sandbox pilots must remain voluntary and informed. Under current protocols, patients engaging with an AI tool operating under our regulatory relief agreements should receive clear disclosures regarding the AI’s involvement, the preliminary and temporary nature of the sandbox authorization, and a direct mechanism to share feedback with our office.

By collaborating with participants to analyze de-identified pilot data, we aim to identify active, informative regulatory data streams. This cooperative approach is intended to help us construct an evidence-based regulatory pathway, allowing us to recommend adaptive, data-driven standards to the Utah Legislature as these technologies continue to unfold.

Evidentiary Expectations for Clinical Sandbox Applicants

To assist prospective applicants and clinical investigators, OAIP has created the following document. This document establishes our provisional negotiation posture, seeking to safeguard patient safety and clinical function quality while balancing technological innovation with structured, measurable, and evidence-based clinical evaluation. Prospective sandbox applicants and clinical investigators should review these expectations to understand our core requirements for clinical competency, continuous quality assurance, and public health benefit.

While this document cannot address every potential clinical use case, it represents our commitment to regulatory transparency by clearly outlining what we expect from sandbox participants. Over the course of these pilots, participants will be expected to continuously provide our office and the public with empirical data to support and evaluate their program’s safety and efficacy hypotheses.

Download Evidentiary Expectations for Healthcare AI Sandbox PDF

The document is available as a PDF: Evidentiary Expectations for Healthcare AI Sandbox.

A Collaborative, Expert-Driven Process

These expectations, authored by Alice Schwarze, PhD and Zach Boyd, PhD in August 2026, represent our office’s current synthesis of clinical and regulatory best practices. We recognize that clinical AI is an evolving science, and our understanding of safe implementation is constantly unfolding. We do not view these guidelines as permanent or immovable, but rather as our best current framework for ensuring patient safety while facilitating beneficial innovation.

We are grateful for the constructive criticism and expert consultation we received from numerous clinical specialists whose peer review significantly shaped these expectations. Recognizing that our empirical understanding of AI performance in clinical environments is constantly expanding, we anticipate that this framework will remain an iterative, living document. To that end, we invite medical professionals, clinical researchers, and patient advocates to continuously review our posture, challenge our assumptions, and collaborate with us in refining these safety standards.

What would you like to do next?

  • Browse our policy work

    Our current learning agenda, past research, and the findings we have published on specific AI policy questions.

    Browse
  • Explore Utah’s AI Sandbox

    How the AI Sandbox works, who qualifies, what an agreement requires, and which pilots are running now.

    Explore
  • Get answers to your questions

    Common questions about what this office regulates, what we’re learning, and what the sandbox means for consumers.

    Read FAQs
  • Engage with us

    Ask a question, share feedback on an AI policy issue, or find out how to join a focus group.

    Engage