From Task Forces to Test Beds
How States Are Piloting AI
Once states establish basic readiness, they face the next challenge: experiment safely through well planned lab activities and pilots. AI pilots can either accelerate responsible learning—or if unorganized, create a fragmented landscape of disconnected tools, unclear risks, and missed lessons.
Code for America’s Government AI Landscape Assessment describes Piloting as the stage where governments move from policy conversations and training into hands-on experimentation. This includes AI innovation labs, sandboxes, pilot projects, and limited deployments with clear guardrails and evaluation timeframes. The Piloting stage can be seen as structured innovation where AI becomes practical. Government agencies learn to test, refine, question, and evaluate in real life use cases that can reveal operational challenges such as data limitations, ethical risks, workforce gaps, and procurement bottlenecks before states scale AI into real services.
Pilots help governments learn safely
In government, AI pilots are not just technology tests. They are institutional learning exercises. A good pilot can reveal:
Whether the use case is actually appropriate for AI
Whether staff understand the AI tooling
Whether the data is adequate to drive AI model learning and thereby effectiveness
Whether the AI system introduces bias or accuracy problems
Whether procurement rules need updating for a fast-evolving set of technologies
Whether residents or frontline workers experience benefits or harm
Whether the tool should be scaled, redesigned, or stopped
The public report notes that experimentation can surface operational challenges such as data limitations, infrastructure gaps, ethical risks, workforce needs, and procurement bottlenecks. That is why pilots should not be treated as public relations exercises (though some have seen more marketing than actual AI tool development). They should be designed to generate evidence for scaling what works.
The first wave: internal productivity
Many states began AI experimentation with internal productivity use cases. Staff used generative AI to summarize documents, draft communications, generate policy analyses, support administrative workflows, and reduce routine burdens. This makes sense. Internal productivity tools can offer relatively low-risk opportunities to learn how AI behaves in a government setting. They also help staff build confidence and judgment before AI is introduced into more sensitive public-facing workflows.
The report observes that many states started by experimenting with internal productivity tools framed as helping staff do their jobs more effectively. This creates a baseline comfortability with AI and awareness of some of its capabilities. This learning is vital for structuring outward facing AI tooling like custom chatbots and more focused internal AI tooling like agentic automated workflows.
But even these broad-based internal tools require guardrails. Government employees handle sensitive information. They need clear rules on what data can be entered into AI systems, how outputs should be verified, and when AI-generated content should or should not be used.
The risk of decentralized experimentation
Energy and creativity often emerge at the agency level. A human services agency may want to test document summarization. A transportation department may explore traffic analytics. A labor agency may test job-matching tools. A tax agency may examine fraud detection. That agency-level experimentation is valuable. But without coordination, it can become fragmented.
The report suggests that decentralized experimentation can produce uneven documentation, inconsistent risk review, and fragmented learning. This is one of the most important governance challenges states face. If every agency pilots AI differently without knowledge sharing across organizational silos, the state may miss the chance to build shared standards, reusable infrastructure, and cross-agency learning.
Sandboxes and innovation labs can help
Structured experimentation environments are one response to this challenge. AI sandboxes and innovation labs give agencies a safer way to test tools before wider deployment. They can create shared rules for data protection, vendor access, human oversight, evaluation, and documentation. The public report highlights formal AI sandboxes and pilot programs as mechanisms that generate better evidence.
These environments are especially valuable when AI touches benefits access, eligibility guidance, case management, or other high-impact services. They allow states to ask practical questions before scaling:
Does the system improve service quality?
Does it reduce administrative burden?
Does it work for people with limited English proficiency?
Does it produce different outcomes across demographic groups?
Can staff override or correct outputs?
Can residents understand when AI is involved?
Can the state explain and audit the system?
As a note, we partner with Social Finance to support state and local government with the AI Learning and Innovation Hub, an AI builder and sandbox for testing models: https://gov.iblueprint.ai.
Pilots should be portfolios, not one-offs
My view on this is that the next generation of government AI experimentation should look less like isolated pilots and more like coordinated portfolios with a range of AI implementation to full scale implementation. A portfolio approach allows states to compare use cases, standardize evaluation, share lessons, and decide where to invest next. It also helps leaders distinguish between flashy demonstrations and high-value public-service improvements.
A strong AI pilot portfolio might include:
Low-risk internal productivity pilots
Medium-risk operational pilots, such as document processing
High-impact public-service pilots with stronger oversight
Benefits access pilots focused on navigation and administrative burden reduction
Cross-agency pilots that test shared infrastructure
Evaluation plans that compare outcomes across use cases
What the 2026 evaluations show
The full report rates 19 states as Established and 4 as Advanced in Piloting. That makes Piloting the second-strongest stage after Readiness. But the report also notes that experimentation is often decentralized. Individual agencies explore use cases independently, which can create energy and creativity—but also inconsistent documentation, risk review, and learning.
The strongest Stage 2 evidence appears where states publish formal case studies, pilot evaluation reports, or structured sandbox environments with training and feedback channels. The report specifically points to Colorado, Pennsylvania, New Jersey, Arizona, and Utah as examples of stronger evidence for structured experimentation.
Case study: Connecticut’s AI Enablement Lab
Connecticut is a useful example of structured piloting. The public report notes that Connecticut established an internal AI Enablement Lab within its IT division to give agencies a safe place to pilot AI use cases with privacy protections. Around 20 AI use cases were deployed or in pilot by 2025, including ChatGPT and Microsoft Copilot trials, document summarization in child services, automated Q&A for resident inquiries, and predictive analytics in health care.
The important lesson is that Connecticut’s lab is not only a technical environment. It is an institutional learning space. It helps the state test use cases, protect data, and build confidence before broader deployment.
Case study: North Carolina’s analytics foundation
North Carolina shows how earlier investments in data analytics can prepare a state for the generative AI era. The public report highlights the state’s Government Data Analytics Center, which had already piloted traditional AI and machine learning before the generative AI wave. Its work included fraud detection for benefit claims and tax evasion, opioid crisis analytics to identify counties at risk of overdose spikes, and generative AI pilots such as document summarization for legislative research staff.
North Carolina’s case is important because it shows that AI maturity is cumulative. States with earlier analytics infrastructure and cross-agency data practices are better positioned to test newer AI tools.
Case study: Utah’s regulated AI Learning Lab
Utah offers one of the most distinctive piloting models. The public report notes that Utah launched the nation’s first Office of Artificial Intelligence Policy and an AI Learning Lab program designed to allow AI innovations to be piloted with private sector under close state oversight. In 2025, Utah approved a pilot with Doctronic to use AI for routine prescription renewals for chronic conditions. The pilot operates in a sandbox with strict monitoring, safety protocols, patient outcome review, and physician involvement for exceptions.
This is a high-stakes use case, and that is exactly why it is valuable as a case study. Utah is not treating AI experimentation as a free-for-all. It is creating a supervised regulatory environment where the state can learn from the private sector while maintaining oversight.
Case study: Pennsylvania’s pilot discipline
Pennsylvania is especially strong because its pilots include structured evaluation. The full report notes that Pennsylvania’s pilot evaluation reports are among the strongest evidence of Stage 2 maturity.
That matters because pilots without measurement can easily become demonstrations. Pennsylvania’s approach shows how experimentation can be designed around evidence: What changed? What improved? What risks appeared? What should scale?
Takeaway
Pilots are where AI becomes real. The full report warns that Stage 2 is the hinge in the AI journey. States that treat experimentation as disciplined learning can move confidently into implementation. States that experiment without structure risk fragmentation and stagnation. That is the key lesson for states: do not just pilot more. Pilot better. The best AI pilots are not technology showcases. States that define goals, document outcomes, involve cross-functional teams, and make clear go/no-go decisions will be better positioned to move from experimentation to real public value. The most successful states will be those that turn pilots into learning systems—not just scattered experiments.


