What goes on an AI POC checklist? Five gates, each with an artifact, an owner and a pass condition, and a final decision (go, stop or re-scope) based on what the gates produced. AWS's guidance for generative AI proofs of concept names the problem: "Too often, PoCs are treated as technical demos that are designed to impress rather than as rigorous experiments that are designed to promote learning." NIST's voluntary AI Risk Management Framework 1.0 expects early work to produce knowledge "to inform an initial go/no-go decision about whether to design, develop, or deploy an AI system."
An AI POC checklist exists so that decision has its evidence on paper. Most of its artifacts carry over to the next use case, which turns a first POC into the start of an AI factory instead of a one-off project.
What an AI POC Checklist Is For
An AI POC checklist is a gate-by-gate list of the artifacts an AI proof of concept must produce, each with an owner and a pass condition, so the go/no-go decision rests on evidence and the work carries into production. Call it an AI proof of concept checklist or a POC plan; the test at every gate is the same. Which document exists, who signed it, and does it meet the condition set at the start?
Most published lists describe activities (define the goal, prepare data, pick a model, build, measure) and leave out who signs each step and what it leaves behind. The gates work the same way for predictive ML and generative AI; only the contents of the evaluation set differ. The step to a pilot with real users starts at Gate 4.
Which use case deserves a POC is decided earlier, in how use cases are scored and funded before a POC starts. For generative AI, the build method is in our guide to GenAI-specific exit criteria and architecture choices.
The AI POC Checklist, Gate by Gate
Five gates, each closing with one of three calls: pass, fix, or stop. Copy the table into your POC charter.
| Gate | Item | Artifact (evidence) | Owner | Next use case |
|---|---|---|---|---|
| Gate 0 | Business baseline | Baseline value, source and signature | Business owner | Reused as template |
| Gate 0 | Decision date and decision maker | One-page POC charter | Sponsor | Reused as template |
| Gate 0 | Stop thresholds | Limits for quality, latency and cost | Sponsor and tech lead | Reused as template |
| Gate 0 | Data access and privacy screening | Access granted; DPIA screening record | Data owner and DPO | Reused as asset |
| Gate 0 | Target environment | POC account on the production stack | Platform lead | Reused as asset |
| Gate 0 | Evaluation set owner | Named domain experts and a labeling plan | Domain lead | Stays with this use case |
| Gate 1 | End-to-end path | Deployed via the pipeline on real integrations | Tech lead | Reused as asset |
| Gate 1 | Cost per request | Usage and cost log from the first call | Tech lead | Reused as asset |
| Gate 1 | Feedback channel | Ratings stored with each output | Product owner | Reused as asset |
| Gate 2 | Evaluation result | Versioned report against Gate 0 thresholds | Tech lead and domain lead | Reused as template (harness only) |
| Gate 2 | Error analysis | Failures grouped by cause | Tech lead | Stays with this use case |
| Gate 2 | Context and risk record | Intended use, users, known limits, applicable rules | Risk owner | Reused as template |
| Gate 3 | Readiness review | Dependencies, monitoring, incident response, capacity, change process | Future operator | Reused as template |
| Gate 3 | Model card and runbook | Documentation for the operator | Tech lead | Reused as template |
| Gate 3 | Fallback and data export | Exit card: fallback option and export path | Architect | Reused as template |
| Gate 4 | Decision memo | Go, stop or re-scope, with reasons | Sponsor | Reused as asset |
Our checklist, not a standard. An item without its artifact is not done.
The last column separates a POC from the first increment of an AI factory: almost every item leaves a template or an asset for the next use case. An AI POC plan that lists activities without artifacts cannot be audited at the gate. Each gate gets a date agreed at Gate 0; data access, more than modeling, decides how long it takes.
Gate 0: What Must Be True Before Kickoff
Baseline and metric come before any model. Martin Zinkevich's Rules of Machine Learning for Google put them in that order: "Rule #1: Don't be afraid to launch a product without machine learning," and "Rule #2: First, design and implement metrics." A heuristic that beats the model is a valid POC result. The business owner signs the baseline; if your use case intake already recorded one, the POC inherits it.
The charter names a decision date, a decision maker and stop thresholds. AWS tells teams to "Set clear thresholds for quality, latency, and cost" and to pivot or end the POC when results miss them.
Data and law come first. The AWS default is "synthetic or fully anonymized data," with personal data only after "explicit approval from your information security and legal teams." In the EU, add a DPIA screening. The Article 29 Working Party guidelines on data protection impact assessment, endorsed by the EDPB, list nine criteria for high-risk processing, new technology among them, and state that "in most cases, a data controller can consider that a processing meeting two criteria would require a DPIA." An AI POC on personal data often meets one criterion already, so screening belongs here, not after the demo. Whether a full DPIA is required is for your data protection officer (DPO) to decide.
The POC runs on the stack production would use, and the evaluation set gets a named owner on the domain side; for generative AI, AWS notes, "preparation of an evaluation dataset replaces traditional curation of training data." Gate 0 is the part of an AI POC checklist most teams skip, and without it nothing later can be judged.
Gates 1 and 2: First Path, Then Evidence
Gate 1 asks for one working path from input to output, deployed through the real pipeline on real integrations, not from a notebook. Log the cost of every request from the first call. AWS defines unit economics as "the per-request costs, which includes token usage, compute resources, and storage" and says to track them early. Add a feedback channel; AWS suggests even "a minimal UI with positive and negative feedback ratings."
Gate 2 is the evidence. The evaluation result is compared with the Gate 0 thresholds, not with how the demo felt. Failures get grouped by cause, so the sponsor sees what a fix would cost. Then the context and risk record: NIST's MAP 1.1 asks that "intended purposes, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented." That record becomes the first draft of production documentation.
Gate 3: Production Readiness Belongs Inside the POC
Readiness gets checked inside the POC because a gap there is still cheap to close. Google's SRE book describes a Production Readiness Review meant to "Verify that a service meets accepted standards of production setup and operational readiness, and that service owners are prepared to work with SRE." A POC needs a reduced version covering dependencies, monitoring, emergency response, capacity and change management, run with the future operator.
For the model, The ML Test Score by Breck and colleagues at Google (2017) offers "28 specific tests and monitoring needs" as a rubric for production ML. It predates large language models, and a POC need not pass every test; it needs to know which ones it fails.
Documentation is an artifact too. A model card, per Mitchell and colleagues, is one of the "short documents accompanying trained machine learning models that provide benchmarked evaluation in a variety of conditions." A datasheet does the same job for the evaluation data. Add a runbook and a line on the fallback option and data export path.
Then the handover. In a 2022 interview study with 45 practitioners from 28 organizations, Nahar and colleagues found that most collaboration challenges center on "communication, documentation, engineering, and process." A survey of deployment case studies by Paleyes and colleagues adds that practitioners "face issues at each stage of the deployment process." Hence the item in our AI POC checklist teams resist most: the future operator signs Gate 3.
Gate 4: The Go/No-Go Decision, With Stop as a Valid Outcome
NIST's MANAGE 1.1 frames the question: "A determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed." In an AI POC checklist, that determination is an item with a signed document attached.
We allow three outcomes. Go comes with a named production owner and a budget. Stop comes with the reason in writing, so nobody revives the project without a new argument. Re-scope needs a new question and a new decision date.
Stop is a valid result. The GOV.UK Service Manual, written for UK public services, says of the alpha phase: "If you get to the end of your alpha and you're not confident you could do these things, you could stop altogether or decide to repeat discovery or alpha." An alpha is not a POC, but the logic transfers: a clear stop costs less than a vague go.
The decision memo fits on one page: results against the thresholds, cost per request, the readiness tests the POC failed, a recommendation and signatures.
What Your First AI Factory POC Leaves Behind
After Gate 4, publish a reuse register: the artifacts the next use case can start from. It holds the charter and threshold templates, the data access path, the screening record, the evaluation harness without domain data, the cost log, the model card template and the runbook. Use case two starts from that list. That is the practical difference between a series of POCs and an AI factory.
The register gets written even after a stop: artifacts from Gates 0 to 2 usually outlive their use case, so an AI POC checklist that ends in a documented stop still pays off.
At DS Stream we run the first use case on this checklist in the client's own cloud and hand over the reuse register with the decision memo. It is the first step toward an AI factory where use case two starts from these artifacts. Bring us the use case you want to put through Gate 0.
FAQ
What should an AI POC checklist include?
An AI POC checklist should include five gates: kickoff conditions (signed baseline, decision date, stop thresholds, data access and DPIA screening), a first end-to-end path with a cost log, evidence on the evaluation set, a production readiness review with a named operator, and a decision memo.
How do you set success criteria for an AI POC?
Write thresholds for quality, latency and cost before the work starts, as AWS recommends, and measure them against a signed baseline that includes what a simple heuristic achieves.
Who should sign off an AI POC go/no-go decision?
The business sponsor makes the call after the data owner, the data protection officer and the future operator sign their own items. A documented stop is a valid outcome, in line with NIST's MANAGE 1.1 and the GOV.UK alpha guidance.
What is the difference between a POC gate review and a production readiness review?
A production readiness review, as Google's SRE book describes it, checks whether a service can be handed over for operation. A POC gate review also asks whether the problem is worth solving. In our checklist it is Gate 3, in reduced form.
