AI POC Checklist for Your First AI Factory Use Case

Paweł Szczepanik
Paweł Szczepanik
September 28, 2026
•
9 min read
Loading the Elevenlabs Text to Speech AudioNative Player...

What goes on an AI POC checklist? Five gates, each with an artifact, an owner and a pass condition, and a final decision (go, stop or re-scope) based on what the gates produced. AWS's guidance for generative AI proofs of concept names the problem: "Too often, PoCs are treated as technical demos that are designed to impress rather than as rigorous experiments that are designed to promote learning." NIST's voluntary AI Risk Management Framework 1.0 expects early work to produce knowledge "to inform an initial go/no-go decision about whether to design, develop, or deploy an AI system."

An AI POC checklist exists so that decision has its evidence on paper. Most of its artifacts carry over to the next use case, which turns a first POC into the start of an AI factory instead of a one-off project.

What an AI POC Checklist Is For

An AI POC checklist is a gate-by-gate list of the artifacts an AI proof of concept must produce, each with an owner and a pass condition, so the go/no-go decision rests on evidence and the work carries into production. Call it an AI proof of concept checklist or a POC plan; the test at every gate is the same. Which document exists, who signed it, and does it meet the condition set at the start?

Most published lists describe activities (define the goal, prepare data, pick a model, build, measure) and leave out who signs each step and what it leaves behind. The gates work the same way for predictive ML and generative AI; only the contents of the evaluation set differ. The step to a pilot with real users starts at Gate 4.

Which use case deserves a POC is decided earlier, in how use cases are scored and funded before a POC starts. For generative AI, the build method is in our guide to GenAI-specific exit criteria and architecture choices.

The AI POC Checklist, Gate by Gate

Five gates, each closing with one of three calls: pass, fix, or stop. Copy the table into your POC charter.

GateItemArtifact (evidence)OwnerNext use case
Gate 0Business baselineBaseline value, source and signatureBusiness ownerReused as template
Gate 0Decision date and decision makerOne-page POC charterSponsorReused as template
Gate 0Stop thresholdsLimits for quality, latency and costSponsor and tech leadReused as template
Gate 0Data access and privacy screeningAccess granted; DPIA screening recordData owner and DPOReused as asset
Gate 0Target environmentPOC account on the production stackPlatform leadReused as asset
Gate 0Evaluation set ownerNamed domain experts and a labeling planDomain leadStays with this use case
Gate 1End-to-end pathDeployed via the pipeline on real integrationsTech leadReused as asset
Gate 1Cost per requestUsage and cost log from the first callTech leadReused as asset
Gate 1Feedback channelRatings stored with each outputProduct ownerReused as asset
Gate 2Evaluation resultVersioned report against Gate 0 thresholdsTech lead and domain leadReused as template (harness only)
Gate 2Error analysisFailures grouped by causeTech leadStays with this use case
Gate 2Context and risk recordIntended use, users, known limits, applicable rulesRisk ownerReused as template
Gate 3Readiness reviewDependencies, monitoring, incident response, capacity, change processFuture operatorReused as template
Gate 3Model card and runbookDocumentation for the operatorTech leadReused as template
Gate 3Fallback and data exportExit card: fallback option and export pathArchitectReused as template
Gate 4Decision memoGo, stop or re-scope, with reasonsSponsorReused as asset

Our checklist, not a standard. An item without its artifact is not done.

The last column separates a POC from the first increment of an AI factory: almost every item leaves a template or an asset for the next use case. An AI POC plan that lists activities without artifacts cannot be audited at the gate. Each gate gets a date agreed at Gate 0; data access, more than modeling, decides how long it takes.

Gate 0: What Must Be True Before Kickoff

Baseline and metric come before any model. Martin Zinkevich's Rules of Machine Learning for Google put them in that order: "Rule #1: Don't be afraid to launch a product without machine learning," and "Rule #2: First, design and implement metrics." A heuristic that beats the model is a valid POC result. The business owner signs the baseline; if your use case intake already recorded one, the POC inherits it.

The charter names a decision date, a decision maker and stop thresholds. AWS tells teams to "Set clear thresholds for quality, latency, and cost" and to pivot or end the POC when results miss them.

Data and law come first. The AWS default is "synthetic or fully anonymized data," with personal data only after "explicit approval from your information security and legal teams." In the EU, add a DPIA screening. The Article 29 Working Party guidelines on data protection impact assessment, endorsed by the EDPB, list nine criteria for high-risk processing, new technology among them, and state that "in most cases, a data controller can consider that a processing meeting two criteria would require a DPIA." An AI POC on personal data often meets one criterion already, so screening belongs here, not after the demo. Whether a full DPIA is required is for your data protection officer (DPO) to decide.

The POC runs on the stack production would use, and the evaluation set gets a named owner on the domain side; for generative AI, AWS notes, "preparation of an evaluation dataset replaces traditional curation of training data." Gate 0 is the part of an AI POC checklist most teams skip, and without it nothing later can be judged.

Gates 1 and 2: First Path, Then Evidence

Gate 1 asks for one working path from input to output, deployed through the real pipeline on real integrations, not from a notebook. Log the cost of every request from the first call. AWS defines unit economics as "the per-request costs, which includes token usage, compute resources, and storage" and says to track them early. Add a feedback channel; AWS suggests even "a minimal UI with positive and negative feedback ratings."

Gate 2 is the evidence. The evaluation result is compared with the Gate 0 thresholds, not with how the demo felt. Failures get grouped by cause, so the sponsor sees what a fix would cost. Then the context and risk record: NIST's MAP 1.1 asks that "intended purposes, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented." That record becomes the first draft of production documentation.

Gate 3: Production Readiness Belongs Inside the POC

Readiness gets checked inside the POC because a gap there is still cheap to close. Google's SRE book describes a Production Readiness Review meant to "Verify that a service meets accepted standards of production setup and operational readiness, and that service owners are prepared to work with SRE." A POC needs a reduced version covering dependencies, monitoring, emergency response, capacity and change management, run with the future operator.

For the model, The ML Test Score by Breck and colleagues at Google (2017) offers "28 specific tests and monitoring needs" as a rubric for production ML. It predates large language models, and a POC need not pass every test; it needs to know which ones it fails.

Documentation is an artifact too. A model card, per Mitchell and colleagues, is one of the "short documents accompanying trained machine learning models that provide benchmarked evaluation in a variety of conditions." A datasheet does the same job for the evaluation data. Add a runbook and a line on the fallback option and data export path.

Then the handover. In a 2022 interview study with 45 practitioners from 28 organizations, Nahar and colleagues found that most collaboration challenges center on "communication, documentation, engineering, and process." A survey of deployment case studies by Paleyes and colleagues adds that practitioners "face issues at each stage of the deployment process." Hence the item in our AI POC checklist teams resist most: the future operator signs Gate 3.

Gate 4: The Go/No-Go Decision, With Stop as a Valid Outcome

NIST's MANAGE 1.1 frames the question: "A determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed." In an AI POC checklist, that determination is an item with a signed document attached.

We allow three outcomes. Go comes with a named production owner and a budget. Stop comes with the reason in writing, so nobody revives the project without a new argument. Re-scope needs a new question and a new decision date.

Stop is a valid result. The GOV.UK Service Manual, written for UK public services, says of the alpha phase: "If you get to the end of your alpha and you're not confident you could do these things, you could stop altogether or decide to repeat discovery or alpha." An alpha is not a POC, but the logic transfers: a clear stop costs less than a vague go.

The decision memo fits on one page: results against the thresholds, cost per request, the readiness tests the POC failed, a recommendation and signatures.

What Your First AI Factory POC Leaves Behind

After Gate 4, publish a reuse register: the artifacts the next use case can start from. It holds the charter and threshold templates, the data access path, the screening record, the evaluation harness without domain data, the cost log, the model card template and the runbook. Use case two starts from that list. That is the practical difference between a series of POCs and an AI factory.

The register gets written even after a stop: artifacts from Gates 0 to 2 usually outlive their use case, so an AI POC checklist that ends in a documented stop still pays off.

At DS Stream we run the first use case on this checklist in the client's own cloud and hand over the reuse register with the decision memo. It is the first step toward an AI factory where use case two starts from these artifacts. Bring us the use case you want to put through Gate 0.

FAQ

What should an AI POC checklist include?

An AI POC checklist should include five gates: kickoff conditions (signed baseline, decision date, stop thresholds, data access and DPIA screening), a first end-to-end path with a cost log, evidence on the evaluation set, a production readiness review with a named operator, and a decision memo.

How do you set success criteria for an AI POC?

Write thresholds for quality, latency and cost before the work starts, as AWS recommends, and measure them against a signed baseline that includes what a simple heuristic achieves.

Who should sign off an AI POC go/no-go decision?

The business sponsor makes the call after the data owner, the data protection officer and the future operator sign their own items. A documented stop is a valid outcome, in line with NIST's MANAGE 1.1 and the GOV.UK alpha guidance.

What is the difference between a POC gate review and a production readiness review?

A production readiness review, as Google's SRE book describes it, checks whether a service can be handed over for operation. A POC gate review also asks whether the problem is worth solving. In our checklist it is Gate 3, in reduced form.

Share this post
Artificial Intelligence
Paweł Szczepanik
MORE POSTS BY THIS AUTHOR
Paweł Szczepanik

Curious how we can support your business?

TALK TO US

More insights

More news

No items found.