AI pilot readiness scorecard
Template · decision scorecard
Decide whether an AI pilot should go wider, be extended, or stop, against a bar you wrote down before it started.
Most AI pilots hit the same wall a few weeks in: is this working well enough to roll out further, or should it be extended, changed, or killed? This scorecard replaces one manager’s gut feeling with a defined bar. Fill in the “Bar” column before the pilot starts.
Pilot
| Field | Answer |
|---|---|
| Tool / use case | |
| Pilot group and size | |
| Start date / review date | |
| Decision owner |
1. Output quality
Define three or four concrete criteria and score a sample of real outputs against each (1–5). A second model can do the scoring (“LLM-as-judge”): it reads the output and the criteria and returns a score with a one-line justification.
| Criterion | Bar | Average score | Notes |
|---|---|---|---|
| e.g. Followed the existing code style and conventions | |||
| e.g. Did not introduce an obvious regression or bug | |||
| e.g. Saved the reviewer meaningful time versus doing it by hand | |||
A minimal scorer:
import json
RUBRIC = [
"Followed the existing code style and conventions",
"Did not introduce an obvious regression or bug",
"Would have saved the reviewer meaningful time versus writing it by hand",
]
def score_output(judge_model, sample_output, task_context):
prompt = f"""You are scoring an AI coding assistant's output against a rubric.
Task context: {task_context}
Output to score: {sample_output}
Rubric: {json.dumps(RUBRIC)}
For each rubric item, return a 1-5 score and a one-sentence justification, as JSON."""
response = judge_model.generate(prompt)
return json.loads(response)
2. Cost per use
Pull cost from the provider’s billing data rather than estimating it, and divide by actual usage, not seat count.
| Measure | Bar | Actual |
|---|---|---|
| Total cost for the period | ||
| Licensed seats | ||
| Active users | ||
| Uses in the period | ||
| Cost per use |
3. Adoption
Week-one growth just means people are curious. What matters is whether usage holds or grows by week three or four.
| Week | Active users | Uses |
|---|---|---|
| 1 | ||
| 2 | ||
| 3 | ||
| 4 |
Decision
| Signal | Weight | Met the bar? |
|---|---|---|
| Output quality | ||
| Cost per use | ||
| Adoption (weeks 3–4) |
Recommendation: Go / Extend the pilot / No-Go
Reasoning (the part a manager needs to defend the decision to their own leadership):