What it does. Score any AI tool a vendor is pitching against 8 questions and get a Trial, Caution, or Walk verdict before you spend.
AI Tool Vetting Scorecard
Answer honestly about the tool being pitched
Verdict
Every box starts at No on purpose. Most tools fail at least one non-negotiable, so make the vendor earn each Yes.
Detailed Overview
Every week there is a new AI tool that will read your compressor, book your calls, or run your plant, and every pitch looks like a demo built to dazzle. This scorecard turns the gut-check into a structured one. You answer 8 plain questions about the tool in front of you, and it returns a verdict, a score out of 16, and a short list of exactly what to push the vendor on before any money moves.
Purpose
Most AI buying decisions in the trades happen on the strength of a slick demo and a confident salesperson. The demo shows the tool on its best, most-rehearsed question; the edge cases that break it never make the screen. This tool replaces the vibe with a repeatable test built around the same principle a good tech uses on a meter: do not trust a reading you cannot verify. Three of the eight questions are treated as non-negotiable. If a tool fails any one of them, it fails outright, no matter how well it scores on the rest. That keeps a strong demo from papering over a missing source citation, no human override, or a vendor that quietly owns your data.
When and Where to Use It
- A vendor is pitching an AI CSR, dispatcher, or diagnostic tool. Run the 8 questions live during the call and make them answer on the record.
- You got an AI pitch in your DMs. Score it before you reply, and use the push-list to separate a real company from a 150 dollar ghost.
- You are comparing two tools. Score each one and compare the verdicts and the gaps side by side.
- A tech was handed an AI app to use in the field. Score whether it shows its source and flags uncertainty before relying on it on a call.
How to Use It
- Open the scorecard with the tool or pitch in front of you.
- Answer each of the 8 questions Yes, Partly, or No. Every box starts at No on purpose, so the vendor has to earn each Yes.
- Read the verdict, the score, and the red-flag count.
- Work the push-list. Each No or Partly becomes a specific question to put to the vendor in writing.
- Use the Copy result button to drop the verdict and push-list into an email or a notes app for the rest of the team.
Outputs
- Recommendation. One of three verdicts: Worth a trial, Proceed with caution, or Walk. The verdict reflects both the total score and whether any non-negotiable is missing.
- Score. A point total out of 16 (Yes is 2, Partly is 1, No is 0) so you can compare tools and track whether a vendor closed your gaps on a follow-up.
- Red flags. A count of how many non-negotiables (source, human override, data ownership) are unmet. Any value above zero forces a Walk.
- Push-list. A tailored set of the exact questions to send back to the vendor, with the non-negotiable items marked.
Context: Where This Tool Lives in HKIA’s Content
The tool was built to accompany the AI-literacy pieces in the Tech Edition cycle on dismantling AI hype and learning to evaluate AI for yourself. Specifically:
- “Pattern Matcher, Not a Brain: What HVAC AI Actually Does”. argues that AI is pattern matching, not understanding, and closes with a five-question test for any AI pitch. This tool turns that test into a scored verdict.
- “Skill-Based Dispatch from the Tech’s Seat”. argues that AI dispatch lives or dies on data quality and guardrails, and that you should make a vendor show the escalation rules, override, and decision log live. This tool folds those buyer checks into the score.
It pairs naturally with a Speed-to-Lead Cost Calculator (which puts a dollar figure on slow lead response) and an AI Dispatch Readiness Scorecard (which grades your own shop’s data before you adopt a routing tool). All three share the same skeptical, evaluate-before-you-buy framing.
Math & Logic
- Scoring: each of the 8 questions scores Yes = 2, Partly = 1, No = 0, for a maximum of 16 points. The percentage is the point total divided by 16.
- Non-negotiables: three questions are flagged critical: source transparency, human-in-the-loop with override, and data ownership and export. These are drawn from the NIST AI Risk Management Framework’s emphasis on explainability, human oversight, and data governance.
- Verdict thresholds: any critical question answered No forces a Walk regardless of score. Otherwise, below 50 percent is a Walk, 50 to 74 percent (or any critical question answered Partly) is Proceed with caution, and 75 percent or higher with all non-negotiables clear is Worth a trial.
- Why a critical Partly caps at Caution: a half-answer on a non-negotiable is exactly where buyers get burned later, so even a high overall score is held to Caution until that item is a clear Yes in writing.
Limitations
- It scores your inputs, not the tool. Honest answers in, useful verdict out. A vendor’s confident “yes” is not the same as proof, which is why the push-list exists.
- It does not price the tool or model ROI. Cost, payback, and lead-response value are separate questions; pair this with a cost calculator for the dollar side.
- It is vendor and product agnostic. The questions apply to AI CSRs, dispatchers, diagnostic apps, and optimization tools alike, so it does not capture category-specific details like control-system compatibility.
Sources Used
- AI Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology, 2023. The govern, map, measure, and manage structure behind the non-negotiable questions on oversight, explainability, and data governance.
- Generative AI Profile (NIST-AI-600-1). National Institute of Standards and Technology, 2024. The generative-AI-specific risks (hallucination, data privacy, lack of explainability) that shape the source, uncertainty, and ownership questions.
