AI Tool Vetting Scorecard

What it does. Score any AI tool a vendor is pitching against 8 questions and get a Trial, Caution, or Walk verdict before you spend.

AI Tool Vetting Scorecard

8 questions · based on NIST AI RMF

Answer honestly about the tool being pitched

Will it show its source for an answer (manual, bulletin, doc and revision)?Non-negotiable
Is it grounded in data fit for your work (manufacturer specs, your records), not just the open web?
Does it flag when it is unsure, instead of always answering with confidence?
Is there a clear line for what it does alone vs what needs your sign-off, with an override?Non-negotiable
Can you check its output faster than just doing the task yourself?
Does it tie into your field-service and accounting software without double entry?
Do you own your data, can you audit it, and export it if you cancel?Non-negotiable
Can the vendor show a real before and after or reference customers, not just a demo?
How it scores: three answers are non-negotiable. If a non-negotiable is a No, the tool fails no matter how it scores elsewhere. Source, human override, and data ownership are the ones you do not bend on.

Verdict

Recommendation
answer the 8 questions
Score
out of 16 points
Red flags
non-negotiables missing
Answer the questions above

Every box starts at No on purpose. Most tools fail at least one non-negotiable, so make the vendor earn each Yes.

Detailed Overview

Every week there is a new AI tool that will read your compressor, book your calls, or run your plant, and every pitch looks like a demo built to dazzle. This scorecard turns the gut-check into a structured one. You answer 8 plain questions about the tool in front of you, and it returns a verdict, a score out of 16, and a short list of exactly what to push the vendor on before any money moves.

Purpose

Most AI buying decisions in the trades happen on the strength of a slick demo and a confident salesperson. The demo shows the tool on its best, most-rehearsed question; the edge cases that break it never make the screen. This tool replaces the vibe with a repeatable test built around the same principle a good tech uses on a meter: do not trust a reading you cannot verify. Three of the eight questions are treated as non-negotiable. If a tool fails any one of them, it fails outright, no matter how well it scores on the rest. That keeps a strong demo from papering over a missing source citation, no human override, or a vendor that quietly owns your data.

When and Where to Use It

  • A vendor is pitching an AI CSR, dispatcher, or diagnostic tool. Run the 8 questions live during the call and make them answer on the record.
  • You got an AI pitch in your DMs. Score it before you reply, and use the push-list to separate a real company from a 150 dollar ghost.
  • You are comparing two tools. Score each one and compare the verdicts and the gaps side by side.
  • A tech was handed an AI app to use in the field. Score whether it shows its source and flags uncertainty before relying on it on a call.

How to Use It

  1. Open the scorecard with the tool or pitch in front of you.
  2. Answer each of the 8 questions Yes, Partly, or No. Every box starts at No on purpose, so the vendor has to earn each Yes.
  3. Read the verdict, the score, and the red-flag count.
  4. Work the push-list. Each No or Partly becomes a specific question to put to the vendor in writing.
  5. Use the Copy result button to drop the verdict and push-list into an email or a notes app for the rest of the team.

Outputs

  • Recommendation. One of three verdicts: Worth a trial, Proceed with caution, or Walk. The verdict reflects both the total score and whether any non-negotiable is missing.
  • Score. A point total out of 16 (Yes is 2, Partly is 1, No is 0) so you can compare tools and track whether a vendor closed your gaps on a follow-up.
  • Red flags. A count of how many non-negotiables (source, human override, data ownership) are unmet. Any value above zero forces a Walk.
  • Push-list. A tailored set of the exact questions to send back to the vendor, with the non-negotiable items marked.

Context: Where This Tool Lives in HKIA’s Content

The tool was built to accompany the AI-literacy pieces in the Tech Edition cycle on dismantling AI hype and learning to evaluate AI for yourself. Specifically:

  • “Pattern Matcher, Not a Brain: What HVAC AI Actually Does”. argues that AI is pattern matching, not understanding, and closes with a five-question test for any AI pitch. This tool turns that test into a scored verdict.
  • “Skill-Based Dispatch from the Tech’s Seat”. argues that AI dispatch lives or dies on data quality and guardrails, and that you should make a vendor show the escalation rules, override, and decision log live. This tool folds those buyer checks into the score.

It pairs naturally with a Speed-to-Lead Cost Calculator (which puts a dollar figure on slow lead response) and an AI Dispatch Readiness Scorecard (which grades your own shop’s data before you adopt a routing tool). All three share the same skeptical, evaluate-before-you-buy framing.

Math & Logic

  • Scoring: each of the 8 questions scores Yes = 2, Partly = 1, No = 0, for a maximum of 16 points. The percentage is the point total divided by 16.
  • Non-negotiables: three questions are flagged critical: source transparency, human-in-the-loop with override, and data ownership and export. These are drawn from the NIST AI Risk Management Framework’s emphasis on explainability, human oversight, and data governance.
  • Verdict thresholds: any critical question answered No forces a Walk regardless of score. Otherwise, below 50 percent is a Walk, 50 to 74 percent (or any critical question answered Partly) is Proceed with caution, and 75 percent or higher with all non-negotiables clear is Worth a trial.
  • Why a critical Partly caps at Caution: a half-answer on a non-negotiable is exactly where buyers get burned later, so even a high overall score is held to Caution until that item is a clear Yes in writing.

Limitations

  • It scores your inputs, not the tool. Honest answers in, useful verdict out. A vendor’s confident “yes” is not the same as proof, which is why the push-list exists.
  • It does not price the tool or model ROI. Cost, payback, and lead-response value are separate questions; pair this with a cost calculator for the dollar side.
  • It is vendor and product agnostic. The questions apply to AI CSRs, dispatchers, diagnostic apps, and optimization tools alike, so it does not capture category-specific details like control-system compatibility.

Sources Used

  • AI Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology, 2023. The govern, map, measure, and manage structure behind the non-negotiable questions on oversight, explainability, and data governance.
  • Generative AI Profile (NIST-AI-600-1). National Institute of Standards and Technology, 2024. The generative-AI-specific risks (hallucination, data privacy, lack of explainability) that shape the source, uncertainty, and ownership questions.
Share this Tool on:
Follow us on:

Save 6% on purchases at TruTech Tools with code knowitall (excluding Fluke and Flir products)

Save 8% at eMotors Direct with code HVACKNOWITALL

Subscribe Now!

Subscribe now and stay up to date with the latest industry trends and HVAC tips and tricks!

Subscribe Now!

Subscribe now and stay up to date with the latest industry trends and HVAC tips and tricks!