How to Select an AI Vendor in Healthcare
Most healthcare AI purchases fail before the contract is signed. The selection process, not the technology, is usually the problem.
Healthcare organizations are now fielding dozens of AI vendor pitches a quarter: ambient scribes, prior-auth automation, real-world data analytics, imaging triage, revenue-cycle copilots. The demos are uniformly impressive. That is precisely the problem. In a market where every vendor can produce a compelling demo on curated data, demo quality carries almost no signal. Selection has to be built on things vendors cannot stage.
This post lays out a practical framework we use with payers, health systems, and investors: five evaluation dimensions, the diligence questions that actually discriminate between vendors, and the failure modes that show up twelve months after go-live.
Start with the problem, not the vendor
The single most common selection error is running a "which AI vendor should we pick?" process instead of a "what operational problem are we solving, and is AI the right tool?" process. Before any RFP goes out, you should be able to state:
The baseline. What does this workflow cost today, in dollars, FTE hours, denial rates, or turnaround time? If you can't quantify the current state, you will never be able to prove the vendor improved it, and neither will they.
The decision the AI is changing. Is the model informing a human decision (decision support), replacing a decision (automation), or generating work product a human edits (drafting)? These carry radically different risk profiles, regulatory exposure, and validation requirements. A vendor that blurs this distinction in their own materials is telling you something.
The kill criteria. What result at 90 days would make you walk away? Agreeing on this internally before vendor conversations start is the cheapest insurance you can buy.
The five dimensions that matter
1. Clinical and operational validation, on data that looks like yours
Peer-reviewed publications are nice. What you actually need is evidence of performance on populations, payer mixes, coding patterns, and documentation styles similar to your own. Model performance in healthcare degrades across sites: a sepsis model tuned in one health system can miss badly in another, and an NLP pipeline trained on academic-center notes can stumble on community-hospital documentation.
Discriminating questions: How many external validation sites, and how different were they from the development site? What did performance look like at the worst site? Will you support a silent-mode pilot on our data before we commit?
A vendor confident in their model will run in shadow mode on your data. A vendor who resists is quoting you their best site.
2. Workflow integration, not just interoperability
"We integrate with Epic" is a checkbox. The real questions are where the output lands, who sees it, and how many clicks it adds. AI tools in healthcare fail operationally far more often than they fail statistically: alert fatigue, output that arrives after the decision was already made, results dumped into a queue nobody owns.
Discriminating questions: Show us the tool inside the actual workflow of the role that will use it, not a standalone portal. What is the median additional time per encounter or per claim? What percentage of your live customers' users touch the tool weekly six months after go-live?
That last number, sustained weekly active usage among the intended users, is the single most honest metric in this market. Vendors track it. Few volunteer it.

3. Data rights, privacy, and security posture
In healthcare the data questions are existential, not procedural. At minimum: a signed BAA covering every subprocessor that touches PHI (including the underlying model provider, if the vendor is building on a foundation model); clarity on whether your data trains their models, and whether you can opt out without losing functionality; data residency, retention, and deletion terms; and SOC 2 Type II plus HITRUST or equivalent, with recent pen-test results available under NDA.
The recurring failure mode here is the subprocessor chain. A vendor can be impeccably compliant while a model API in their stack is not covered by a BAA. Map the full data flow, every hop, before signing.
4. Regulatory and liability clarity
Know whether the product is, or should be, regulated as Software as a Medical Device (SaMD), and whether the vendor's claims match their regulatory status. A vendor marketing "diagnostic" capability with no FDA pathway and a disclaimer that it's "for informational purposes only" is transferring their regulatory risk to you.
Contract-side, look for: indemnification that covers model-output errors (not just IP and data breach); clear allocation of liability when a clinician acts on a wrong output; and audit rights over model changes. Ask how model updates are communicated and validated. A silent model swap that shifts performance is a patient-safety and compliance event, not a routine release.
5. Vendor durability and switching costs
Healthcare AI is consolidating fast. Many of today's vendors will be acquired, repriced, or gone within your contract term. Assess runway and revenue quality, not just funding announcements; the customer roster's retention, not its logos; and, critically, your exit. What do you get back if you leave? In what format? Does your fine-tuned configuration, prompt library, or labeled data leave with you?
The strategic question underneath: is this vendor a durable capability or a feature your EHR or clearinghouse will ship natively in 24 months? Paying enterprise-platform prices for a soon-to-be-commodity feature is one of the quietest ways to destroy ROI in this market.
Weight the dimensions before you look at vendors
Different buyers should weight these dimensions very differently. A payer automating prior-auth intake is not making the same bet as a health system deploying clinical decision support. Fix the weights before scoring vendors, or the weights will quietly bend toward whichever vendor gave the best demo.
A reasonable starting point for a clinical-adjacent deployment: validation 30%, workflow integration 25%, data and security 20%, regulatory and liability 15%, vendor durability 10%. For back-office automation, integration and durability rise while regulatory weight falls. The point is not the specific numbers, it's that the argument about what matters happens before the beauty contest, not after.

Run the process like diligence, not procurement
A selection process that reliably separates real vendors from demo-ware looks like this:
Problem definition
Baseline metrics, decision type, kill criteria, dimension weights. No vendor contact.
Structured evaluation
Same scenario and same de-identified sample data given to every finalist. Score against the weighted rubric. Reference calls with churned customers, not just current ones; ask the vendor for one, and treat how they respond as data.
Paid, bounded pilot
Silent mode first, then limited live use. Pre-agreed success metrics tied to your baseline. Pilot pricing that doesn't obligate an enterprise contract.
Enforce the kill criteria
Enforced by someone other than the project's internal champion. Sunk-cost momentum is how mediocre pilots become five-year contracts.
The bottom line
The healthcare AI market rewards buyers who can distinguish evidence from theater. Vendors control the demo; you control the baseline, the pilot design, and the kill criteria. Anchor the process on validation against your own data, sustained usage in the real workflow, a fully mapped PHI chain, contractual clarity on liability, and a sober view of whether the vendor will exist, as an independent capability worth paying for, at renewal.
The best vendors will welcome this process. The rest will tell you it's not how they usually work. Believe them.