Tutorial
IntermediateIs Your AI Vendor Just a Wrapper? A Buyer's Guide
Learn to tell an AI wrapper from real AI capability, ask the five questions that reveal what you are buying, do the price math against direct models, and decide whether to keep, replace, or rebuild.
Time needed: About two hours to evaluate one vendor
Before you start:
- One current or proposed vendor contract (or sales proposal)
- Someone who can ask the vendor questions and expect answers
Vendors Procurement Governance
By the end of this tutorial you will be able to look at any “AI-powered” product your association pays for, or is being pitched, and answer three questions: what model is actually doing the work, what your data is worth to the vendor, and whether you should keep paying for the wrapper or go direct. No machine learning background required.
Start with the shape of the thing. Many products sold as AI are a thin layer: your data goes into the vendor’s interface, the vendor forwards it to someone else’s model, the answer comes back, and you pay the vendor a per-seat markup for the round trip. That is a wrapper. Wrappers are not fraud; a good one adds real workflow value, like the deterministic decision layer in our Laya tutorial. But a wrapper you cannot see through is a black box you are renting, and this tutorial teaches you to see through it.
The evaluation borrows its shape from NIST’s AI Risk Management Framework, the voluntary federal framework organized around four functions: Govern, Map, Measure, Manage. You do not need the whole framework. You need its instinct: map what the system is, measure what it does on your data, and govern the relationship before you sign.
Step 1: Draw the wrapper stack
On one page, draw what you believe happens when a staff member clicks the AI button: your data, the vendor’s interface, the model that answers, and the money. Then ask the vendor to correct your drawing. A vendor who can name the model, describe where your data travels, and explain the pricing in one breath is selling something real. A vendor who answers with “our proprietary AI” has told you something too: they either do not know or do not want you to know. Both are information.
Step 2: Ask what model powers it
This is the Map step. Ask for the model name and version, whether it is the vendor’s own or someone else’s, and what changes when the underlying model is updated. If it is someone else’s model, ask which one, because your data terms are partly their data terms. If the vendor will not say, treat the product as unmeasurable: you cannot evaluate what you cannot name, and NIST’s framework is explicit that validity is contextual, measured against the intended use, not the vendor’s benchmark slide.
Step 3: Ask where your data goes and what it trains
Ask four questions and write down the answers: Where is our data processed and stored? Is it used to train the vendor’s models or anyone else’s? Who at the vendor, or their subcontractors, can read it? What happens to it when we cancel? The third question surprises people; the fourth is the one vendors hope you skip. A product that cannot export your history, your prompts, and your configurations in a usable format is not a tool you use. It is a tenancy you occupy, and the rent goes up.
Step 4: Do the price math
This is the step that ends most wrapper relationships. A direct subscription to a capable model costs about $20 a month per power user at current prices, and most staff need no paid seat at all (see our three-routes guide). So ask: what are we paying per seat per year, how many seats actually use the AI features, and what would the direct model cost for those same people? If the wrapper costs ten times the model and adds a chat box, you are paying for the box. If it adds a real workflow, integrations with your AMS, audit trails, member-data handling your IT team approved, the markup may be worth it. The math does not decide. It clarifies what you are deciding about.
Step 5: Measure it on your data before you sign
NIST’s Measure function asks whether the system was evaluated for the use it actually sees, not a benchmark chosen by its builder. Run the vendor’s demo on your hardest real cases: the ambiguous member email, the event description with the tricky sponsor tier, the policy question with the exception. Do it three times; language models are non-deterministic, and our Laya tutorial exists because the same input can produce different outputs. If the vendor will not let you test on your data, or the answers wobble on the cases that matter, you have measured all you need to.
Step 6: Decide: keep, replace, or rebuild
Three outcomes, and all are respectable:
Keep the wrapper when it adds real workflow value: integrations, audit trails, support, and data handling your team has reviewed. Pay for the value, knowingly.
Replace with direct models plus a harness when the wrapper is a chat box with a markup. Our tutorials teach the direct route for drafting, triage, and redaction, and the local model lab runs it in the browser for free.
Rebuild the thin layer when the workflow is specific to your association and simple. A form that sends a prompt and files the answer is a weekend project for a developer, not a per-seat contract.
Step 7: Govern the relationship you keep
For the vendors you keep, write down the answers from steps 2 and 3, the price math from step 4, and the test results from step 5, and review them once a year. That one page is your Govern function: the moment the organization chose its AI posture on purpose. Vendors get acquired, models change under the interface, and data terms get rewritten. The annual review is how you notice.
The red flags
“Proprietary AI” with no named model. You cannot map what you cannot name.
No data-processing terms you can read. If the contract does not say where data goes and what it trains, assume the worst answer.
Per-seat pricing that dwarfs the underlying model cost with no workflow to show for it. Do the step 4 math out loud in the sales meeting.
Your data as the lock-in. If leaving means losing your history, prompts, and configurations, the product is a tenancy, not a tool.
No trial on your data. A vendor confident in the product wants you to measure it. A vendor selling the demo wants you to buy the demo.
The two-minute check
Pick your most expensive AI line item and run the four answers:
- What model does the work?
- Where does our data go, and what does it train?
- What do we pay per active user per year?
- What would the direct route cost?
If you can answer all four, you are governing. If you cannot answer the first two, you are renting a black box. This tutorial just gave you the questions; the answers are one vendor call away.
Sources
- NIST: AI Risk Management Framework: voluntary framework organized around four functions (Govern, Map, Measure, Manage), applicable to organizations of any size and sector; validity is contextual, measured against intended use.