Before You Sign
Twelve questions to ask an AI implementation partner before you sign, grouped by what they protect you from. Four of them would disqualify us.
Arrête de tout configurer. Place à la construction.
Des templates SaaS avec orchestration IA.
Problem: Four vendors are pitching you an AI project. All four demos worked. You cannot tell which one hands you a working thing and which one hands you a deck, and somewhere between five and sixteen people inside your own company have to agree before you sign.
Quick Win: An AI implementation partner is an outside firm you pay to build and install an AI system inside your company, as opposed to selling you software you run yourself. The twelve questions below are contract-shaped: each exists to force one specific promise into writing before money moves. Know the odds first. Gartner estimates only about 130 of the thousands of vendors now claiming to sell AI agents (software that takes steps on its own instead of waiting for a click) actually have them. The rest are rebranded chatbots, assistants, and old-style rule-following automation (Gartner).
The Checklist Is Really an Internal Alignment Tool
The checklist's job is not to grade the vendor. It is to get your own people to agree on what "done" means before you pay someone to deliver it.
Gartner surveyed 632 business buyers and found 74% of buying teams show what it calls "unhealthy conflict" during the decision: members holding conflicting objectives, disagreeing on the right course of action, or being overruled by someone outside the group. Those groups now run from five to sixteen people across as many as four departments, and the ones that reach agreement are 2.5 times more likely to say the deal they made was a good one (Gartner).
So the failure to fear is not a bad vendor. It is your CFO, your head of operations, and your IT lead each signing off on a different definition of success, then finding the gap in month four when nobody will accept the work. Run the twelve questions internally first. Wherever your own people give different answers, that is the clause that has to be written in painful detail.
The 12 Questions, Grouped by What They Protect You From
Group 1: Protection from buying a demo
1. Can you run this on our data, on a case we pick? A demo built on the vendor's sample data proves only that the vendor can build a demo. Your data is messier: wrong formats, missing fields, three spellings of the same customer name. A real partner says yes, or asks for your ten worst records and a week. A deck-seller explains why that needs a paid discovery phase first.
2. How often does it get the answer right, measured how, and who checked? "Highly accurate" is not an answer. What you want has a number, a test set, and a judge: "on 200 real cases your team labeled, it matched your expert 91% of the time." If nobody on your side did the checking, the number is marketing.
3. What does it do when it does not know? This separates a real system from a rebranded chatbot faster than anything else. A real one detects the case it has not seen, stops, and hands it to a person. A rebranded one answers anyway, fluently, and is wrong.
Group 2: Protection from an undefined finish line
4. Write "done" in one sentence a non-technical person can check. Not "improved efficiency." Something like: "The system produces a first-draft proposal in under ten minutes that a salesperson sends with light edits, eight times out of ten." A vendor who cannot compress the outcome into one checkable sentence has not decided what they are building.
5. Who signs off that it is done, and by what date? One named person, one date. Acceptance criteria, meaning the written test that decides whether the work is finished, belong in the scope of work document, not in an email. Standard contract language for AI work puts the burden on the supplier: it runs the testing, provides evidence on request, and reworks anything that misses at its own cost, inside a fixed window (Tascon Legal). A project with no named acceptor never finishes. It runs out of budget.
6. What is the smallest version that pays for itself, and what would make you tell us to stop? A partner who cannot name a stopping condition is selling an open-ended engagement. Gartner projects more than 40% of AI agent projects will be canceled by the end of 2027, driven by rising costs, unclear business value, and weak risk controls (Gartner). None of those three is a technology problem. All three are scoping problems.
Group 3: Protection from owning nothing at the end
7. If we stop working with you on day 91, what exactly do we still have? The answer should be things you can point at: the code, running in your accounts, under your logins, documented well enough for your team to read. If the honest answer is "you lose access," you are renting an outcome, not installing one.
8. Who owns the code, the prompts, the outputs, and any model trained on our data? Ownership here is genuinely unsettled, which is why it has to be explicit. The US Copyright Office's January 2025 report holds that material generated entirely by AI is not copyrightable, that only human contributions in a mixed work can be protected, and that writing detailed prompts is not by itself enough (US Copyright Office). So "we own the output" may mean less than you think. What you can own outright is the code, the configuration, and the right to keep running it.
9. Where does our data go, who can see it, and does it improve anything you sell to others? The third part is the one vendors dodge. Ask whether your documents and your corrections feed anything offered to other customers, and get the answer in the contract rather than in an email.
Group 4: Protection from the day it breaks
10. When a wrong answer reaches a customer, who is responsible? Under the EU AI Act, the company using a high-risk system (the EU's label for AI applied to sensitive areas like hiring, credit, and education), not just the one that built it, must assign human oversight to named people with the competence, training, and authority to do it, and must keep the system's logs for at least six months (EU AI Act, Article 26). Even outside Europe that is the right shape. Name the human before you need one.
11. What breaks it, and how will we find out before a customer does? These systems degrade. Models change, connected software changes, your own data changes. NIST's AI Risk Management Framework treats this as a governance requirement, calling for policies covering AI risk from outside vendors and contingency processes to handle failures or incidents in third-party data or AI systems deemed high-risk (NIST AI RMF Playbook, GOVERN 6.1 and 6.2). Ask what the alarm is and who it wakes up.
12. What does year two cost, and what makes that number go up? The build price is the small number. Ask for the running cost, the maintenance cost, and the events that raise them: more volume, a new document type, a model price change. A partner who has done this before answers in thirty seconds.
Four Questions That Would Disqualify Us
A checklist that only flatters its author is a brochure. Here are four where the honest answer takes us out of the running.
"Can we call three named customers in our industry?" No. Our work sits under confidentiality and our case studies are anonymized. If named references are a hard requirement, we lose to a firm that can supply them, and that is a legitimate reason to pick someone else.
"Do you carry the enterprise security certifications our vendor review requires?" Probably not. We are a small senior team, not a certified enterprise supplier. If a certification is a hard gate in your review process, ask every vendor on day one and get a yes or no. A small senior team and a large certified vendor are different purchases, and pretending otherwise wastes a month.
"Can you staff round-the-clock support with a named on-call team and a one-hour response?" No. We build, install, hand over, and support on a defined arrangement. If you need a help desk running through the night, buy that from someone who runs one.
"Can you run a company-wide program across ten departments starting next month?" No, and we would not take it if we could. We do one bottleneck at a time, prove it pays, then move.
All four are questions about fit, not quality. The useful vendor conversation ends with both sides knowing which of you is wrong for this.
How to Spot a Deck-Seller in the First Meeting
- They ask your budget before they ask your bottleneck. A builder wants to know what breaks in your week. A deck-seller wants the number so the proposal can match it.
- The demo is a slide. If the screen shows an architecture diagram instead of the thing running, the thing is not running.
- They present a maturity model. Five-stage journeys with your company placed at stage two exist to sell stage three.
- Everything routes through a paid discovery phase. Fine when it produces something you would have bought anyway. A warning sign when the only deliverable is a proposal for more work.
- They say "platform" more than "output." Ask what you will be looking at on the Monday after it ships. If the answer is a dashboard rather than finished work, you are buying a tool and the adoption problem attached to it.
If you have not yet chosen between categories of vendor, done-for-you versus software versus consultants covers that. This post assumes the category is settled and you are scoring one firm.
What "Done" Has to Say in Writing
The written test of "done" usually fails for a boring reason: it is written in adjectives. Here is the shape that survives a dispute.
| Element | Weak version | Version that holds up |
|---|---|---|
| The output | "An AI proposal tool" | "A first-draft proposal in our template, populated from the deal record" |
| The standard | "High quality" | "Sent with light edits only, no rewriting of scope or pricing" |
| The sample | "It works" | "Across 40 real deals chosen by us, not the vendor" |
| The hit rate | "Accurate" | "32 of 40, with a miss defined before we start" |
| The judge | "The team" | "Our head of sales, named in the contract" |
| The date | "Q4" | "Day 30, with a go or no-go decision that day" |
| The miss | Unspecified | "Reworked at the vendor's cost, or the engagement ends" |
Every row on the right can be checked by an executive without asking an engineer. That is the only real test.
Scope Shape: The Smallest Thing That Pays, Then Stop
The best predictor of an AI project that ships is not the vendor. It is the size of the first bite.
Big programs fail because they need a lot of people to agree, and the Gartner conflict data tells you how likely that is across five to sixteen people. Small builds succeed because they need agreement on one thing and produce a result somebody can hold. Same reason a 30-day pilot beats a twelve-month program, and the same reason most department automation projects die of scope rather than technology.
Shape the contract like that: one bottleneck, one output, one judge, one date, one decision. Then stop and decide again. A partner who pushes back is telling you their business model needs a longer engagement than your problem does.
Red Flags That Are Fine, and Green Flags That Are Not
| Looks bad, usually fine | Looks good, usually bad |
|---|---|
| A small senior team with no famous logos | A big logo that turns out to be a one-week workshop |
| Refusing to quote a fixed price before seeing your data | A firm price on day one, before anyone has seen a record |
| Saying "we do not know yet" about a specific step | A 60-page proposal delivered in 48 hours |
| Insisting a person still checks the output | "Fully autonomous, no human needed" |
| Declining part of the scope | Accepting all of it without a single objection |
| A short contract with a stopping point | A three-year agreement with an annual uplift built in |
The pattern: confidence about process is a good sign, confidence about outcomes nobody has measured is not. Gartner is blunt about why. Most AI agent offers lack real value or return, because today's models are not mature enough to chase complex business goals on their own or follow detailed instructions over time (Gartner).
Frequently Asked Questions
Which of the twelve matters most?
Number three: what does it do when it does not know. Everything else is recoverable with a good contract. A system that answers confidently when it should stop produces wrong work at scale, and you hear about it from a customer rather than a log.
What if the vendor will not commit to an accuracy number?
Ask them to commit to a method instead. "We test on 40 cases you choose, you define what counts as a miss, and we agree the bar before we start" beats a confident percentage invented on the spot. Refusing both is the disqualifier.
How long should the evaluation take?
Weeks, not quarters. The twelve questions run in two meetings. If your evaluation is taking six months, the delay is almost never the vendors. It is the conflict inside your own group, and more vendor research will not fix it.
If you are running this checklist against us, start with question four and question six. We will write "done" in one sentence, name the smallest version that pays for itself, and tell you plainly if your problem is the wrong shape for us. See what we install inside companies →
Arrête de tout configurer. Place à la construction.
Des templates SaaS avec orchestration IA.