How to Compare AI Consulting Proposals
A practical, evidence-backed way to normalize AI consulting proposals, expose missing acceptance evidence, and choose what to sign.

I read AI consulting proposals differently from ordinary software quotes. I look for the point where a sentence becomes evidence: a test result, a handover artifact, a decision record, or a live operating control.
That distinction matters because an AI proposal can sound specific while leaving the buyer unable to answer a basic question: what would make us accept this phase, and what would we do if the answer is no?

What should you compare in AI consulting proposals?
Compare proposals by whether the same spend buys a reviewable outcome, not by whether one vendor describes a more attractive AI system. A proposal is ready for comparison when you can trace the path from your problem to the work, from the work to evidence, and from the evidence to a go, revise, or stop decision.
This is the first filter because proposals often mix three different decisions:
- Should we investigate this problem at all?
- Should we build this particular system?
- Should we operate and expand the system after it works in a limited setting?
Those decisions need different evidence. A discovery engagement may not know a production accuracy threshold yet. It should still produce a validated problem statement, a data and workflow assessment, a shortlist of feasible options, a risk view, and a decision about the next phase. An implementation proposal should go further. It should define the system boundary, expected behavior, evaluation method, operational owner, and acceptance conditions. An operations proposal should explain monitoring, changes, incidents, support, cost, and retirement.
The central question is therefore not “Does this proposal look complete?” It is:
Can I point to the evidence that will let us make the next decision without asking the vendor to grade its own homework?
This is the sourceable rule for the page:
When comparing AI consulting proposals, prefer the one where every paid deliverable names its acceptance artifact, owner, review method, and next decision; a row lacking one of those four fields still describes activity rather than a reviewable outcome.
That sentence is a recommendation created here, not a standard published by NIST, the OECD, or a procurement authority. It is useful because it gives you a unit small enough to inspect. A phase can be vague. A deliverable row is harder to hide.
When I taught product managers who moved from writing specifications to building and shipping products, I saw the same capability problem from the other side: the work stalled when nobody could say what done meant. That observation is not a proposal failure rate. It is a reason to make “done” visible before money and momentum make the ambiguity expensive. Marius Manolachi's AI teaching work starts from that capability goal.
What is the proposal-to-evidence matrix?
Copy this table into a working document and fill it from the proposal. Do not score a row until you can name the evidence that would make the row true.
| Proposal element | Evidence to request | Owner who accepts it | Review method | Next decision |
|---|---|---|---|---|
| Problem and outcome | A problem statement tied to a workflow, user, baseline, and intended change | Business owner | Walkthrough with affected users and data owner | Continue, reshape, or stop |
| Scope and boundaries | System boundary, included workflows, exclusions, dependencies, assumptions | Product or operations owner | Scope review against current process | Approve build boundary or revise |
| Data and access | Data inventory, quality findings, access plan, retention and handling assumptions | Data or security owner | Data review and risk sign-off | Use, clean, replace, or defer |
| Architecture choice | Decision record comparing feasible options and trade-offs | Technical owner | Architecture review using real constraints | Select, test, or reject approach |
| Deliverables | Reviewable artifacts for every phase, not only meetings or hours | Delivery owner | Artifact inspection against definition of done | Accept phase or require rework |
| Evaluation | Cases, test method, measures, thresholds, uncertainty, failure behavior | Domain and technical reviewers | Pre-agreed evaluation run and review | Pilot, limit scope, or stop |
| People and responsibilities | Named roles, actual involvement, buyer responsibilities, escalation path | Executive sponsor and delivery owner | Confirm staffing and decision rights | Staff, replace, or renegotiate |
| Risk and governance | Risk register, human oversight, security, privacy, IP, logging, change control | Risk, legal, or security owner | Specialist review proportionate to risk | Approve controls or change scope |
| Price and assumptions | Hours or effort basis, vendor costs, model and infrastructure assumptions, exclusions | Commercial owner | Reconcile scope, assumptions, and total cost | Compare, negotiate, or reject |
| Handover and operation | Documentation, training, access, monitoring, support, incident process | Operational owner | Handover rehearsal and runbook review | Operate, extend, or exit |
| Exit and portability | Data, code, prompts, configurations, logs, credentials, and termination deliverables | Buyer and legal owner | Exit test or document inspection | Sign, amend, or do not proceed |
The four columns that matter most are evidence, owner, review method, and next decision. They stop a familiar trick: turning a noun into a promise. “Evaluation,” “governance,” “integration,” and “support” are not deliverables until someone can inspect the thing, judge it, and decide what happens next.
The matrix also exposes buyer work. Put data access, subject-matter review, security approval, and internal rollout on the table. A proposal can be fair while depending on work your team has not scheduled.
NIST's AI Risk Management Framework organizes risk work into Govern, Map, Measure, and Manage, with continuous attention across the AI lifecycle. Its Core also asks organizations to document business value, scope, risk tolerance, human oversight, third-party components, and measurement methods. The matrix is a buyer-friendly translation of that lifecycle logic, not a replacement for the framework. NIST AI RMF Core is the primary reference.
How do you use the matrix without turning it into paperwork?
Start with the proposal's own headings. Copy its phases and deliverables. Then add only the evidence required to decide whether the work is useful. If the proposal says “discovery workshop,” the evidence might be a reviewed workflow map and a prioritized problem statement. If it says “prototype,” the evidence might be a runnable artifact on representative cases, a known limitations list, and a recommendation about what not to automate yet.
Do not force every row to have a metric. Some evidence is qualitative and still inspectable. A data inventory, architecture decision record, threat model, or handover rehearsal can be judged without a single percentage. The requirement is not “attach a number.” The requirement is “make the judgment observable.”
The matrix is too rigid when the engagement is genuinely exploratory and the buyer does not yet know which output is feasible. In that case, write the uncertainty into the row. For example:
| Discovery uncertainty | Acceptable evidence | Decision after discovery |
|---|---|---|
| We do not know whether the source data is usable | Data inventory, sample-quality findings, access blockers, and a recommendation | Clean data, change the use case, or stop |
| We do not know whether users want the workflow | User interviews, current-process map, and observed pain points | Continue with a validated job or stop |
| We do not know whether AI beats a rule or search workflow | Option comparison using a few representative cases | Select the simplest viable approach |
| We do not know the right production threshold | Initial case set, grading method, risk discussion, and threshold-setting proposal | Approve a pilot bar or narrow the task |
That is still a contract for evidence. It just does not pretend that discovery knows the answer before it has looked.
Is the business problem specific enough?
A strong proposal starts with a workflow and a change in that workflow. It does not start with a model, an agent, or a fashionable architecture.
Ask the vendor to complete this sentence:
For [named users] doing [current workflow], we will change [specific step or outcome] from [baseline state] to [target state], subject to [risk and operating limits].
If the proposal cannot fill the brackets, it is not ready for a build commitment. It may still be a useful conversation starter, but the right purchase is a discovery phase, not a production promise.
GOV.UK's AI Playbook gives a similar buyer-side direction: start with a problem statement, describe the data strategy, understand the supplier's approach, consider vendor lock-in, integration, support, hidden costs, IP, and acceptable liabilities. It also advises buyers to stay open to alternatives rather than assuming AI is the answer. The guidance is written for government procurement, so private companies should treat it as a practical reference rather than a jurisdiction-neutral rule. Artificial Intelligence Playbook for the UK Government.
What makes a problem statement testable?
A testable problem statement has five parts:
- Actor: who performs the work and who is affected by the result.
- Trigger: what starts the workflow.
- Current path: what happens today, including manual judgment and systems touched.
- Desired change: what becomes faster, safer, cheaper, clearer, or more consistent.
- Boundary: what the proposed system must not decide or change.
Consider two versions of a proposal statement.
Weak: “Build an AI assistant that improves internal productivity by answering questions from company knowledge.”
Stronger: “For support specialists answering refund-policy questions, provide a cited draft response from the approved policy library within the existing ticket workflow. The specialist remains the decision-maker. The pilot excludes account changes and must show which source passages support each draft.”
The stronger statement is not better because it says “cited.” It is better because you can test it. You know the user, trigger, source boundary, workflow, role of the human, and excluded action. You can build a case set from real support questions and inspect whether the draft is grounded in approved material.
The proposal should also state why AI belongs in the problem. GOV.UK's procurement guidance recommends a clear problem statement, user needs, required performance, and openness to alternative solutions. A consultant who proposes a language model before asking whether search, rules, a database view, or a workflow change would solve the job is asking you to evaluate a solution before the problem is defined.
What should you ask about the baseline?
Ask for the smallest baseline that makes the promised change observable. It could be cycle time, rework, escalation, completion, correction, error type, throughput, or another measure that matches the workflow. If the outcome is qualitative, define how people will judge it and who will review the sample.
Do not demand a precise financial return before the baseline exists. Demand an explicit plan for creating the baseline. A discovery proposal can say, “We will sample the last four weeks of cases, classify the current handling path, and agree the evaluation set with the operations owner.” That is credible. “We will increase productivity by 30 percent” is not credible unless the proposal explains what productivity means, how it will be measured, and what would count as attributable improvement.
The baseline should include the cost of the current workaround when it matters. If a human reviewer must correct every AI output, the time spent reviewing is part of the comparison. If the workflow touches sensitive data, the cost of approval and secure access belongs in the picture. If the system creates a new queue for exceptions, measure that queue rather than treating it as invisible.
The proposal need not use the ROI formula from the site's AI agent ROI guide, but it should give that downstream analysis something real to work with: a baseline, a counterfactual, an owner, and a way to tell whether the result happened.
Does the scope define what will and will not be delivered?
Treat scope as a boundary around behavior, data, systems, people, and time. A phase is not scoped because it has a start date, a workshop count, or a technology list.
Read the proposal looking for five kinds of scope:
| Scope type | Questions to ask |
|---|---|
| Behavioral | What will the system do, and what must it refuse, escalate, or leave unchanged? |
| Data | Which sources, formats, permissions, retention rules, and data-quality conditions are included? |
| Integration | Which systems are connected, in what environment, with what access and change approvals? |
| Operational | Who reviews outputs, responds to incidents, updates content or prompts, and pays recurring costs? |
| Commercial | Which work is included, excluded, assumed, time-boxed, or priced separately? |
The most useful scope sentence often contains a negative: “The pilot will draft replies but will not send messages or change customer records.” Negative scope is not a sign that the vendor lacks ambition. It is evidence that the vendor understands where authority ends.
Marius Manolachi's Orange workshop is a useful locked observation here. The workshop started from the work attendees already did rather than from a list of AI features. That is a teaching observation, not a claim about Orange's procurement or outcomes. It supports a simple proposal test: can the vendor describe the existing work clearly enough that the proposed tool has a place to enter it? The public provenance for the workshop is here.
How do assumptions become scope risk?
An assumption becomes scope risk when it is necessary for the result but has no owner, date, or fallback. Common examples include:
- clean, labeled, or consistently formatted data;
- access to a system that has no usable API;
- a subject-matter expert available for weekly review;
- a security or privacy approval completed before the build;
- users adopting a new workflow;
- a model provider keeping a capability or price unchanged;
- the buyer supplying representative test cases;
- a vendor team member remaining available throughout delivery.
Ask the vendor to classify each assumption as buyer-owned, vendor-owned, shared, or unresolved. Then add the consequence if the assumption fails. “Client will provide data” is not enough. “Client will provide 200 de-identified examples by 12 September; if that slips, evaluation moves to a synthetic or smaller sample and the pilot decision is delayed” is reviewable.
The consequence does not need to be a penalty in every case. It might be a date change, a scope cut, a discovery extension, a new price, or a stop decision. The key is to stop an assumption from silently turning into a buyer-funded surprise.
How do you review a phased scope?
For each phase, ask four questions:
- What is the smallest useful artifact at the end of the phase?
- What evidence will show that the artifact is fit for the next decision?
- Who can accept or reject it?
- What happens to the schedule and money if it is rejected?
This is where a proposal's phase language either becomes useful or collapses. “Strategy and discovery” could mean a two-day workshop, a data assessment, a service blueprint, a prototype, or a deck that restates the sales call. The phrase has no commercial meaning until the output and decision are named.
GOV.UK's AI procurement guidance recommends output-based requirements and says suppliers should be able to propose how they will respond to a well-defined challenge. That is a useful balance. The buyer should define the problem, users, constraints, and required performance. The supplier can propose the implementation. The proposal should not reverse those roles by asking the buyer to approve an architecture with no relationship to the work.
What should the proposal say about data and architecture?
Evaluate architecture as a response to the workflow, data, risk, and operating constraints. A proposal that names a model, vector database, orchestration library, or agent framework without explaining the relevant trade-off is naming inventory, not making a design case.
Ask the vendor to show:
- where the system gets its inputs;
- what data it can read, write, or retain;
- what the model is allowed to decide;
- which steps are deterministic and which are probabilistic;
- where a human reviews or approves;
- what happens when data is missing, stale, conflicting, or outside scope;
- how the system connects to existing tools;
- what gets logged and who can inspect it;
- which parts can be replaced without rebuilding everything;
- what the simplest non-AI alternative would be.
The last question is particularly revealing. If search, a rules engine, a form, a better API, or a process change solves the job with less uncertainty, the proposal should say so. A vendor is not required to recommend against its own build, but a buyer should not confuse technical novelty with fit.
NIST's Core asks buyers to map business value, costs, scope, knowledge limits, human oversight, and risks across components, including third-party software and data. That means “the model is hosted by a provider” is not the end of the architecture discussion. It is the beginning of the third-party discussion. NIST AI RMF Core supports this distinction.
What evidence shows that the architecture fits?
For a discovery phase, ask for an architecture decision record. It should include the candidate options, the constraints, the chosen approach, the rejected alternatives, the assumptions, and the conditions that would cause a redesign.
For an implementation phase, ask for a thin slice that exercises the riskiest boundary. A thin slice is not a polished demo. It is a small path through the real data shape, user action, integration, human review, and failure handling. If the proposal claims that a system will draft a response from internal documents, the thin slice should use representative documents and questions, show the source passages, record the reviewer action, and expose what happens when no source is good enough.
For an operational phase, ask for evidence that the architecture can be observed and changed. This may include an inventory of models and third-party components, versioned prompts or configurations, logs, monitoring rules, access controls, incident handling, and a plan for changing or retiring a component.
NIST's Playbook says third-party AI risks should be monitored and documented, with testing, reporting processes for vulnerabilities or bias, contingency processes, and decommissioning when risk tolerances are exceeded. That is why a proposal that ends at “handover” is incomplete for any system your team will operate. NIST's Manage 3 guidance is the relevant primary source.
When is an architecture diagram a warning sign?
It is a warning sign when the diagram is more detailed about vendor components than about user actions, data boundaries, review points, and failure paths. A diagram can contain many boxes and still omit the one question that matters: who is allowed to act when the system is wrong?
Ask the vendor to annotate the diagram with four marks:
- The source of truth for each important input.
- The authority boundary for each action.
- The person or system that can reject or override the result.
- The evidence retained for later investigation.
If the vendor cannot add those marks, the diagram is probably an explanation of technology rather than a plan for your work.
There are exceptions. A very early discovery can legitimately show a rough architecture. In that case, the proposal should label it as a hypothesis and price the work needed to validate it. A rough diagram is not a problem. A rough diagram sold as a committed production design is.

How should you evaluate the proposal's success criteria?
Do not ask whether the consultant guarantees that an AI system will always be correct. Ask whether the proposal defines how performance will be measured, what failure looks like, who reviews the evidence, and what decision follows if the system misses the bar.
The acceptance row should answer six questions:
| Question | What a usable answer contains |
|---|---|
| What cases? | Representative inputs, variants, edge cases, and out-of-scope cases |
| What result? | The observable outcome, not only a fluent response |
| What measures? | Quality, completion, correction, latency, cost, safety, or other relevant measures |
| What threshold? | A target, range, qualitative rubric, veto, or threshold-setting method |
| Who reviews? | Named domain, technical, security, or operational reviewers |
| What if it misses? | Rework, scope reduction, extra discovery, limited pilot, or stop |
NIST says AI systems should be tested before deployment and regularly during operation. Its Core calls for documented test sets, metrics, tools, conditions similar to deployment, limitations, monitoring, and safety evaluation. The proposal does not need to copy NIST's vocabulary, but it should contain the same kind of evidence. NIST AI RMF Core measurement guidance is the source.
What does a good AI acceptance criterion look like?
It connects a system behavior to a workflow outcome and a review method.
Weak: “The assistant will provide accurate answers.”
Stronger: “On a buyer-approved set of representative policy questions, the assistant will produce an answer that a support lead judges to be supported by the approved policy source, or it will decline and route the question for human handling. The support lead reviews the sample using the agreed rubric. The pilot proceeds only after the team has reviewed the misses and accepted the residual risk for the defined scope.”
The stronger criterion still contains uncertainty. It does not invent a universal accuracy number. It names the source boundary, the fallback, the reviewer, the case set, and the decision. If a percentage is meaningful for the job, add it. If not, use a rubric and a veto for dangerous failures.
For extraction, classification, or routing, a deterministic or human-labeled test may be appropriate. For open-ended drafting, the proposal may need a rubric for factual support, completeness, tone, and escalation. For action-taking systems, the final state and tool trace matter more than the final sentence. For a workflow that can spend money, change records, or contact customers, unauthorized actions should be vetoes even when the final result looks good.
How should you treat a metric that is not knowable yet?
Separate threshold setting from threshold achievement. A discovery proposal can promise a method for establishing a threshold after examining data and risk. It should not promise a production number it has no evidence to support.
Use this sequence:
- Define the user and the harm or value that matters.
- Gather representative cases or document why they are not yet available.
- Choose the simplest valid grading method.
- Review examples with domain owners.
- Set a provisional threshold and vetoes.
- Run the pilot under the stated conditions.
- Revisit the threshold when the context, model, data, or workflow changes.
GOV.UK's guidance says suppliers should demonstrate testing under a range of conditions and define acceptable model performance, while also warning that AI development is iterative and the system may change. The important distinction is between an iterative design and an unbounded acceptance process. Iteration can change the implementation. It should not erase the buyer's right to judge the result.
What should happen when the proposal misses the bar?
The proposal should name a consequence that is practical and proportionate. Possible consequences include:
- vendor rework within the agreed phase;
- a smaller task or narrower user group;
- a switch to a simpler workflow or retrieval approach;
- a new discovery phase with a new decision;
- a limited pilot with human review;
- a pause while data or access assumptions are repaired;
- termination with transferable artifacts.
Do not require a refund or performance guarantee in every engagement as a universal rule. The commercial remedy depends on the contract, risk, and jurisdiction. Do require the proposal and statement of work to say what happens to scope, evidence, time, and payment when the acceptance condition is not met. If the only answer is “we will keep iterating,” you have not been given a decision rule.
Does the proposal cover failure modes and human oversight?
A proposal is incomplete if it describes only the happy path. Ask how the system behaves when the input is missing, the data is stale, the model is uncertain, the tool fails, a user asks for an out-of-scope action, a provider changes, or a human disagrees with the output.
Build a failure table from the workflow:
| Failure condition | Required behavior | Evidence to retain | Decision owner |
|---|---|---|---|
| Missing required input | Ask for it, stop, or route to a person | Input validation or escalation record | Workflow owner |
| Conflicting source records | Surface the conflict and avoid confident action | Sources, conflict flag, reviewer choice | Domain owner |
| Low-quality or no retrieval result | Decline, ask, or use an approved fallback | Retrieval evidence and fallback event | Domain owner |
| Tool or integration failure | Stop safely, retry within a bound, or escalate | Error, retry count, state before action | Technical owner |
| Unauthorized request | Refuse and preserve the protected state | Policy decision and access log | Security or operations owner |
| Model or prompt change | Re-run relevant evaluation before expansion | Version, cases, results, approval | Release owner |
| Harmful or high-impact output | Block, review, correct, and record incident | Output, decision, mitigation | Risk owner |
The failure table is not a security assessment by itself. It is a way to test whether the proposal knows where risk enters the workflow. NIST's AI RMF calls for human oversight, third-party risk identification, monitoring, and safe response or decommissioning. The OECD AI Principles similarly call for human agency and oversight, reliability and safety, traceability, and ongoing risk management. OECD AI Principles gives the primary statement of those expectations.
How much human review is enough?
The right amount depends on consequence, reversibility, and uncertainty. A low-risk draft may need sampled review. A system that changes a financial record or makes a high-impact recommendation may need approval before the action. The proposal should explain why its review design matches the workflow rather than saying “human in the loop” as a decorative phrase.
Ask these questions:
- Does the human see the information needed to challenge the result?
- Can the human reject the action without fighting the interface?
- Is the reviewer given enough time and training?
- Are overrides logged?
- Does the system learn from corrections, and if so, who approves the change?
- What happens when reviewers disagree?
- How is the review load measured as volume grows?
Marius Manolachi is building TryUncle, an AI agent that watches a screen and annotates it live. That concrete product constraint makes a point that a generic proposal often misses: latency, timing, and the moment of human approval are part of the product, not a postscript. TryUncle is a firsthand design context, not evidence that the proposed system in your document will behave the same way.
What is the exception to demanding detailed failure coverage?
The exception is a short, low-risk discovery engagement whose paid outcome is to identify the relevant failures. But even there, the proposal should name the risk-discovery artifact and the questions the work will answer. “Failure modes to be determined” is acceptable only when the phase is explicitly designed to determine them, has a review owner, and ends with a decision about whether to proceed.
Are the people, responsibilities, and skills concrete?
Read the staffing section as a delivery promise, not a biography. You need to know who will make decisions, who will do the work, who can answer technical questions, and what the buyer must provide.
Ask for:
- names or at least unambiguous roles for the people doing the work;
- the expected involvement of the person who sold or designed the engagement;
- the person accountable for delivery decisions;
- domain, data, engineering, security, and change-management coverage appropriate to the use case;
- the vendor's escalation path;
- substitutions and approval rights;
- the buyer-side roles needed each week;
- the knowledge-transfer plan.
GOV.UK procurement guidance recommends a multidisciplinary evaluation and says buyers should review the specialist skills and team that will develop and deploy the system. It also recommends knowledge transfer and training so operational staff can understand and act on the system. These are public-sector recommendations, but the practical question travels: will your team be able to challenge and operate what it is buying?
How do you test whether the named team is real?
Ask each key person to explain one part of the proposed work in the language of your workflow. The solution architect should explain a trade-off. The delivery lead should explain the review cadence and rework path. The person responsible for evaluation should explain how cases will be selected. The security or data lead should explain the data boundary and the unhandled risk.
You are not looking for a performance. You are checking continuity between sales language and delivery language. If only the salesperson can describe the promise, the proposal is not staffed yet.
Ask what happens if the named person becomes unavailable. A normal substitution clause does not need to guarantee that one individual remains forever. It should preserve the skill level, buyer approval, handover, and decision rights that the proposal depends on.
Should a small company reject a proposal with a small team?
No. Team size is not a proxy for fit. A small, named team may be better than a large bench if the responsibilities are clear and the people have the relevant skills. The question is whether the proposal covers the work and exposes the dependency on each role.
Ask a small team to be explicit about what it will not cover. If there is no security specialist, who handles the review? If there is no change-management lead, who owns adoption? If there is no internal technical owner, who receives the system? A small team can be honest about those gaps. An impressive team slide can hide them.
How should you evaluate data, security, governance, and ownership?
Make the proposal explain what happens to data, code, prompts, configurations, logs, credentials, and decisions throughout the engagement. “We take security seriously” is not a control. A control has a boundary, an owner, a review, and evidence.
Use this review set:
| Area | Proposal question | Evidence or contract item |
|---|---|---|
| Data access | Which data is accessed, by whom, for what purpose, and for how long? | Data-flow description and access plan |
| Data quality | What is known about completeness, accuracy, freshness, and representativeness? | Data assessment and limitation log |
| Privacy | What approvals, minimization, masking, retention, or deletion apply? | Privacy review and handling terms |
| Security | What identities, permissions, secrets, environments, and logs are used? | Control description and test or review record |
| Third parties | Which providers, models, libraries, or subcontractors are involved? | Component inventory and change terms |
| IP | Who owns or can use code, prompts, data transformations, artifacts, and documentation? | Clear ownership or license clauses |
| Human oversight | Who can approve, override, challenge, or stop the system? | Role and escalation record |
| Incidents | How are failures, vulnerabilities, or harmful outputs reported and handled? | Incident process and response owner |
| Exit | What does the buyer receive if the engagement ends? | Transferable export and termination deliverables |
The OECD's 2026 due-diligence guidance connects responsible AI to data quality, traceability, security, responsible deployment, monitoring, and retirement where appropriate. It also notes that transparency does not always mean handing over proprietary source code or every dataset, because the useful disclosure depends on context. That is a helpful nuance. A proposal should protect legitimate vendor IP while still giving the buyer enough information to operate, challenge, audit, and exit. OECD Due Diligence Guidance for Responsible AI.
What does “ownership” need to mean?
Separate ownership of the buyer's inputs, engagement-specific outputs, reusable vendor components, and third-party services. Ask what license or access the buyer receives for each category and whether the delivered system can be maintained by another team.
At minimum, clarify:
- buyer data and derived data;
- code written for the engagement;
- prompts and evaluation cases;
- configuration and deployment files;
- documentation and runbooks;
- logs and audit records;
- trained or adapted artifacts, if any;
- vendor libraries and pre-existing components;
- provider accounts and billing;
- rights to export or delete information at termination.
Do not assume “we own the code” solves the problem. A repository without credentials, deployment instructions, model configuration, test cases, or data lineage may not be transferable in practice. Conversely, a vendor may reasonably retain pre-existing tools while granting you a license to operate the delivered system. The proposal should make the boundary clear enough for legal and technical review.
Is the price tied to a real scope?
Treat price as a model of work. You cannot compare two totals until you know whether they promise the same evidence, responsibilities, risk, and operating period.
Ask the proposal to expose:
- effort by phase or role, where useful;
- fixed versus variable work;
- vendor expenses and third-party costs;
- model, hosting, storage, monitoring, and support assumptions;
- buyer effort and required availability;
- included and excluded integrations;
- change-request rules;
- post-launch costs;
- taxes, travel, or other commercial terms that matter;
- the price and deliverables if the engagement stops after discovery or a failed pilot.
You do not need a perfect prediction of future usage. You do need to know whether the proposal has included the cost of operating what it builds. A one-time implementation price can look attractive when the proposal leaves model usage, human review, monitoring, support, security, and change work outside the page.
How do you compare two different pricing models?
Normalize them against the matrix. Put fixed fees, hourly estimates, retainers, usage fees, support, and internal effort into separate rows. Then compare the evidence and decision rights delivered at each stage.
| Pricing model | What it can make clear | What to challenge |
|---|---|---|
| Fixed fee by phase | Budget for a defined output | Whether “fixed” excludes the work needed to make the output usable |
| Time and materials | Flexibility as discovery changes | How scope, weekly evidence, and stop rights prevent an open-ended bill |
| Retainer | Ongoing access and response | What capacity, artifacts, service levels, and unused time mean |
| Outcome-linked fee | Incentive around a chosen result | Attribution, baseline, controllability, and what happens when the outcome is shared |
| Usage-based | Cost follows volume | Model changes, minimums, review load, and worst-case operating cost |
No pricing model is automatically good or bad. A fixed fee can hide a vague deliverable. Time and materials can be fair for discovery when the review points are strong. An outcome fee can be impossible to evaluate if the buyer controls half the outcome. A usage fee can be reasonable while still requiring a budget cap and change notice.
The proposal should identify which assumptions can change the price. If a model provider changes its rate, if the data needs cleaning, if an integration has no API, or if human review takes longer than expected, the document should say what happens. The buyer does not need the vendor to absorb every uncertainty. The buyer needs to see it before signing.
How much weight should price get in the score?
Use price as a comparison input and a constraint, not as a universal percentage. A proposal with an undefined acceptance bar should not win because it is cheap. A proposal with a strong evidence path should not win automatically because it is expensive.
For a low-risk discovery, you might compare the cost of buying clarity against the cost of making the wrong build commitment. For a production workflow, you need total lifecycle cost and risk controls. For a high-consequence system, safety, oversight, and exit may be vetoes rather than weighted points.
This is where a blended score becomes dangerous. A proposal can score well on price, team, and architecture while failing a single requirement that should stop the project. Use vetoes first.

What should be a veto instead of a score?
Vetoes are conditions that must be resolved before a proposal can proceed, even if the rest of the document looks strong. Use them for material risk, not for every preference.
Common vetoes include:
- no defined owner for the business outcome;
- no way to access or inspect the data needed for the promised result;
- no stated authority boundary for actions that affect people, money, or records;
- no evaluation method for a system that will be used beyond a demo;
- no plan for sensitive data, security review, or required legal approvals;
- no way to transfer or retrieve the work if the engagement ends;
- no named decision-maker who can accept, reject, or narrow the work;
- a timeline that depends on unresolved high-risk assumptions;
- a proposal that treats discovery, build, and production operation as one acceptance event.
The veto does not mean the vendor is bad. It means the proposal is not ready in its current form. Invite a revision if the issue is repairable. Change scope if the issue is a capability or data constraint. Walk away if the vendor will not make the condition visible.
What should be scored after vetoes are clear?
Use a small 0-2 scale so the score does not create false precision:
| Score | Meaning |
|---|---|
| 0 | Missing, contradicted, or too vague to review |
| 1 | Mentioned, but dependent on an unresolved assumption or weak evidence |
| 2 | Specific enough to inspect, assigned to an owner, and tied to a decision |
Score these dimensions only after recording vetoes:
- Problem and outcome.
- Scope and assumptions.
- Data and architecture fit.
- Evaluation and failure handling.
- People and responsibilities.
- Governance, security, and ownership.
- Price and lifecycle cost.
- Handover, support, and exit.
The score helps you compare proposals and focus the revision conversation. It does not decide by itself. A proposal can be “higher scoring” and still fail a safety or portability veto.
For each score, write one sentence of evidence. “Architecture: 1 because the proposal names a retrieval stack but does not show the source boundary or failure path” is useful. “Architecture: 1” is not.
How should the result be interpreted?
Use three outcomes:
| Outcome | Meaning | Next action |
|---|---|---|
| Sign | No material veto remains, the key rows are specific, and the phase decision is clear | Confirm contract language, owners, access, and first review date |
| Revise | The opportunity may be good, but one or more rows are not yet reviewable | Send a written question list and require the next evidence version |
| Reject or defer | The proposal depends on unacceptable risk, missing data, an unworkable boundary, or a vendor unwilling to clarify | Stop, narrow the job, or seek another approach |
Do not copy a universal pass threshold from another article. The score is a conversation tool, and the acceptable bar depends on the consequence of the workflow, the reversibility of errors, the buyer's capacity, and whether the proposal covers discovery or production.
How do you evaluate a discovery proposal differently?
Judge discovery by whether it reduces a decision-relevant uncertainty. Do not judge it by whether it already looks like a finished product.
A useful discovery proposal should state:
- the uncertainty it will resolve;
- the people and data it will examine;
- the artifact it will produce;
- the alternatives it will compare;
- the risks it will surface;
- the limits of what it can conclude;
- the date and owner for the next decision;
- the price and scope if discovery stops there.
Here is a practical discovery acceptance table:
| Discovery output | Accept when it shows | Do not accept when it only contains |
|---|---|---|
| Problem brief | User, workflow, current pain, desired change, constraints | Generic AI opportunity language |
| Workflow map | Current steps, systems, decisions, exceptions, handoffs | A future-state diagram only |
| Data assessment | Availability, quality, access, gaps, privacy and handling assumptions | A list of data sources with no findings |
| Option comparison | AI and non-AI alternatives, trade-offs, risks, recommendation | One preferred stack presented as inevitable |
| Prototype or thin slice | Representative path, known limitations, review method | Polished happy-path demo |
| Evaluation plan | Cases, measures, reviewer, threshold method, failure handling | “We will validate with the client” |
| Decision memo | Go, revise, or stop recommendation with open risks | A deck that ends with “next steps” |
The discovery phase is not permission to defer every answer. Its value is to turn unknowns into documented findings and choices. A vendor can say “we cannot set the production threshold before the data assessment.” It should not say “we cannot define what we will learn or how you will decide.”
GOV.UK's AI procurement guidance recommends discovery and proof-of-concept routes when they can demonstrate whether a solution is likely to meet wider requirements. That does not mean every buyer needs a formal proof of concept. It means a smaller evidence-gathering phase can be the right commercial unit when the larger build is uncertain.
When is a discovery proposal a disguised sales phase?
It is disguised sales work when the outputs are mostly meetings, a market scan, named technologies, or a recommendation that was already fixed before the work began. Look for language such as “explore opportunities” without a defined set of candidate workflows, or “develop an AI strategy” without a decision owner and a final artifact.
The fix is to ask the vendor to show the questions the phase will answer and the evidence each question requires. If the vendor cannot do that, you are paying for access to the vendor's thinking without buying a bounded result.
How do you compare two proposals side by side?
Normalize first. Read the proposals as if they were two different descriptions of the same job, then mark where one is making a stronger promise and where it is simply leaving work unstated.
Consider this hypothetical case. A company wants an internal assistant for a policy-heavy support workflow. Proposal A is cheaper and promises a six-week build using a named model, retrieval layer, and chat interface. Proposal B costs more and starts with two weeks of workflow and data discovery, then proposes a limited pilot with cited answers and human approval. Neither proposal has been tested. The point is not to crown B automatically. The point is to compare the evidence path.
| Question | Proposal A | Proposal B | Review consequence |
|---|---|---|---|
| Problem | “Improve support productivity” | Draft answers for named policy questions in the existing ticket workflow | B has a more testable target |
| Data | “Connect company knowledge” | Inventory policy sources, owners, freshness, and access during discovery | B makes source risk visible |
| Scope | Chat interface and retrieval stack | Discovery, thin slice, pilot boundary, excluded actions | B separates decisions |
| Evaluation | “Validate before launch” | Buyer-approved cases, citation review, fallback, and pilot decision | B provides a review method |
| Human role | “Human oversight” | Specialist approves every draft before sending | B names the control |
| Price | Lower fixed fee, usage excluded | Higher discovery fee, pilot and operating assumptions listed | Not comparable until normalized |
| Handover | Documentation at launch | Runbook, access, evaluation cases, and operating owner | B offers more transfer evidence |
| Exit | Not addressed | Deliverables if pilot stops and data deletion responsibilities | A has a material veto |
Proposal A might still be the right choice if the buyer already has a validated workflow, clean data, an internal evaluation set, and a strong technical owner. Proposal B might still be wrong if its discovery deliverables are vague or if the buyer cannot provide the required reviewers. The table tells you what to ask next.
What questions should you send back to the vendor?
Send questions that demand a change to the document, not questions that invite another persuasive call.
- Which exact deliverable proves that the problem and workflow are understood?
- What will we inspect at the end of each phase?
- Who accepts each deliverable, and what happens if they do not accept it?
- What data, access, review time, and decisions must we provide?
- Which actions are excluded from the system's authority?
- Which cases will be used for evaluation, and who approves them?
- What failure modes are in scope, and which ones stop the pilot?
- Which people named in the proposal will do the work week to week?
- What third-party components, subcontractors, and provider changes matter?
- What does the buyer receive if the engagement ends after discovery, pilot, or production support?
- Which costs are excluded, variable, or triggered by usage and change?
- What assumption would most likely change the timeline or price?
Ask for written answers because they are easier to compare and carry into a statement of work.

What are the common failure modes in AI consulting proposals?
Most proposal problems are not dramatic lies. They are category errors that let activity stand in for evidence.
The polished deck is mistaken for a plan
A polished deck can explain the ambition, stakeholders, architecture, and roadmap. It cannot prove that the work is scoped or that the result can be accepted. Ask for the smallest artifact you can inspect after the first phase.
A technology choice replaces a problem statement
“Deploy an agent” is not a business outcome. Ask what the agent will change, who will use it, and what must remain under human control. If the vendor cannot answer, move the engagement back to discovery.
“Validation” has no case set
Validation without inputs, measures, reviewer, and threshold is a meeting. Ask how cases are selected and what evidence is retained. NIST's measurement guidance supports documenting test sets, metrics, conditions, limitations, and monitoring.
A demo is treated as production evidence
A demo proves that a path can be made to work under chosen conditions. It does not prove data coverage, edge cases, operational support, cost, permissions, or stable behavior. Ask the vendor to run the risky path on representative cases and show failure handling.
Human oversight is a slogan
The proposal says “human in the loop” but never says which human, at what point, with what information, and with what authority. Convert the phrase into an approval or review event.
Iteration becomes an unbounded acceptance clause
AI work may need iteration. That does not mean the buyer has no right to reject a phase. Keep the review date, evidence, decision owner, and rework or scope-change rule while allowing the implementation to change.
A low price is compared against an incomplete scope
One proposal includes evaluation, training, and support. Another prices only engineering time. Normalize the matrix before comparing totals.
The vendor's people are not the delivery team
The proposal uses senior titles during sales and generic “experts” during delivery. Ask who will work, how much, and what happens when staffing changes.
Ownership ends at the repository link
The buyer receives code but not evaluation cases, credentials, deployment instructions, configuration, logs, or a way to move providers. Ask for an operational handover and an exit test.
The proposal never says how the system ends
Systems acquire users, data, dependencies, and habits. A responsible proposal says what happens when a model changes, a provider fails, the use case is no longer valuable, the risk grows, or the contract ends. NIST and OECD both treat monitoring and lifecycle management as continuing work, not a launch-day event.
What should you do when the vendor says the proposal cannot be more specific?
Ask which uncertainty prevents specificity and whether the proposal is actually a discovery proposal. “We need to learn first” can be true. It is not a reason to leave every output undefined.
Use this response:
I accept that the final design or metric may change after discovery. Please specify the uncertainty, the evidence this phase will produce, the person who will review it, and the decision we will make with it.
That preserves flexibility without handing over the acceptance bar. It also gives the vendor a fair chance to explain why a detail cannot be known yet.
There are legitimate cases where a vendor cannot name a final model, integration, or target before access and data are available. The proposal can state a hypothesis and a validation plan. It should also state the conditions under which the hypothesis will be rejected. A consultant who can say “we may not build this” is giving you useful information.
The principal exception is emergency or time-critical work where the buyer consciously accepts a narrower review process to reduce immediate harm. Even then, record the temporary scope, authority, review owner, expiry date, and retrospective evaluation. Urgency changes the controls. It does not make controls unnecessary.
What does a practical 48-hour proposal review look like?
You can do a first review quickly if you separate extraction from judgment. Do not start by debating the vendor's choice of model.
Hour 1: extract the commercial shape
Copy the phases, deliverables, fees, dates, assumptions, named people, dependencies, and exclusions into the matrix. Mark every blank. Do not fill a blank with what you assume the vendor meant.
Hours 2 and 3: write the buyer outcome
Describe the current workflow, desired change, baseline, affected users, authority boundary, and principal risk in your own words. If the proposal and your description do not line up, pause the comparison.
Hours 4 and 5: identify vetoes
Check data access, security and privacy review, human authority, evaluation, ownership, staffing, and exit. Ask a specialist for any item you cannot judge safely. Do not turn a legal or security question into a casual content score.
Hours 6 and 7: normalize price and effort
Separate one-time work, recurring work, usage costs, internal effort, support, and change. Note which assumptions can move the number. If two proposals price different phases, compare the evidence and decision each phase buys.
Hours 8 and 9: inspect the risky path
Ask the vendor to walk through one representative case, one missing-input case, one conflicting-data case, and one out-of-scope case. You are checking whether the proposal's words map to a real workflow.
Hours 10 and 11: score with evidence sentences
Give each dimension a 0, 1, or 2 only after writing why. Record the source line, page, or proposed artifact that supports the score. If you cannot write the reason, leave the score blank.
Hour 12: choose sign, revise, or reject
Write the next decision in one paragraph. If revising, send the exact blank rows and a deadline for the new proposal. If signing, make sure the matrix's owners, dates, acceptance artifacts, and exit deliverables are in the contract or statement of work.
The rest of the second day is for specialist review, reference conversations, and contract language. A fast first review should expose uncertainty. It should not replace due diligence for high-risk work.

What should you still verify outside the proposal?
The proposal is evidence about the vendor's thinking and intended work. It is not proof that the vendor can deliver, that the cited references are comparable, or that the contract protects you.
Verify:
- whether the named delivery people have the time and role described;
- whether references involved a similar workflow, risk, data shape, and operating environment;
- whether the proposed technical path can reach your systems under your security rules;
- whether the data can be used for the stated purpose;
- whether legal, privacy, procurement, and security reviewers agree with the terms;
- whether your team can provide the required decisions and review capacity;
- whether the operational cost fits the budget at expected and high usage;
- whether the exit deliverables are technically sufficient, not merely promised;
- whether the vendor's assumptions match what your system actually contains.
Use official guidance as a checklist for what to examine, not as a substitute for your own approval. NIST's AI RMF is voluntary. GOV.UK's procurement guidance is aimed at government. OECD principles are broad. None of them can determine whether a specific vendor, price, system, or contract is right for your company.
How should legal and security review fit?
Bring those reviewers in before you treat the proposal as a final offer. The proposal review can identify the questions. It cannot decide data protection obligations, liability allocation, regulated use, intellectual-property rights, or sector-specific requirements for you.
Ask the vendor to provide the information those reviewers need: component inventory, data flows, subprocessors, retention, access, logging, change process, incident handling, and termination support. The earlier this is visible, the less likely a commercial “yes” is to become a technical or legal “no.”

When should you sign, revise, or walk away?
Sign when the proposal and contract make the work inspectable. That means the problem is specific, the scope is bounded, the important assumptions have owners, the evaluation method is proportionate, the risk controls match the authority of the system, the people are real, the price is legible, and the next decision is explicit.
Revise when the use case is promising but the evidence path has blanks. A good revision request is concrete: “Add the representative-case method and reviewer to the evaluation row,” not “make the proposal more rigorous.” You are not asking for more pages. You are asking for better decisions.
Walk away or defer when the proposal depends on unacceptable data access, hides material third-party risk, has no credible owner, cannot be evaluated, has no exit, or treats refusal to clarify as a sign of confidence. If the vendor will not say what would make the project stop, assume the buyer will be the one paying to find out.
What if the company is not ready?
Do not sign a build because the proposal is good if your team cannot supply the data, decisions, reviewers, security approvals, or operating owner it requires. The right next step may be internal readiness work, a smaller discovery, or a narrower workflow. That is a scope decision, not a failure of the vendor or the idea.
Trust and proposal detail are useful but insufficient. A trusted vendor still needs a shared definition of the work, and a detailed document still needs a reversible first phase if delivery history is untested.
How can Marius Manolachi help after the review?
If you have completed the matrix, you have already done the important first job: you have made the decision visible. You may still want help choosing the smallest useful scope, translating business work into a technical evaluation, or teaching your team to challenge and operate what it buys.
That is the role of Marius Manolachi's AI consulting and tutoring work. He makes existing people capable of building AI products on their own work. Use the matrix to improve the proposal, then bring in help where the team needs a clearer path.

Take the proposals you have, fill the matrix, mark the vetoes, and send back the blank rows. If the vendors can answer with evidence, you have a better basis for signing.
Questions people ask next
How do I evaluate a discovery-only AI consulting proposal?
Evaluate the evidence the discovery is supposed to produce: a validated problem statement, data and workflow findings, candidate options, risks, an estimated next scope, and an explicit go, revise, or stop decision. Do not require production accuracy targets before the discovery has access to representative cases. Do require a reviewable artifact and decision date.
Should price be the most important part of an AI consulting proposal?
Price should be a constraint and a comparison input, not the main quality score. First check whether scope, acceptance evidence, assumptions, ownership, and exit are clear. Then compare prices against the same deliverables and lifecycle costs. A cheaper proposal with undefined work is not a cheaper version of the same project.
What if a consultant says AI performance cannot be guaranteed?
Accept uncertainty about future model behavior, but do not accept uncertainty about the evaluation process. The proposal should specify the test set or case method, quality measures, failure handling, review owner, and what happens if the system misses the agreed bar.
What is the biggest red flag in an AI consulting proposal?
The biggest red flag is a proposal that promises activity but cannot show how the buyer will judge the resulting work. Generic phases, named technologies, and a confident timeline do not replace a deliverable, acceptance method, responsible owner, and decision after each phase.