All insights

    AI

    What an AI Pilot Actually Costs, Including the Parts Nobody Quotes

    Big Sky Consulting Group · August 24, 2026 · 7 min read

    The quote is not wrong. It is just answering a different question

    A vendor sends you a proposal for a pilot. Fifty thousand dollars, ten weeks, one process, one team. Your CFO looks at it and sees a number small enough to approve without a board conversation. That is the point of the number.

    Nine months later the pilot has cost something closer to two hundred thousand, most of it spent by your own people, and the vendor is genuinely confused about why you are unhappy. They delivered what they quoted. They quoted the software.

    That gap is not fraud, and treating it as fraud will make you a worse buyer. It is a scoping convention. Vendors price the part they control, because it is the only part they can commit to in writing. Everything on the other side of the line, the state of your data, the shape of your process, the willingness of the people who run it, is your side of the table. Nobody quotes it because nobody can. That does not stop it from being the majority of the bill.

    Here is what sits on your side of the line.

    Data readiness is not a task. It is the project

    Every pilot proposal contains a sentence like "customer will provide access to relevant data." Read that sentence as a purchase order with the amount left blank.

    The number varies but the shape does not. Practitioner benchmarks routinely put data preparation at 30 to 50 percent of build cost on a mid-market dataset, and Gartner's own reading of why generative AI projects die after proof of concept puts poor data quality first on the list, ahead of cost and unclear business value. Their 2024 prediction was that 30 percent of GenAI projects would be abandoned after PoC by the end of 2025. The revised figure is worse.

    What makes this expensive is not the cleaning. It is the discovery. You find out in week four that the field everyone calls "customer" means the billing entity in one system and the ship-to site in another, that four of your six branches stopped using the status codes in 2021, and that the one person who knows which records are real retires in March. None of that is in a data dictionary. It is in someone's head, and getting it out takes their time, not the vendor's.

    Budget the hours of your best operations person, not a contractor's rate. That person is the constraint, and they already have a job.

    Integration is priced per system, and you have more systems than you think

    Integration surveys keep landing in the same place: legacy system connectivity sits just behind data quality as the thing that stalls projects, and integration work commonly runs 15 to 25 percent of a build.

    The trap is the count. You scoped the pilot against the ERP. Then it needs identity, so add the directory. Then it needs to write back, so add the ticketing system. Then someone points out that the process actually starts in a shared mailbox. Each connection is small. The set is not, and the set is discovered sequentially, which means each one arrives as a change order rather than as a line item you compared against another vendor's line item.

    We ask a version of this question in every diligence conversation, and it is the one that most reliably changes the price: how many systems does this process touch on a bad day, not a good one? The exception path is where the systems are.

    Change management is the line item that decides whether any of it was worth it

    MIT's GenAI Divide study, the one that produced the widely repeated finding that 95 percent of pilots showed no measurable P&L impact, is more useful for its second finding than its headline. Ninety percent of the workers surveyed were already using AI tools daily on their own initiative, while only 40 percent of their employers had official subscriptions. The people were not resisting the technology. They were resisting the deployment.

    That distinction is worth money. A pilot that ships into a team with no changed incentives, no changed queue, and no changed definition of done produces a tool people work around. You have then bought a second process running in parallel with the first, and paid for both. S&P Global's 2025 survey found the share of firms scrapping most of their AI initiatives rose from 17 percent to 42 percent in a single year, with the average organization killing 46 percent of its proofs of concept before production. Very little of that is model quality.

    The line items are training hours, a period of dual running where output drops, someone owning the exception queue, and a supervisor whose measured performance does not get worse for using the new thing. Change management does not appear on a software quote because it is entirely your payroll.

    This is the general shape of the problem. Which parts apply to your process depends on answers only your systems can give.

    Put us on it, from $5,000

    Inference is the cost that arrives after the pilot succeeds

    The pilot cost is a one-time number. The system cost is a rate.

    Gartner now expects at least 50 percent of GenAI projects to exceed their budgeted costs by 2028, and their stated mechanism is worth sitting with: inference is projected to be at least 70 percent of a model's lifetime cost, because it recurs on every single call. Their per-user estimates for real deployment run in the region of $11,000 to $21,000 per user per year once you include inference, embeddings, maintenance, and licensing.

    Agentic systems change the exponent, not the coefficient. One user request becomes a plan, five retrievals, three tool calls, and a verification pass. The demo you priced ran one query. The production system runs the whole loop, on every ticket, forever.

    A pilot run on a vendor's promotional credits tells you nothing about this. Ask what a month of full-volume traffic costs at list price, and ask it before the pilot, because after the pilot you are negotiating from inside a system you already depend on. This is the same discipline we describe in how to evaluate an AI vendor when you are not an engineer: the questions that matter are the ones that are cheap to ask in April and expensive to ask in November.

    The overrun is predictable, which means it is budgetable

    Reported budget overruns on AI projects cluster around 40 percent above the original figure, and roughly two thirds of projects exceed their estimate at all. A number that consistent is not a series of accidents. It is a systematic omission in how the estimate gets built.

    So build the estimate differently. The vendor quote is one of five lines, not the total. The other four are your data hours, your integration surface, your change management, and your run rate. You will not get any of them exactly right. You will get all of them closer than zero, which is what they are on the proposal in front of you right now.

    And be honest about which line is load bearing. If the data line dwarfs the software line, you are not buying an AI pilot. You are buying a data remediation project with a model attached, and it should be scoped, sequenced, and priced as one. That is often the right thing to buy. It is rarely the right thing to buy by accident. Call it what it is, or you will end up paying for a model that has nothing good to read.

    We say this to clients often enough that it has become the house position: a pilot that proves your data is not ready has still returned its cost, as long as you find out in month two rather than month nine. The failure mode is not learning it. The failure mode is paying full price for the lesson. The same argument shows up in the case against pilot purgatory and in how institutions should size AI investment against actual return.

    Where this stops being an article

    What we cannot do in writing is tell you which of these five lines is the big one for you. That depends on the state of your master data, the age of the systems the process touches, how many exceptions your team absorbs quietly, and whether the people running the process today believe the change is happening to them or with them. Those are answerable questions. They are just not answerable from here.

    If you have a pilot proposal on your desk and the total looks manageable, that is exactly the right moment to have someone price the other four lines before you sign. Book a consult and bring the proposal. We will tell you what the real number looks like, including the case where the honest answer is that you should not run this pilot at all.

    AI PilotAI BudgetingAutomation ROIBuy-Side Advisory

    Want this applied to your business?

    A paid working session on one of your processes, ending in a written implementation plan. Not a sales call. Consults start at $5,000, and you see the price before you commit.

    Book a consult