
How to Tell Which Work AI Can Actually Take
The most expensive AI mistake I see is not picking the wrong tool. It is picking the wrong process, spending four months on it, and finding out at the end that the thing was never automatable in the first place.
It usually looks like this. Somebody nominates a process that feels painful. Everyone agrees it is painful. A pilot gets built. Then the work runs into something nobody checked at the start: the data lives in a system with no export, or the output is a document a partner has to sign and will not accept from a machine, or the process turns out to have forty exceptions and only eleven normal cases. The pilot works in the demo and dies in the hallway.
Painful and automatable are not the same property. Almost every failed project I have looked at confused them at the start.
The Failure Rate Is a Selection Problem
The scale of this is well documented and consistently misdiagnosed. MIT's GenAI Divide report traced enterprise custom AI tools through their lifecycle and found that 60% of organizations evaluated them, only 20% reached a pilot, and just 5% reached production. Gartner predicted in 2024 that 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, naming poor data quality, inadequate risk controls, escalating costs and unclear business value as the causes. That was a forecast rather than a measurement, and it is worth reading as one, but every reason on the list is knowable before you start.
That is the part worth sitting with. Data quality, risk tolerance, cost and business value are not things you discover in month three. They are things you can assess in an afternoon if you know what to ask.
Six Questions That Predict Whether It Will Work
Over enough of these, the same six factors keep deciding the outcome. I score them, and I weight them, because they are not equally important.
Can the system reach the data, without a person in the middle? This carries the most weight and it fails the most projects. Not "does the data exist" but "is there an API, a database, a scheduled export that already runs." If the honest answer is that Dana downloads a report every Monday and emails it around, that is not a data source. That is Dana.
Will somebody accept the output as it comes out? Also heavily weighted. An output that always needs a human to rewrite it has not saved anything, it has moved the work. Ask who receives the result and whether they will act on it unedited. If the answer is a customer, a regulator or a lawyer, weight the answer accordingly.
Is it clear when the work should start? A process that begins because an email arrives or a record changes state can be automated. A process that begins because somebody noticed something cannot, until you can name what they noticed.
How much does it vary? Count the exception paths honestly. Ten variations on one process is a rules problem. Fifty is a judgment problem wearing a process costume. The tell is whether the person doing it can describe the rules, or only recognize the right answer when they see it.
Does the person who owns it actually want this? Unweighted in most business cases and decisive in practice. A process owner who wants the change will find the edge cases for you. One who feels audited will find them for the post mortem instead.
What happens when it gets one wrong? Low tolerance for error is not disqualifying, but it changes the design. It means a person reviews before anything leaves the building, and that review time comes out of your savings before you count them.
Score those, weight the first two heaviest, and you get a number. The number is not the point. The conversation you have to have to produce it is the point, because it forces someone to say out loud that the data actually lives in a PDF.
The Answer Is Often No, and That Is Useful
The uncomfortable part of running this honestly is how often the answer comes back negative on something everybody wanted to do.
That is a feature. A process that scores badly is telling you something true and specific: this will cost more than it returns, or it will work in the pilot and fail in production, or you will spend the savings on the review step. Knowing that in week one costs you an afternoon. Knowing it in month four costs you the project, and worse, it costs you the organization's appetite for the next one, which is usually the real damage.
There is a second reason to be strict, and Workday's January 2026 research puts a number on it. Surveying 3,200 employees at companies above $100M in revenue, they found that nearly 40% of AI time savings are lost to rework, meaning correcting errors, rewriting output and verifying results, and that only 14% of workers consistently achieve clear, positive net outcomes from their AI tools. Rework is what a bad output-acceptance score looks like after you ship it. Screen for it up front or pay for it forever.
Where This Points You Instead
When a process fails this test, the failure usually names its own fix, and the fix is rarely a different model.
Fails on reachability? The project is a data project first. Fails on output acceptance? The project is a trust and format project, and possibly a smaller scope where the system drafts and a person approves. Fails on variability? Split it. The eleven normal cases are often automatable on their own while the forty exceptions stay with a person, and that split is frequently the whole win.
This is why I keep coming back to the argument in Context Is the Whole Game. Most of these failures are context failures wearing a technology costume. The data is not reachable, or it is stale, or it contradicts itself, and no model choice fixes any of that.
The Takeaway
The 5% of enterprise AI tools that reach production are not the ones with better models. They are, overwhelmingly, the ones aimed at processes that were automatable to begin with.
You can tell the difference before you spend anything. Ask whether the system can reach the data without a person in the middle, whether somebody will accept the output unedited, whether it is clear when to start, how much it varies, whether the owner wants it, and what happens when it is wrong. Weight the first two the heaviest, because they fail the most projects.
Then be willing to hear no. A screen that only ever says yes is not a screen, and the four months you save on a process that was never going to work is the cheapest return you will get all year.
Tracy Thayne* is the founder of Expona, an AI-powered operational intelligence platform for B2B marketing. Read the Expona founder story or subscribe to the blog (below) for weekly insights on context, AI, and the operating model of the next decade.*
Subscribe
Get notified by email when we publish a new post. No spam, unsubscribe anytime.