Insights / AI
What to Automate First With AI: A Business Audit Framework
You do not have an AI problem; you have too many candidate processes and no way to pick one. Five questions produce a single winner you can finish in two weeks.
Why picking the first AI automation is a scoring problem
What to automate first with AI is not a creativity exercise. It is a ranking. Most small companies can name eight ideas and implement none, because none of them was forced to beat the others on volume, error cost, repeatability, data quality, and owner.
Five questions, one winner. The winner must be finishable in two weeks with a human on the send button. If your favourite idea fails that test, it is not first. It might be third. Third is allowed to wait.
I use this when I audit operations — the same habit I use looking at a queue inside a hosting company like VPSGrid, or a software delivery conversation at BlocksGenie. Boring work with an owner beats a glamorous agent (software that takes actions without a person in the loop each time) with a slide.
The five questions to score each candidate
Score each candidate from 1 (weak) to 5 (strong). Do not debate half points. If you cannot answer, score 1. Missing information is a weakness, not a mystery to romanticise.
1. Volume
How often does this happen in a normal week? Daily beats monthly. A process that runs four times a year will not teach you a loop and will not pay back the attention. Count tickets, invoices, reports, or calls. Guessing “a lot” is a 2.
2. Error cost
What happens when it is wrong? If a bad draft is an internal edit, that is a good first target (score high for “cheap to be wrong”). If a bad action refunds money, files a legal claim, or emails a customer something false, score this as expensive — and keep a human in the loop, or leave it manual for now.
Note the direction: you want high volume and survivable errors for the first automation. Catastrophic error cost is a reason to supervise, not a reason to go first with autonomy.
3. Repeatability
Can you write the steps? If every instance is a unique negotiation, it is judgment, not a candidate. If 80 percent of instances follow the same path and 20 percent are exceptions, you can automate the 80 percent and route the rest. If you cannot describe the 80 percent, stop.
4. Data quality
Can you reach the text, fields, or examples without a long research project? Reply templates need past tickets and a policy paragraph. Invoicing needs clean due dates and emails. An agent that “understands the business” needs a corpus you do not have. Bad data scores 1. Do not automate a mess; you will scale the mess.
5. Owner
Who checks output next month, including on holiday? If the answer is “we’ll all keep an eye on it,” score 1. High volume with no owner dies after the workshop. The owner must have time on the calendar, not a title.
How to score and pick a single winner
List three candidates only. More than three is procrastination. Score all five questions. Add the numbers. If two tie, pick the clearer owner. If still tied, pick the cheaper error. Write the winner on the board and put the other two on a dated parking lot — 45 days out, not “later.”
A score below 15 is usually not first. A score with owner = 1 is never first, regardless of the total. A score with data quality = 1 is a data cleanup project, not an AI project. Call it that. Budget it that way.
Finishability is the hidden sixth constraint. If the winner needs a new warehouse, a CMS rebuild (content management system, such as WordPress), or a custom model, it is not a two-week pilot. Demote it. Promote the candidate that can run with tools you already have. Buying software is not the strategy; winning the audit might mean buying nothing.
| Question | Score of 5 looks like | Score of 1 looks like |
|---|---|---|
| Volume | Daily or many times per week | A few times a year |
| Error cost | Wrong draft is an internal edit | Wrong action moves money or legal position |
| Repeatability | You can number the steps | Every case is a unique negotiation |
| Data quality | Examples and fields exist today | You would need a new corpus or warehouse |
| Owner | A named person with calendar time | “We’ll all watch it” |
Worked example: support reply templates vs invoicing vs a shiny agent
A 12-person services firm. Three candidates. Same scoring rules. No vendor in the room.
Candidate A — support reply templates
Volume: ~40 tickets a week that cluster into five types (5). Error cost: a bad draft is caught before send if a human clicks (4). Repeatability: five templates cover most of the queue (5). Data: two years of tickets and a policy page (5). Owner: the support lead already lives in the helpdesk (5). Total: 24.
Candidate B — invoicing reminders
Volume: ~25 invoices a month, reminders on a subset (3). Error cost: a wrong reminder to a client who already paid is embarrassing but survivable (3). Repeatability: dates and templates are clear if the books are clean (4). Data: the books are mostly clean, a few missing emails (3). Owner: the bookkeeper, who also runs payroll and is already at capacity (2). Total: 15.
Candidate C — an unsupervised “company agent”
Volume: undefined; it would “handle questions” (2). Error cost: it would talk to customers with no send button (1). Repeatability: not specified (1). Data: “we’ll connect everything” (1). Owner: the founder, who is in sales all day (1). Total: 6.
Winner: reply templates. Not because support is glamorous. Because it is the only candidate that is frequent, repeatable, data-ready, cheap to get wrong under supervision, and owned. Invoicing is a fair second at day 45 if the bookkeeper’s time is protected. The agent is a demo. It is not first. An AI automation consultant who tries to start with C is selling a demo, not a first project.
Run the same table on your three ideas this week. If C wins on your whiteboard, you have not scored honestly. Recheck owner and error cost.
What to leave manual on purpose
Leave manual what is rare. Exceptions are where your reputation lives. A model that has never seen your worst ticket should not be the first voice on it.
Leave manual what is irreversible until you have a long, boring log of supervised success. Refunds, legal language, medical or financial advice, public posts, access changes. Drafts maybe. Send no.
Leave manual what has no owner. Automation does not create ownership. It creates more output for an absent owner to ignore.
Leave manual what you cannot measure in two weeks. “An AI copilot for the whole company” has no metric. It is a slogan. Slogans do not get kill dates, so they never die. They just bill.
Leave manual the close, the recovery conversation, and the moment a senior has to say no to a client. You can draft the email. A person still decides.
How to run the audit in ninety minutes this week
Minute 0–15. Three candidates on a board. No extras. If people keep adding, park them on a dated list.
Minute 15–45. Score the five questions out loud. Argue with examples, not with hope. “Last Tuesday we had twelve of these” beats “this will scale.”
Minute 45–60. Pick the winner. Write the current steps as they exist. Name the reviewer and the day-14 metric. Put the kill date on the calendar.
Minute 60–90. Open the tools you already have. Do not shop. Sketch the supervised loop: trigger, draft, human click, write-back (saving the result into the software you already use). If that sketch requires a new platform, you picked the wrong winner. Go back to the scores.
After the ninety minutes, run two weeks. Keep or kill. Only then consider a new tool or a build. If you want a third party in the room, use a named package. Implementation, development, licences, and ongoing Slack are not included.
| Package | Price | What you get | Best for | Not included |
|---|---|---|---|---|
| AI Discovery Call | $149 · 45 min | Direction in the session | You are not sure the candidates are AI problems | Written summary, build, licences, Slack after |
| AI Business Audit (featured) | $299 · 60 min + post-session summary | What to automate first, in writing | You want a third party on your artefacts | Development, licences, ongoing Slack |
| AI Implementation Blueprint | $599 · written roadmap | Sequence in writing when implementation is imminent | A winner is picked and someone is about to build | Coding, deployment, licences, ongoing management |
After the winner: write, pilot, measure, then maybe build
Write the as-is steps. Pilot with supervision. Measure. Keep or kill. Train a second person. Then, if built-in tools fail because data cannot be pasted or the step must run unattended, design custom work. That order is the whole discipline.
Skipping to custom because a developer has availability is how you automate the wrong thing at the highest rate. Skipping to a new seat because a demo was clean is how you delay the audit. Both feel like motion. Neither is the first automation.
If you do nothing else: score three candidates tomorrow with the five questions, pick one winner, and refuse to discuss agents until that winner has a day-14 number. You will already be ahead of firms that are still collecting tools. When you need a person to apply this to artefacts you do not want to sort alone, hire an AI consultant for small business on a capped package — not an open retainer.
Common scoring mistakes
Scoring hope instead of last Tuesday. “This will be huge when we launch the new offer” is not volume. Volume is what already hits the queue. Automate the present. The future offer can have its own audit when it exists.
Treating error cost as a reason to go first with autonomy. High harm means more supervision, or manual. It does not mean “we should put our best model on it.” Your best judgment stays with the people who already carry the licence, the relationship, and the regret.
Giving owner a 5 because a founder cares. Care is not calendar time. If the founder is in sales, they are a 1 unless they delegate. Delegate in the room, in writing, with a name. Then score that name.
Averaging scores to keep the peace. If marketing wants content, support wants reply templates, and finance wants reminders, you still pick one. A committee winner is usually the shiny agent. The framework exists to prevent that. One winner. Dated parking lot for the rest.
Skipping data quality because a vendor promised to “connect everything.” Connection is a project. If the emails are missing or the articles contradict each other, you are scoring a cleanup job. Cleanup can be the first project. It is not AI. Name it honestly so it gets a mop, not a model.
The two weeks after you pick
Day 1: write the as-is steps and the review rule. Day 2: run five live items supervised. Days 3–10: the owner samples daily and logs corrections into the prompt or the knowledge article — not into a private chat. Day 14: keep, change the rule, or kill. No new tool during these fourteen days unless the process cannot run without it, which usually means you picked a 1 on data quality and should have stopped.
If you kept it, day 15 is for a second person, not for a platform. If you killed it, take the next candidate off the parking lot and score again. Do not skip scoring because you are embarrassed. A kill is data. Data is how a 20-person agency stops paying for a demo that never became a process.
This is also when people call an AI automation consultant too early. You do not need one to run reply templates in a helpdesk you already own. You need one when the winner requires a loop across systems, a tool-versus-build choice, or a pilot you will not design yourself. Until then, the five questions are the product. Use them.
Questions people ask
What should I automate first with AI?
High-frequency, repeatable work with decent data and a named owner — support reply templates, lead research, reporting summaries, invoice chasers. Not your most sensitive decisions and not a shiny agent with no loop.
What are the five questions?
Volume, error cost, repeatability, data quality, and owner. Score each candidate. Automate the highest score you can finish in two weeks. If two tie, pick the one with the clearer owner.
Should sales be automated first?
Only the repetitive parts: research, drafts, logging. Do not automate the close, the recovery conversation, or unsupervised booking on day one.
What should stay manual?
Irreversible actions, rare exceptions, work with bad data, work with no owner, and anything you cannot measure in two weeks. Manual is a valid outcome of the audit.
Support macros vs invoicing vs an agent — which wins?
Usually reply templates: volume, repeatable, data you already have, a reviewer in the helpdesk. Invoicing wins if volume and data quality are high. An unsupervised agent almost never wins as the first project.
How do I use this without a consultant?
List three candidates. Score the five questions. Pick one winner. Write the current steps. Run a two-week supervised pilot. Keep or kill. Hire only if you are about to spend on tools or a build.
When is a low error-cost process still a bad first pick?
When there is no owner, the data is a mess, or it happens twice a year. Volume without ownership dies after the workshop. Rare work never teaches you the loop.
Does this replace an AI Business Audit?
It is the same thinking. The $299 session applies it to your artefacts and sends a summary. The framework is free to run this week on a whiteboard.