I spent years in digital transformation at Bitdefender before moving into innovation. Part of that job was rebuilding our e-commerce platform and subscription billing stack. If you want to see what twenty years of business logic looks like, open up a billing system. Every odd renewal rule, regional exception and manual override is there for a reason. Usually nobody remembers the reason, and usually it still matters.
So when Mark Ajzenstadt of Limestone Digital posted a thread this week about where he'd start an AI transformation at a company with 20 years of legacy systems, I read it as someone who has been asked that question from the inside. His answer is good. Below I go through it, adding where my own experience agrees with him and where it pulls a different way.
Ajzenstadt starts from a point most transformation programs skip. Two decades in business leave behind working software, experienced people, and rules nobody ever wrote down. Understanding what already makes the business work comes first.
I'd go further. Those unwritten rules are some of the most valuable knowledge a company has, and also the most fragile. They live in people's heads and in spreadsheets with names like final_v3_USE_THIS. A transformation that starts from a target architecture runs straight over them, and you meet them again later, in production, usually at a bad moment.
His first step is to sit with the people doing the work and follow a real customer request, order or report from start to finish. He wants to learn several things:
- What starts the process, and what counts as finished.
- Which systems and spreadsheets people actually use.
- Where work waits.
- Which exceptions need someone experienced.
- Who owns the outcome when several teams are involved.
Then he checks that account against the records, the code and the documentation.
He makes a sharp point about the gaps. When the process document and the person doing the job disagree, you could be looking at three different things:
- an outdated instruction,
- a workaround the business depends on, or
- a real problem.
Each needs a different response. Treating them all the same is how projects go wrong.
This is where I'd adjust his method. I'm biased toward prototypes, so I wouldn't keep observing and building as separate phases. I'd put a rough version of the workflow in front of the people doing the job while you're still watching them. Nothing surfaces an unwritten rule faster than a prototype breaking it, with the expert sitting next to you saying "no, that's not how it works." Interviews tell you how people describe their work, and a prototype shows you where that description is incomplete.
For the first workflow, Ajzenstadt looks for a recurring bottleneck with accessible data, a clear owner, and an outcome the business cares about. Before building anything, he would record four things:
- how much work gets done,
- how long it takes, including waiting time,
- how much checking and rework it needs, and
- what it costs to run.
He would also collect real cases and ask the people responsible to define an acceptable result, including when the system should stop and ask for help. Those cases become the first evaluations.
I agree with all of it, and I'd add one thing. Deciding when the system should stop is a business decision, not a technical one. An agent that hands a case to a human at the right moment is worth more than one that is usually right and confidently wrong the rest of the time. Settle that boundary with the people who own the outcome on day one, not after the first incident.
His third step gives each part of the system a clear job:
- Existing systems hold the records.
- Proven code handles calculations and predictable steps.
- AI interprets documents, assembles context and proposes actions.
- People handle approvals, judgment and exceptions.
Permissions, approvals and spending limits are enforced in code. The agent's instructions explain the boundaries, but code enforces them.
This is the part I care about most, because it's close to my day job. My team works on scam detection and on the security of AI agents that browse the web. The basic problem is simple. An agent reads content, and content can contain instructions. A web page, an email or a PDF can tell an agent to do something its owner never intended. If the only thing standing between that agent and a refund, a transfer or a data export is a line in its prompt, that line is a request. An attacker can write a competing one.
Older companies have an advantage here that they tend to overlook. Their legacy systems often already enforce permissions, approval chains and limits, because someone required it years ago. Don't route around those systems to make the agent faster. Make the agent go through them.
There's a second security point that rarely shows up in transformation plans. Once an AI workflow handles customer documents or payments, criminals will probe it. Scammers adapt fast. The test cases should include hostile inputs from the start, not just the clean ones your best operator handles on a good day.
Step four is to run the new workflow against the old one on the same cases, then hand it to a small group of users. He tracks completed work, quality, review effort and operating cost. He also insists that the comparison include the time people spend correcting the agent. Failures and corrections become new test cases. A named person owns the workflow, with instructions for reviewing changes, handling incidents and rolling back a release.
Correction time is the number demos leave out, and it decides whether anything actually improved. Ownership is what leadership tends to leave out. From what I've seen, these systems rarely fail on launch day. They fail months later, when the project team has moved on and nobody is sure who is allowed to change anything. Name the owner, write the rollback plan, and run the workflow like the production system it is.
Ajzenstadt says the first project should leave behind three things: a useful capability, evidence of its value, and people who can operate it. The bigger redesign comes only after that evidence exists. That is when you decide which handoffs disappear and where approvals still matter.
In the replies, Jacob Schulman pushed back. He argued it's easier to just start, building for three or four areas at once and doubling down on whichever works. On speed, I'm closer to Schulman. On scope, I'm with Ajzenstadt. Prototype quickly and in parallel while you're learning. Then commit properly to one workflow, with real measurement, limits enforced in code, and a named owner. Fast exploration and careful production can live in the same company. The trouble starts when a demo quietly becomes the production system.
The models will keep getting better and cheaper, and swapping one for another will keep getting easier. That part mostly takes care of itself. The harder work is understanding a twenty-year-old company well enough to change it without breaking what made it work.