Start With the Operation, Not the Software

Starting from the operation sounds like advice until you try it, at which point it becomes a question about method. What exactly do you look at, and what do you do with what you see.

Ink drawing of a distribution warehouse aisle: a worker with a clipboard walking beside a pallet on a pallet jack, tall racking on both sides, a roller conveyor to the right and roof trusses overhead.

The operation is a unit you have to choose

Nobody can study an organization. It is too large and it does not hold still. What you can study is one flow of work with a beginning and an end, and choosing which one is the first real decision.

A good unit is small enough to follow end to end in a day and consequential enough that the organization cares whether it moves. A customer request from arrival to fulfillment. A hire from requisition to first day. A change from proposal to production. Pick one and follow it. The instinct to map everything first produces a diagram, which is a different deliverable and not a useful one.

Choosing badly is recoverable but expensive. Units that are too large stop having a clean beginning, and you spend the first three days arguing about where the process starts instead of watching it run. Units that are too small are tidy and prove nothing, because the interesting failures happen at the seams between units and a unit with no seams has no findings in it. The test is whether you can name the moment the work arrives and the moment somebody would say it is done. If two people in the room give different answers to either question, you have found something already, and you should write it down before you go looking for anything else.

Follow one, not the average

Trace an actual instance, with its actual timestamps and its actual detours. Not the typical case, because the typical case is a reconstruction, and reconstructions are tidy in exactly the places you need them to be messy.

Averages hide the thing you are looking for. A process that takes four days on average and eleven days when it touches finance does not have a duration problem. It has a finance seam, and the average is actively concealing it. Follow individual instances until you can predict which ones will be slow, because that prediction is the finding.

The distribution is usually the whole story. Take a request queue where the reported median is two days and leadership is happy with two days. Trace thirty individual requests and the shape typically comes apart into two populations: most clear in under a day, and a minority sit for a week or more. The median is a fact about neither group. Nobody experiences two days. Half the requesters experience a fast process and a quarter experience a broken one, and the quarter are the ones filing complaints that get read as anecdote because they contradict the dashboard.

What separates the two populations is almost never the work itself. It is a property of the request that routes it somewhere different: a dollar amount above a threshold, a customer in a regulated category, a field left blank on intake. Those properties are knowable at the moment the work arrives, which means the slowness is predictable, which means it is designed. That is a much more actionable finding than a duration.

What to write down

Four columns are usually enough. Where the work went. How long it sat before someone touched it. Who decided, and what they did not know when they decided. What a person supplied by hand that the system should have supplied.

The last column is the important one and it is the one people skip, because supplying things by hand stops registering as an event after the second week. You have to watch for it rather than ask about it. The question "what did you have to go and find out" recovers more than "what are the problems with the process".

The manual supply shows up in recognizable forms once you know to look. Someone keeps a spreadsheet next to the system of record because the system does not hold a field they need. Someone maintains a mental list of which approver is responsive and routes around the one who is not. Someone re-keys the same customer detail into a second tool because the two tools do not speak. None of these appear in a process document and all of them are load-bearing. Remove the person and the flow stops, which is the definition of a component, and it is worth noticing that the organization has quietly made a person into one.

Write down who, not just what. A workaround owned by a named individual with fourteen years of context is a different risk than the same workaround owned by a rotating queue. The first one is fragile and invisible. The second one is expensive and at least legible.

Nobody can argue with a timestamp, which is why the trace persuades a room that an evaluation matrix cannot.

Separate the delay from the work

Once you have a trace, split the elapsed time into two piles. Time somebody was doing the work, and time the work was sitting.

The proportions are usually not close. Follow enough of these and the pattern is consistent: the doing is a small fraction of the elapsed time and the rest is waiting, queuing, and being handed over. This matters because the improvement instinct usually targets the doing. Faster tools, better templates, more training. All of it aimed at the smaller pile, while the larger pile is untouched because nobody owns it and it does not appear on anyone's dashboard.

The arithmetic is worth doing out loud in front of a sponsor. If a flow takes eleven days and the accumulated hands-on time is six hours, then a tool that makes the doing forty percent faster removes about two and a half hours from eleven days. It is a real improvement and it is nearly invisible to anyone downstream. Meanwhile a single approval that sits for three days because the approver reviews in a weekly batch is worth more than the tool, and it costs nothing but a decision about batching.

Waiting is also under-owned in a specific and predictable way. Every step in the flow has an owner, and every gap between steps has none. Ask who is accountable for the three days between submission and review and the honest answer is usually that the question has never been asked. Work sitting in a queue is not on anyone's list because it is not in anyone's queue. It is between queues, which is precisely why it is where the time goes.

Then, and only then, ask what it needs

With a trace in hand the question changes shape. It stops being "what should we buy" and becomes "what is the smallest change that removes this wait".

Sometimes the answer is software. Often it is a change to who decides, or to what information travels with the work, or to when a check happens relative to the thing it checks. Those are cheaper, faster, and harder to procure, which is precisely why they get skipped by organizations that start at the other end.

Moving a decision down one level removes a wait entirely if the person one level down already has the information. Attaching the missing field at intake removes a round trip that was costing two days every time it happened. Moving a compliance check from after the work to during it turns rework into a constraint, which is cheaper because nothing has to be undone. None of these require a vendor and all of them are hard to put in a budget line, which is a fact about procurement rather than a fact about value.

When the answer genuinely is software, the trace tells you which software and, more usefully, how much. A flow with one bad handoff needs an integration, not a platform. The difference between those two purchases is often an order of magnitude, and the only thing that lets you tell them apart is having watched the work.

Why this order is hard to hold

The method is not complicated. Holding to it is, because tracing looks like nothing is happening.

There is no vendor, no schedule, no launch, and for the first week no deliverable a sponsor can look at. The pressure to convert the ambiguity into a purchase is real and it comes from good people who need to report progress. The counter is to make the trace itself visible early. A single annotated flow, showing where a real request waited and for whom, is more persuasive in a room than any evaluation matrix, because everyone recognizes it immediately and nobody can argue with a timestamp.

There is a second source of pressure that is less often named. A trace produces findings about people and their decisions, and some of those findings are uncomfortable in a way a software evaluation never is. It is easier for an organization to conclude that its tooling is inadequate than that its approval structure is. Starting from the operation keeps surfacing the second kind of answer, and the method survives contact with the organization only if whoever is sponsoring it has agreed in advance to hear that.

Hold the order anyway. The cost of starting at the other end is not that you buy the wrong thing. It is that you buy something plausible, deploy it into the same structure, and lose the one window where the organization was willing to look at how it actually works.

Similar operation?

If this describes an operation you are responsible for, the diagnostic is where that conversation starts.

How engagements begin
Continue reading