One product, seventeen squads shipping into it, and banks reselling the result to their own customers. More than a hundred and fifty engineers across four regions built it, and the business now wanted it run as a service instead of installed on premises, bank by bank. Nobody could say, precisely, how far it was from ready.
Opinions, yes. Evidence, no.
That is a worse position than it sounds. Running a platform as a service is not a lift-and-shift, and a bank regulator does not care that your containers are elegant. The question is part of the trap, too. "Are we ready?" has no answer until someone has said what ready means, and until then it gets answered from wherever each person happens to stand.
- Why "are we ready?" cannot be argued, and what I asked the room to argue about instead.
- What a line had to carry before it went on the list, and what became of the lines that did not last.
- How a list this size stayed alive for months: where it lived, who could see it, and what the weekly meeting was for.
Argue About the Test, Not the Verdict
I call it a religious debate. Programmers use the word for a fight that gives off heat and no light: editors, tabs against spaces, whose language is better. A fight goes that way when the two sides share no test, so nothing either of them could produce would count against them. The readiness argument had that shape. The question could be answered; nobody had written the test.
So I did not ask anyone for a verdict. I went to each team and asked what ready would look like for the part they knew, in terms somebody could check. That was the first move, and the assessment that grew out of it was a list.
I have written about the time I ran an experiment instead of an argument. A platform is too large for a clean experiment, so the equivalent is a test that is written down, with the argument moved onto it.
The assessment turned a religious debate into a checklist.
What a Line Had to Carry
It was not a desk exercise. The criteria came out of those conversations and covered architecture, security, data, operations and compliance, more than a hundred of them by the time we were done. Every criterion was scored, and every criterion had an owner. Every topic carried a measurable goal.
Each topic was written twice, once as what the platform had to do and once as how well it had to do it. The three most important non-functional requirements sat beside every part, and they were built into the solution design itself.
My standing example is going from ten transactions to a thousand: how, and what is measurable about it? The database shows the shape.
Each of those is a sentence you can test. The architecture literature has a formal name for the move, a quality attribute scenario, which ends in a response measure. The Software Engineering Institute's report on quality attribute workshops shows why with a modifiability case: "modify the system" says nothing about how well the system took the change, while "in less than two person-weeks" is something an architect can fail. The report's point is that given enough time and money any change is possible, so the measure is what makes a requirement mean anything.
Every Line Gets a Name
The rule I cared about most was the shortest one. I worked with Architecture, Development, QA, DevOps, SRE and Product until every criterion had a name next to it, and on the cloud side the work was split into five areas, each with its own owner: compute, file storage, persistence, security and network.
A criterion without a name next to it is an opinion. Nobody can be asked about it, nobody can fail it, and nobody is the one who gets it fixed. A name turns a line on a list into something a person has agreed to answer for.
The List Was Allowed to Be Wrong
The first version was not the last. As we learned, some criteria were thrown out and others were added, and the count settled a little above a hundred.
That is the part I would defend hardest. A position can be held forever; a checklist can be corrected, and that is where its authority comes from. It earns trust the way a test suite does, by being changed in the open when it turns out to be wrong. Had the list been closed on the first day, it would have been one more opinion with better formatting.
I have argued on this site that adopting microservices is a readiness decision, across people, processes and technology. The list is what the measuring half of that decision looked like on a real platform.
Where the List Lived
Out of the list came a three-tier roadmap with milestones, action items, owners and deadlines. The first phase was six months long and went to the most critical topics; the second carried optimisation and improvement.
The whole roadmap lived in Jira, and every ticket could be followed on dashboards that senior management could open. Once a week we met on the hard problems and the progress, and what was said went back into the tickets, so that anyone reading them could see what was really going on. Once a month there was a session with senior management to confirm the work still matched what they wanted.
None of that was reporting for its own sake. It kept the list as the place where the argument happened. A criterion that was hard stayed on the board with a name on it, and did not go back to being a mood.
What the List Did Not Do
The API-first work that followed is its own story, told in an earlier essay on treating an API as a product. This one stays with the list, and the list has limits.
I have no number for what it saved, and I will not make one up. The move to AWS that followed cut the platform's running cost by about a fifth, and that belongs to the move. What the assessment did was smaller, and I think it came first. It turned "how ready are we?" from a question people answered with a mood into one they could answer with a list. Once the list was scored, I could tell you which lines were failing and whose they were.