Back to work

HMRC: both sides of a multi-tenant platform

In one sentence: On HMRC's HAWK platform I worked on the platform itself and I was the platform team's point of contact for one of its tenants, and doing both at once taught me more about why tenant onboarding is slow than either job would have on its own.

Client
HM Revenue & Customs (UK)
Sector
Government
When
June 2023 to March 2024
Role
Platform engineer, HAWK platform team, remote
Stack
AWS, EKS, RDS, Terraform, HashiCorp Vault, Keycloak, GitLab CI, Helm, Prometheus, Grafana

The context

HAWK is a turn-key runtime platform built by HMRC for HMRC. It provides Kubernetes (EKS), databases (RDS), secrets management (Vault), identity (Keycloak), monitoring (Prometheus and Grafana) and CI/CD, so that other HMRC projects can bring their applications and not have to build all of that themselves. Those projects are the platform's tenants.

I joined the HAWK platform team as a platform engineer. About half of my time went on the platform itself: new features, improvements, bug fixes, documentation. The other half went on being the platform team's point of contact for one particular tenant.

The problem

Every multi-tenant platform has the same tension. The platform team wants tenants to use the platform the way it was designed. The tenants want to ship their application and they experience the platform as a set of rules that get in the way. Both sides are right, and the gap between them is where onboarding slows down, tickets pile up, and, in the worst cases, tenants quietly start building their own infrastructure on the side.

HAWK was a good platform. It still had this tension, because all platforms do.

What I did on the platform side

Platform work is mostly unglamorous and that is fine. I implemented features and improvements that tenants had asked for, fixed bugs, and updated the documentation. When you are also talking to a tenant every day, you get a very direct sense of which bugs matter and which documentation pages people actually read, and I used that to decide what to work on first.

What I did on the tenant side

I was in charge of looking after one tenant and making sure they could use the platform to its full abilities. In practice this meant:

  • Writing and maintaining their CI/CD pipelines in GitLab, so that their code got built, tested and deployed to the right environments without anybody having to think about it.
  • Deploying and maintaining their environments on the platform.
  • Troubleshooting the issues they ran into across those environments, which is where you learn what the platform's error messages look like to somebody who did not build it.
  • Explaining how the platform works, and how to get the most out of it, to engineers whose job was their application and not the platform.

The tenant provided excellent feedback on the work, and in particular on the communication. On a multi-tenant platform, the quality of the communication between the platform team and the tenants is not a soft skill. It is part of the platform itself.

What I learned that I still use

Looking after a tenant while building the platform is the best way I know to find out what actually makes onboarding slow. A few things I took away, and that I now look for first when I review somebody else's platform:

The documentation is written from the inside. Platform teams document what the platform does. Tenants need to know what to do. These are different documents, and most platforms only have the first one.

The first pipeline is the hardest. Once a tenant has one working CI/CD pipeline that they understand, they can copy it forever. Until then, every deployment is a support ticket. Giving a new tenant a working, minimal pipeline on day one pays for itself many times over.

Errors should name the fix. When something goes wrong on the platform, the tenant sees an error message. If that message tells them what to do, they do it. If it does not, they raise a ticket, and the platform team spends some precious time working out what they would have known instantly.

Somebody on the platform team should own each tenant. Not to do their work for them, but to be the person they call. Rotating this role around the team means everybody on the platform team eventually understands what the platform feels like from the outside.

The result

The tenant was able to use the platform to its full abilities, which was the brief, and their feedback on the work was excellent. The platform got a steady stream of fixes and improvements that came directly from watching a real tenant use it. And I came away with a set of opinions about tenant onboarding that I have since applied at a bank and at a pharmaceutical company, which suggests they are not specific to HMRC.

What this means if you run a multi-tenant platform

If your tenants are slow to onboard, if your ticket queue is full of questions the documentation was supposed to answer, or if you suspect that some teams are routing around the platform rather than through it, the problem is usually in the gap I described above, and it is usually fixable. I am happy to come and look at it from both sides.