Back to work

NatWest: per-tenant egress for a multi-tenant Kubernetes platform

In one sentence: I designed and built the per-tenant egress solution for NatWest's next-generation platform for business-critical services, so that tenants could reach the on-premise services they need without stepping on each other, and so that the platform team would know straight away when something was wrong.

Client
NatWest Group (London, UK)
Sector
Banking
When
December 2024 to September 2025
Role
Platform engineer in the core platform team
Stack
AWS, EKS, Istio, Terraform, OpenTelemetry, Amazon Managed Prometheus, Amazon Managed Grafana, CloudWatch

The context

NatWest was building the next generation of its multi-tenant platform to run what the bank calls Important Business Services, which is the regulator's term for the services that must not fail. The platform runs on EKS. The tenants are the bank's own application teams, each of them with their own workloads, their own on-premise dependencies and their own view on how urgent their requests are.

I joined the core platform team, mostly to work on tenant isolation and observability. This case study is about the first of those. The observability work has its own write-up.

The problem

On a shared platform, egress is one of the places where tenants end up sharing more than they should. If everybody goes out through the same gateways, one tenant with a chatty application or a misbehaving retry loop can slow everybody else down. This is the well-known "noisy neighbour" problem. In a bank there is a second problem on top of it: each tenant should only be able to reach the on-premise services it is actually allowed to reach, and this has to be provable.

Consequently, the platform needed egress that was segregated per tenant, where each tenant could declare which on-premise services their applications need, and where a failure in one tenant's egress would be detected quickly and would not affect anybody else.

What was hard

Three things, in increasing order of difficulty.

The first was making it self-service. A design where the platform team has to hand-edit configuration every time a tenant wants to reach a new destination does not scale, and it makes the platform team the bottleneck. Tenants had to be able to configure their own egress within limits set by the platform.

The second was knowing when it breaks. An egress gateway that is silently down is worse than one that is loudly down, because the tenant finds out before the platform team does.

The third was plain TCP. Istio is very comfortable with HTTP, but some of the on-premise services the tenants depend on used plain TCP protocols, and the traffic still had to be encrypted in transit and attributable to a tenant. Getting non-HTTP traffic to flow cleanly through Istio egress gateways, wrapped in mTLS, was the technically tricky part of this work.

What I did

I designed, implemented and tested a solution based on Istio egress gateways, one set per tenant. Each tenant gets its own gateway, so a problem with one tenant's egress stays with that tenant. The configuration is all in code and goes through the normal review process.

For self-service, I built a mechanism that lets a tenant declare the on-premise services their applications are allowed to reach, in their own configuration, without needing the platform team to intervene. The platform team still have to approve the changes, so the tenant can't just do what they want.

For detection, I wrote continuous health checks that exercise each gateway and alert the SRE team when one stops working. These are not just "is the pod running" checks. They test that traffic actually gets through.

For plain TCP, I implemented a way of routing TCP traffic by encapsulating it into Istio mTLS, so that non-HTTP protocols get the same encryption and the same per-tenant segregation as everything else. I will spare you the details here, but it took a fair amount of reading of Istio's documentation and a fair amount of testing.

The result

The egress solution went into production and the tenants used it as intended. The plain TCP encapsulation worked, which was not a given when I started. The health checks do what they were meant to do: the SRE team finds out about a failing gateway from an alert, not from a tenant.

What I would do differently

I built the gateways first and the health checks second, which is the natural order and the wrong one. The health checks are a small amount of work and they are what turns an egress gateway from something you hope is working into something you know is working. If I were doing it again I would write the checks first, run them against the gateways as I built them, and I would have caught a couple of configuration mistakes earlier than I did.

I would also have pushed for the self-service configuration to be part of the design from the very first conversation with the tenants, rather than something I added once the gateways existed. Tenants who have been told they can declare their own destinations start thinking about what they need much earlier, and that saves everybody time.

What this means if you are building something similar

Per-tenant egress is not an exotic requirement, but it is one that gets bolted on late and badly if it is not designed in. If you are building a multi-tenant Kubernetes platform in a regulated environment and egress is still "one NAT gateway and hope", I am happy to talk it through.