Pfizer: getting rid of database passwords on a Kubernetes data platform
In one sentence: I moved the applications on Pfizer's data and BI platform from database passwords to RDS IAM authentication, in a way that the development teams barely noticed, and then fixed the platform's logging while I was at it.
The context
Pfizer runs a number of its data and BI applications on a platform built on EKS and AWS managed services. The platform team was small relative to the number of development teams depending on it, and the day-to-day requests were taking longer to answer than anybody wanted. I was brought in as an additional platform engineer, and I made it my business to learn the platform quickly so that I could take some of that load, and then to take ownership of a couple of problems that had been waiting for someone to have the time.
The problem
The applications were connecting to their Aurora databases with passwords. Passwords need to be stored somewhere, rotated regularly, and distributed to the applications that need them. Every one of those steps is a place where something can go wrong, and a large organisation in a regulated industry does not want to explain to an auditor why a database password has not been rotated since the application was written.
Pfizer wanted the applications to access their databases without passwords at all. AWS supports this through RDS IAM authentication, where the application proves its identity to AWS and gets a short-lived token instead of a password. The difficulty is not turning it on. The difficulty is that every application has to be changed to use it, and there were many applications, owned by many teams, none of whom had asked for this.
What was hard
If the solution required every development team to rewrite how their application connects to its database, it was not going to happen, or it was going to happen slowly and unevenly. Indeed, the platform team would have spent the next year chasing teams to make the change.
So the real requirement was: passwordless database access, with as close to zero changes for the development teams as possible.
What I did
I implemented RDS IAM authentication across the platform, and I hid the mechanics behind a sidecar container. The sidecar container hid away the mechanics of obtaining access to the databases, which meant very few changes for the apps.
In order to make adoption easy, I also wrote a small reference application that shows how to use the new mechanism end to end, so that a team could look at a working example rather than a document. In my experience a working example is worth a lot of documentation.
The second piece of work was observability. The platform's logging and alerting had grown organically and there was room to make it a lot more useful. I defined a standard format for structured logging across the platform, so that logs from different applications can be searched and correlated in the same way. I deployed Grafana Loki in a high-availability configuration, configured the OpenTelemetry Collector to collect container logs and ship them to Loki, and registered Loki as a data source in Grafana. All of this is in infrastructure-as-code, so it can be reviewed and reproduced.
I also onboarded a new team whose job was to use AI to extract insights from data held in Snowflake. I worked out what they needed, created the repositories, permissions and GitHub pipelines, and spent time explaining how the platform works. They were productive within a couple of weeks, which is roughly what onboarding should take and often does not.
The result
Applications on the platform access their databases without passwords. The development teams did not have to rewrite their database code. The auditors have one less thing to ask about, and the platform team has one less thing to rotate.
Logs across the platform now share a common structure and a common home, and can be queried from Grafana next to the metrics. The AI team was onboarded quickly. And the original reason I was brought in, faster answers to the development teams' day-to-day requests, improved noticeably, partly because I was there and partly because the platform got easier to operate.
What I would do differently
The sidecar approach is pragmatic and it worked, but a sidecar is one more container per pod, and on a large platform that adds up. If I were doing it again from a blank page I would look at whether the token handling could live in a shared library or in the connection pooler instead, and pick based on how heterogeneous the applications are. On this platform, with many teams and many languages, the sidecar was the right call. On a platform where everything is written in the same language, a library might be.
What this means if you are in a similar position
If your applications still connect to databases with passwords, and you are in an industry where somebody eventually asks about rotation, this is a well-bounded piece of work with a clear result. It can be done without a big migration project and without upsetting the development teams. I am happy to talk about how.