Case Study — Docker Hub

Hub
Vision

Docker Hub does 20 billion downloads a month, and even the people who run it struggle to define it.

Role
Lead Product Designer
Focus
Blue-Sky Vision
Year
2026
Team
Myself

Overview

Docker Hub is built for downloads, not decisions. Engineers Google for the image, ask in Slack, keep a spreadsheet for the security audit, and only land on Hub when they already know what they want. It does 20 billion downloads a month and is almost never where someone makes up their mind.

Hub is old, and it shows. A decade of teams layering features on features, and an interface that reads like the history of those decisions rather than a tool for making one. Underneath the surface problem is a structural one. It's a storage system with a search bar, which works as long as you know the exact name. The harder questions, which image to trust, which one to buy, which one to enforce across the team, don't live in Hub. It started as one question. What would Hub need to look like for a security-conscious engineer to trust it enough to make a real decision on it? The answer grew into a reframe of what Hub even is. Not a storage system. A marketplace where trust is the currency.

The answer grew into a reframe of what Hub even is. Not a storage system. A marketplace where trust is the currency.

And then the actors in that marketplace are changing, and changing quickly. Engineers are already running teams of agents that write code and open pull requests. Soon those agents will be reading Hub directly, choosing images, enforcing policies, and spending against budgets their humans set. Publishers will run agents on the other side. Same Hub, two new audiences, and real money moving between them. Hub barely serves the humans it has today, and it isn't set up for agents at all. I tested that directly. Five agent personas, real API calls against the real Hub. None of them could complete their task. Later I rebuilt Hub's data layer and ran the same five again. All of them could. The gap isn't the agents. It's what Hub refuses to tell them.

The gap isn't the agents. It's what Hub refuses to tell them.

This wasn't Docker's first run at a Hub vision. I'd been part of earlier ones. This one is different, and it took three weeks, in the cracks between other projects. Not all of it could move at once. Two pieces are moving now: Docker's AI in Hub search, and a Hub MCP for agents. They're the pieces that start unlocking the rest.

Problem

Hub today is one surface trying to do two fundamentally different jobs. Discovery on one side, the moment someone is choosing what to use. Management on the other, the work of governing what a team has already chosen. Mixing them produces a surface that does neither well, and the revenue follows the same shape. The high-margin moment, where a platform engineer chooses what to trust and pay for, lives in the discovery half. Hub today is mostly the management half, billing for storage and data transfer. Docker is monetizing the wrong end of its own product.

The humans Hub is supposed to serve already have a bad time. Only 26% of users reach the browsing section. Only 15% scroll past the first row of results. The trust signals that do exist aren't legible to the people who need them. A participant from Bell Canada, looking at Docker's single-letter security grade, summed it up in nine words: "That C in a shield is a security grade, that went over my head."

The person Hub is failing here has a name inside Docker. Patrick, the platform engineer persona, the one choosing which tools lock down what his organization runs. He can't get auditable answers from the catalog his employer depends on.

Search makes the same point at scale. The 2023 usability research put it plainly: the only way to find something is to already know its name, so users fall back to Google. Named as a problem in 2022, still on the open list in the August 2025 audit. And search isn't a side feature. It's the front door, which makes it both the biggest failure on the surface and the biggest lever on it.

Then there are the new actors. A human reading a README can infer enough to make a call. Agents ask structured questions and need structured answers. My experiment made the gap concrete: five agent personas, real API calls. Three blocked outright. Two degraded to guessing. Zero completed with verification. The Security-First Agent refused to recommend a Node.js image because Hub couldn't tell it four things in a form a program could read: signed, vulnerability-free, verifiable origin, actively maintained. Worse, some endpoints returned empty success responses for questions Hub had no answer to, quietly telling agents everything was fine when nothing was.

All of this had been a slow-moving problem for years. Agents put a clock on it.

Solution

Three weeks isn't enough time to audit 250 images, study ~75 marketplaces and registries from outside developer tools, design across seven surfaces, and prep a 19-slide deck. And the focused time was less. I was running two other projects in parallel. So I didn't work the old way. I worked with Loupe, Docker's internal AI Design Agent, with access to every research transcript, spec, and strategy doc Hub has ever produced. A terminal back-and-forth: I described a direction, Loupe challenged it against the research, I refined. Loupe didn't make decisions. It made the depth possible.

The biggest thing it did was push me outside developer tools. How Amazon surfaces trust. How Steam communicates quality through review volume and sentiment instead of a single score. That research unlocked the defining decision. Tear Hub in half. A Marketplace for discovery and evaluation. A Console for management and policy. The split was risky. Docker's strategy calls Hub a marketplace; users treat it as a registry. Splitting it meant leading users toward a definition they didn't yet hold. The reason to do it anyway: the unified shape was costing Docker the high-margin half of its own product.

Tear Hub in half.

The work moved through three pressure-test sessions with the Hub PM, engineers, designers, and leadership. The goal wasn't consensus. It was finding what wouldn't survive scrutiny.

The first resolved the audience question. Senior engineers pull from the command line. Junior engineers execute someone else's decision. The team landed on Patrick and Sam, two existing internal personas. Patrick is the platform engineer who chooses which images become the company standard. Sam is the dev manager who approves the spend and answers when the audit comes. Patrick picks. Sam pays. The cost was the rest of Docker's audience — Hub had spent years trying to meet every developer where they were, and the breadth was why it served no one particularly well.

Patrick picks. Sam pays.

The second challenged the explore and detail pages. The click data was unambiguous: users arrive with intent and go straight to search. But search needed to do more than match strings. Hub knows almost nothing about its users beyond whether they're signed in. Pulling Docker's AI into search is how Hub starts learning. Every query becomes context. The same session surfaced a legal line that reshaped the trust signals: showing factual security data and letting users decide is defensible. Claiming an image is safe is not.

The third pulled the agent work forward. The team started treating the agent vision as future state. By the end, the call had flipped: the future is coming fast enough that designing near-term Hub without designing for agents was already wrong. That decision is why there's an agent journey in this case study at all.

One assumption got killed early. The first framing had Patrick discovering vulnerabilities for the first time. Wrong. A platform engineer preparing for a compliance audit already runs scanning tools. The gap isn't awareness. It's evidence. Hub isn't a security scanner. It's an audit evidence machine.

Hub isn't a security scanner. It's an audit evidence machine.

Then I built a stripped-down Hub, same shape as the real one, but with trust signals, structured security data, and machine-readable endpoints exposed. The same five agents that failed before all approved in under two minutes. And to show where this leads, I built the agent journey as a full prototype: Patrick sets a budget and permissions, the agent finds the vulnerabilities, hits its spend ceiling, escalates for approval, and applies the policy across the team. Four minutes, two human decisions, an audit that would have taken weeks. It's the same thesis the Hub MCP work now makes real: agents can transact inside Hub if the data supports them, and humans stay in control of what matters.

What landed was a designed argument for what Hub could be, credible enough to be believed and specific enough to be challenged. That fidelity is what gave the team the confidence to start.

My
Takeaways

Before this, I would have called this scope a six-month project. With Loupe in the terminal, it took three weeks, while I was working on other things. The cost of being thorough collapsed, and that's what staff-level design is going to mean from here on.

What I'd do differently is move faster. My own anchoring on what Hub is today was the bottleneck, and tighter feedback loops in week one would have gotten me to the blue-sky work sooner.

I presented the vision to the VP of Product Management, the VP of Engineering, the Head of Design, my design manager, and the PM. Two things caught the room: the search experience with Docker's AI, and making Hub agent-ready, with an MCP as the first step. But the vision landed in a consumed moment. A new CPO, and every ounce of leadership attention pointed at Docker's AI push. A full Hub restructure wasn't getting decided in that room, and waiting for the org to be ready would have killed the momentum. So we didn't wait.

So we didn't wait.

The team started with the pieces that had enthusiasm behind them and, by the numbers, were the right pieces anyway. Both are now real work: Docker's AI in Hub search, and a Hub MCP that lets users connect their own agents to Hub before Docker builds official connectors with the major AI providers. The marketplace vision doesn't get built without context, and search is how Hub finally starts learning who's asking. If that lands, it's the spark that unlocks the rest.