Acquired Taste
What building an AI tool for brand development taught us about creative judgment—and whether it stays human.
Earlier this year, we built an internal experiment at A Color Bright we called taste. It takes a brief and helps explore brand narratives, visual directions, and references.
The starting point was a gap that anyone working with these models on design might recognise. In text, everything sounded convincing. The model could describe a direction, explain its strategic rationale, and give it a name that looked perfectly at home in a presentation.
Then we would get to the images.
The references were obvious. They lacked cultural context. Often, they didn’t express what the text had just promised. A direction could sound specific and interesting, then turn into a moodboard you’d seen a hundred times before.
At the time, that felt like a reassuring boundary for a design studio. Generating things was getting cheaper. Knowing what was interesting, what belonged together, and what was right for a particular brand still seemed to depend on us.
taste started as an attempt to work on that gap. What happened made us less confident that the boundary would hold.
Giving the system something worth looking at
Finding a useful reference involves more than matching a description.
An image can contain all the right ingredients and still be the wrong choice. Another can come from an apparently unrelated world and open up the whole direction. The connection might be in its attitude, its cultural associations, or the way it resolves a tension. Those are difficult things to capture with a few adjectives.
We approached this from several sides.
One experiment was to make the system think through analogies. We asked taste to look beyond the client’s category—to music, architecture, rituals, subcultures—and find relationships that could change how we approached the brief. Comparing a fintech brand to a Swiss bank, for example, gets you familiar territory. Comparing it to a Japanese convenience store gives you somewhere else to go. The useful part is working out which relationships carry across and how they might change the direction. The connection had to go deeper than borrowing a look.
We also worked on the criteria the system used. The project includes a library of design criticism and principles distilled from it: conceptual coherence, restraint, appropriateness, avoiding category clichés. We tried to articulate some of the things that make us accept one direction and dismiss another.
But a major improvement came from giving the tool a reference library.
The original collection grew to roughly 2,200 images, bringing together our studio’s reference collections we had built over the years. This changed the range of material the system could work with. Selection was already happening before it generated a board, through what entered the library.

That mattered enormously. It also raised a second problem: how would the tool find its way through those images?
From describing images to comparing them
We used tags and descriptions to connect references to a territory. But describing images introduces another translation step. Two references can share words such as “confident,” “warm,” or “technical” while feeling very different.
We wanted to use an image itself as the starting point for a search. If a reference felt right, could the system find others with related visual qualities—even if they depicted completely different things? That would let us explore connections our tags and descriptions missed, and give the system something concrete to work with when we selected a reference.
That’s why we then brought in CLIP, an image model running locally. It turns each image into a numerical representation, called an embedding. Comparing those representations lets the system retrieve visually related images without requiring us to describe every quality we were responding to.

In taste, that supported several different actions. We could find more images like a selected reference. Images similar to ones we’d liked could receive a higher ranking. We could also look for an alternative that remained relevant while avoiding something too visually similar.
Alongside that, we added selection rules to prevent boards from being dominated by the same sources or clusters of references.
The results improved massively.
This wasn’t us training a new model. We were combining existing models with a reference collection, retrieval logic, and feedback. But the difference was substantial enough to change how I thought about the original limitation.
The system still needed our choices. It had become much better at doing something useful with them.

Where was taste?
I don’t think this means we had given a machine cultural judgment.
Our judgment was present in the collection, in the criteria we supplied, and in what we selected or rejected. Visual similarity doesn’t establish that a reference is culturally appropriate. Nor does an unexpected pairing automatically make an interesting idea.
The models still struggle with the non-obvious connections that make a direction worth pursuing.
But we had made parts of our judgment available to the system. We could express some of it through examples, preferences, and relationships between images. And the system could use those signals to produce better results.
That is harder to reconcile with the comforting idea that taste is an indivisible human quality which sits permanently beyond the model.
Our experiment doesn’t establish how far this can go. It does make me hesitant to draw a firm line around where it must stop.

A temporary advantage?
I’ve been thinking about this again after looking at Taste Labs. They describe working with frontier labs on taste and design capabilities, and explicitly include post-training in their work.
Their approach operates at a different level from our little tool. We supplied references and feedback around existing models. Post-training can change the model itself.
It would be a leap to say that this solves the problems we encountered. It doesn’t prove that models can reliably find culturally interesting references today, or that better visual output amounts to better judgment.
But it does suggest that the gap is being treated as something to work on.
Having seen how much our own results improved with better material and a way to navigate it, I increasingly expect some of the capabilities we assembled around the model to become part of the model.
For a design studio, that complicates the claim that AI will handle production while we retain taste and judgment. Those qualities still matter. They may become more important as the volume of possible work increases. Their importance, however, doesn’t guarantee that they remain exclusively ours.
Building taste made me less comfortable positioning our future around the assumption that this gap will stay open. It also made me more interested in what we can do as it closes.
In our own process, we’re using taste to give creative development more breadth and depth. We can explore more interpretations of a brief, earlier in the conversation, look beyond the client’s category, and follow an unexpected connection further before committing to a direction. A reference can become the starting point for another search; an analogy can open up a different narrative. The collection we’ve built over years becomes something we can keep exploring in new ways.
That gives us more to exercise our judgment on. Comparing directions helps us articulate what makes one more appropriate, where another becomes predictable, and which deserves more work. The value is in how far we can develop an idea through that process, as well as how many possibilities we can consider.
That’s what makes this exciting for us as a studio. Better models give us room to ask more of the creative process: explore further, question our first answers, and develop directions we might otherwise have left unexplored.