We made AI design 320 logos. Here’s what surprised us.
by Toby Trembath, Founder / CEO
by Toby Trembath, Founder / CEO
Pioneering platform designed to empower artists, entrepreneurs, and small businesses with the tools to create magnetic and cohesive visual identities without stretching their budgets.
Read moreWe were thrilled to win the prestigious Creative Catalyst funding award from Innovate UK. Read the related press release here
Read moreLoading map...
AI can generate a logo in seconds. The harder question is whether you'd ever want to use one. The demo reels never show the other thirty-one attempts. While building BrandSpark.app, our AI platform for creating brand identities, we needed to know what these models can actually do. Measured, not cherry-picked. So we invented four fake brands and tested eight AI image models on all of them, in a total of twelve different configurations.
The main campaign produced 320 finished logo files. A follow-up added 96 more. Total damage: over £60 in image-generation API tokens, which turned out to be the cheapest design education we've ever bought. One result shocked us straight away: the model that cost eighteen times more produced logos about five per cent better. It wasn't the only surprise.
Each of our four invented brands came with a personality, a fixed palette, and two constraints chosen because we could score them: a geometry rule and a hidden negative-space request. Real briefs are rarely this tidy. Ours had to be measurable.
Artisan sourdough bakery. Warm, hand-made, heritage craft
Overlapping circles and soft arcs · a wheat grain hidden in the negative space
Bookkeeping platform for freelancers. Precise, calm, trustworthy
Strict 45° angles on a square grid · an upward arrow implied in the counter-space
Children's gymnastics club. Loud, playful, energetic
Bouncy rounded forms, no sharp corners · a star hidden between the shapes
Minimalist architecture studio. Restrained, structural, monochrome
Rectilinear, single stroke weight · an N implied by the structural voids
Every approach did two jobs for every brand: a text-free symbol mark, and a wordmark (the name set in type). Four variations of each, because consistency matters and flukes flatter. A vision model (gpt-5.4-mini) scored every output from 0 to 4 on seven weighted criteria and justified each score in writing. Humans reviewed the contact sheets and held the veto.
Here is the same Hearth & Crumb bakery brief, as interpreted by nine of those configurations:









Prompt: Logo mark for Hearth & Crumb, an artisan sourdough bakery. Personality: warm, hand-made, heritage craft. The mark should evoke the warmth of a wood-fired hearth and hand-shaped bread. A text-free symbol only — absolutely no text, letters, numbers, or typography of any kind. Use only these colours: rust (#B4552D), cream (#F4E9DA), charcoal (#2E2A26). Geometry: built from overlapping circles and soft arcs. Negative space: a wheat grain hidden in the negative space. Flat, solid shapes on a plain white background, suitable as a scalable brand asset.
Yes, one of them wrote "SOUNDGHBOURGH BAKERY" on a badge. We asked for no text at all. More on obedience later.
Recraft's premium vector model topped the symbol-mark table with a weighted 3.4 out of 4, and it earned the crown: best in class for editability and consistency. Then you read the price column. The runner-up scored 3.3, a gap of five per cent on the raw scores, and cost $0.017 per usable logo against $0.30. The last five per cent of quality costs eighteen times more. Generate three candidate logos and you're choosing between five cents and nearly a dollar. Build a product on it and that choice compounds into the difference between viable and extravagant.
One vendor sells the same model at three quality tiers: $0.017, $0.064 and $0.222 per usable logo, as we measured it. We re-ran one of the front-runners at all three and braced for regret. On symbol marks all three tiers scored 3.2. That is not a typo: to one decimal place they are identical, and the raw scores actually fall slightly as the price rises, a spread within judge noise. On wordmarks the cheapest tier won outright. In fairness, the dearest tier did match the requested brand colours most reliably. It still lost, at thirteen times the price.
I suspected as much going in; it's why I insisted marks and wordmarks were scored separately. The data confirmed it emphatically. Our marks champion finished sixth of ten on wordmarks. The wordmark winner, Ideogram v3 (2.8, typography reputation intact), placed fourth on marks. Not one approach made the podium in both events. If someone sells you a single model for your whole visual identity, quality is leaking somewhere.
The wordmark:

A symbol mark at the same sizes:

To be fair to the machines, this is typography physics, not an AI defect: a human-designed wordmark dies at 16 pixels too, which is why favicons and app icons are monograms and symbols rather than full brand names. The data adds the useful part. Every configuration lost around a full point of small-size silhouette and one-colour quality when it moved from marks to wordmarks (one leader fell from 3.6 to 1.7). The top strip is our winning wordmark model, porridge by 16 pixels. The bottom strip is a symbol mark for the same brand, from the model with the best silhouette scores in the test, and it's still legible at favicon size. So the warning is about workflows, not models: plenty of "logo in seconds" tools derive the favicon from the wordmark. Insist on a mark for the small stuff.
Every brief named its palette precisely, hex codes and all. One model obeyed: Google's Nano Banana Pro hit all three requested colours as its median performance, comfortably the obedience champion. Two more typically managed two. The rest managed zero or one. An approved palette usually exists before the logo does, and most generators paint straight over it.


The brief asked for deep navy , signal green and white . Nothing else.
Price and quality are barely correlated in AI tooling. The best value in our test cost eighteen times less than the best score and delivered 95% of the quality. That's uncomfortable, because the expensive option feels safer; our data says that instinct gets expensive quickly. Ask whoever runs your generation how they chose their tools. "Evidence" is the right answer.
Ask your provider how they measure output quality. "It looks great" is not a methodology. A tool or agency that can't tell you how they validate AI output, against what criteria and scored by whom, hasn't checked.
AI accelerates concepting; people still own typography and craft. Generated type fails exactly where a brand needs it to work hardest: small, one colour, everywhere. The winning workflow pairs generative range with a designer's judgement.
Eight image models in twelve configurations, all normalised to SVG output for comparison (raster models' outputs were vectorised). Two task classes (symbol marks and wordmarks), four briefs, four variations each: 320 outputs, plus the 96-output quality-tier follow-up.
We rendered every output at 512, 64, 32 and 16 pixels, plus a forced one-colour version, then scored each on seven weighted axes: distinctiveness, freedom from generic-stock clichés, small-size silhouette, one-colour survival, brief adherence (palette, geometry, negative space), consistency across variations, and editability of the vector file.
Those aren't machine-learning metrics. They're a brand designer's acceptance criteria. Will it survive embroidery on a polo shirt? One-colour print? A favicon? Some of those tests are close to pass/fail. Others, distinctiveness and creativity above all, are judgement calls however you score them, and we treated them that way: the judge (gpt-5.4-mini) justified every score in writing, and human corrections overrode the machine.
That veto earned its keep twice: the campaign's two worst scoring bugs, both ours rather than the models', were caught by human eyes on the contact sheets, and one vendor's real per-image price came in seventeen times under our conservative estimate, flipping the cost ranking. It's why every price was verified against live vendor usage data instead of published tables, and justifies our working rule – the bot is advisory; the human is authoritative.
The four brands are fictional and no clients were harmed in the making of these benchmarks. Here is the wordmark winner's best take on each of the four brands, picked by human eyes:




AI didn't replace brand design. It replaced the slowest part of it. The craft now lives in choosing the right idea, refining it relentlessly, and knowing when the machine is wrong. That lesson is built into BrandSpark.app, which you can read more about in our launch announcement. And if you want a brand identity built with this level of rigour, with AI where it measurably helps and craft where it doesn't, talk to us.