If you have four AI concepts but can’t explain which one is better
When you ask several AI models to create a website or logo, the results pile up quickly. The problem comes next. One concept is clean but does not fit the brand; another stands out but has a weak information hierarchy. It is easy to end up choosing based on a model’s name or a first impression.
This ambiguous choice is exactly what DesignArena explores. It hides the model names behind results made with the same prompt, has people repeatedly choose between two at a time, and turns those choices into ranking data. Rather than asking people to explain their “taste,” it collects which of the two they chose.
This approach can help narrow down model candidates, but being ranked first does not automatically mean a model is best at accessibility or conversion. DesignArena’s rankings measure the preferences participants show in comparison situations.
One evaluation includes five choices
The official methodology is more structured than a simple popularity vote. In one session, four different models receive the same prompt, and users compare two results at a time without knowing the model names. The first two choices separate winners from losers, followed by additional matchups and a tiebreaker, for five choices in total that determine ranks one through four.
The design principle to take away is “anonymous, same input, repeated choices.” Hide model names to reduce brand preconceptions, keep inputs identical to align what is being compared, and compare several pairs rather than making a single choice to produce a full ordering.
The cumulative leaderboard gives every pairwise comparison the same weight and calculates ratings with a Bradley–Terry model. Its main charts exclude models with fewer than 50 comparisons, and results are generally marked as provisional until they reach 200 comparisons. These are operating rules for ranking stability, however, not criteria for validating design usability or business outcomes.
A 12x ARR increase in six months is a growth figure disclosed by the company
Grace Li, co-founder of operator Intelligence, announced that annual recurring revenue (ARR) grew from $5 million to $60 million in six months and that the service reached 5.5 million people in more than 190 countries. That is a 12x increase by calculation. The company also announced a $7.9 million seed round led by Index Ventures in August 2026, and TechCrunch likewise reported $60 million in ARR based on the company’s description at the time.
The figures stand out, but they should not be read as equivalent to recognized revenue or profitability. The ARR formula, contract terms, customer retention, and gross margin have not been disclosed, and the user count differs between the founder’s stated 5.5 million and TechCrunch’s reported 5.3 million. The supported conclusion goes only as far as this: ARR rose quickly according to the company’s own disclosure.
The shutdown of Yupp, which operated a similar comparison service, is another reason not to oversimplify DesignArena’s success formula. Yupp compared 800 models and aimed to sell preference data to labs; according to the company, it acquired 1.3 million users and some customers. It raised $33 million in 2024 but closed before its service had been live for a year. Its founders said product-market fit was not strong enough and that the model landscape changed quickly. Public information alone cannot establish a single cause such as “it failed because it was general-purpose and succeeded because it focused on design.”
Start your first experiment with evaluation criteria, not a leaderboard
DesignArena is set up to compare results in multiple formats, including websites, slides, mobile apps, games, images, logos, and SVGs. Supported formats and models may change, so check the currently available options on the start screen first.
The point of the exercise below is not to find the highest-ranked model, but to compare anonymous results all the way through using the same criteria.
- Prepare only inputs you can share. The service is for people aged 18 and over. Prompts, uploaded media, generated results, and votes may be collected and shared with third-party AI providers for evaluation, improvement, development, or potential training, so do not enter customer materials, personal information, or unreleased brand assets.
- Set three evaluation criteria first. For example, for a B2B dashboard, write down criteria that you will not change after seeing the results, such as “information hierarchy,” “error prevention,” and “WCAG contrast.”
- Log in from the DesignArena start screen. You need to log in to receive generated results. If this is your first visit, click Sign In before generating, then sign in with an existing account or create one as instructed.
- Select one format and start generating with the same prompt. You can narrow the first input like this: “Create a desktop onboarding dashboard for a B2B analytics tool. It needs left-hand navigation, data connection progress, and a card that highlights one next action. Do not use real customer names or data.”
- Do not guess the model names; make five choices using your preset criteria. Under the official procedure, you compare anonymous results from four models two at a time, then receive an ordering from first to fourth.
- Record the result only as a signal for reducing candidates. It is not wrong if the order differs from the public leaderboard. It becomes evidence for a product decision only after separately passing accessibility checks, usability tasks, or conversion experiments with your actual target audience.
Your first success criterion is not one pretty concept. It is enough to complete five choices while keeping the input and criteria consistent, and to explain why you chose first place using those three criteria. If the format you want is unavailable or a needed model is missing, practice only the evaluation structure with an available format and do not carry that ranking into a real project.
If you want to dig deeper
Methodology Review the matchup format that compares four models five times, the Bradley–Terry ratings, and the criteria for provisional results. designarena.ai
Design Arena creators raise $7.9 million to bring taste to AI models Covers the investment and, through founder interviews, the background behind the product selling evaluation data. techcrunch.com
PRIVACY NOTICE Explains why prompts and uploaded materials may be collected and shared. Read it before entering actual work materials. designarena.ai



