As part of our product studio, we created generalizable matching software called Matchbox. Hosts have now used it to match 100,000s of people at events in 100+ countries around the world. Usually, matches talk for 10–60 minutes after matching. We’ve obtained 20,000 labels from day-after match ratings. We still consider these to be “first-impression” ratings, but note they differ significantly from typical speed-dating datasets:
| Measure | Typical speed dating datasets | Matchbox dataset |
|---|---|---|
| Dataset size | n = tens to hundreds | n = 20,000 |
| Duration of conversation between matches | 2–5 minutes | 10–60 minutes |
| Time before labels were collected (more time = less bias from fleeting emotions) | 0–1 hours post-match | 16–24 hours post-match |
Note a few things about the following data:
You’ll see the correlations of just one question vs the final match score
Many more factors contribute to that final match score (usually each participant answers 20+ other questions, and the in-person interaction plays a significant role)
It’s therefore stunning that a strong effect from a single question is easily visible to the naked eye
N is so large that these are all extremely statistically significant

While the “similarity effect” exists, it’s weak. The clearer picture is in the 7×7 chart, which shows the protagonist’s rating of their match, organized by the protagonist’s own response to the question (x axis) vs their partner’s response (y axis). What we see: the more either person is comfortable with their partner being friends with an ex, the better for the match rating. Note this is true regardless of if it’s the protagonist or it’s the partner who’s comfy!
Lesson: the fewer neurotic people there are in any relationship, the better for relationship outcomes. This is an example of a “vertical trait”: one where it is universally better to have (or not have) the trait. This presents a challenge (albeit an overcomeable one) for matchmaking: neurotic people will consistently rank lower in searches/matches.

Here, again, there is a strong effect in the 7×7 grid: people who disagree with this statement typically rate other nontraditional people highly. This belies their underlying tolerance of unconventionality, which makes you more open-hearted to your matches.
But unlike last time, this time there is a massive similarity effect: matches simply aligned—anywhere—on this question outperform matches unaligned on this question by +22% on an absolute quality scale (vs +5%, give or take, on the previous question).

The truth is that opposites almost never attract. But in matters of polarity, like with sexual roles, opposites do attract. Here, we can quantify the exact extent to which that is the case in romantic matches: compatible responses to this question predict a +13% increase in match satisfaction after a 10-to-60 minute in-person conversation, when compared with a “mismatch of desires” at the extremes.
Please forgive noise in the 7×7 visual—we have “only” 2,632 relationships labeled here.