Filtering a Database to Find Comparable Companies
Using data · 9 min read ·
How to use filters in a startup database to build a fair comparison set: choose criteria, combine them, check data quality and avoid selection traps.
A startup database is useful when it helps you answer a question: who else is doing this, how do they compare and what can I learn from them? Filters are the tool for narrowing a long list to a short one. Used well, they make a fair comparison set in minutes. Used carelessly, they produce a set that looks authoritative and is misleading.
This guide explains how to choose and combine filters, what to check about the data behind them and how to avoid common selection traps.
Start with the question
Before you touch a filter, write down what you want to know.
- "Which companies sell scheduling software to small clinics in the UK?"
- "How big are companies at about our stage in this field?"
- "What do comparable companies charge, and what do they offer?"
- "Who might be a competitor, a partner or a customer?"
The question tells you which dimensions matter. A pricing comparison needs similar products and customers. A hiring benchmark needs similar size and location. Without a question, you may choose filters that are convenient rather than relevant.
Know what you can filter on
Most databases offer several kinds of filter.
What the company does. Category, tags, product type, business model, target customer.
Where. Country, city, region, remote status.
When. Founding year, launch date.
How big. Team size, funding stage, funding total.
How it is doing. Status (active, acquired, closed), recent news, update dates.
How well documented. Verified status, completeness of the profile.
Each filter narrows the set along one dimension. The art is choosing a combination that captures what "comparable" means for your question.
What "comparable" means
Comparability is about the dimensions that drive the thing you are studying.
- For product and pricing comparisons, category, customer type and business model matter most. Size and age matter less.
- For benchmarks of team size or growth, stage, age and sector matter most.
- For market maps, category and region.
- For competitor research, the customer and the problem.
Wikipedia describes startups as ventures seeking a scalable business model under uncertainty. That means two companies of the same age and size may be at very different points in their search. A comparison set needs to reflect that: filter on stage as well as age where you can, and look at what the companies actually say they do.
A step-by-step method
1. Start broad on the core dimension
Begin with the category or topic that defines your question. Look at the number of results. If there are thousands, you need more filters. If there are three, you have probably been too narrow.
2. Add the most important second dimension
Often this is the customer or region. Add it and watch the count.
3. Add stage or size,** if relevant
These help you compare like with like. Use ranges, not exact values.
4. Check what you lost
Each filter excludes companies. Look at a few that were removed and ask whether they should have been. Sometimes a company is excluded because a field is blank, not because it fails the criterion.
5. Read the profiles
A filter gives you a list. A comparison needs understanding. Open each profile in your set and read the description, the dates and the sources. Remove companies that do not fit, and note why.
6. Record your criteria
Write down the filters you used, the date and the number of results. This makes your set reproducible and lets you update it later.
Understand how blanks behave
Databases differ in how they treat missing values. If you filter for team size 6 to 10, are companies with no team size listed excluded or ignored? Most exclude them, which means that the more fields you filter on, the more companies disappear because of missing data, not because they fail your criteria.
To check, run the filter in two ways: with and without the field, and compare the lists. If there is a large difference, consider whether the missing companies might belong in your set. You may need to review them by hand.
A good database shows unknown values clearly and lets you filter for "not stated" as well as for specific values.
Mind the data quality
Filters are only as good as the data. Consider:
- Freshness. Is the team size from last month or three years ago? Look for last-checked dates.
- Definitions. What does "team size" count? What does "stage" mean? Read the explanation page.
- Verification. Is the founding date verified against a register or company-reported?
- Consistency. Are categories assigned by a controlled list or free text?
- Coverage. Is the database likely to include companies like the ones you want? A database biased toward certain regions or sectors will give a biased set.
Where possible, filter on fields that are verified and fresh, and treat company-reported fields as approximate.
Beware of selection traps
Survivorship
Databases often over-represent companies that still exist. If you compare funding amounts among listed companies, you may miss the many that raised money and closed. Wikipedia notes the high failure rates among startups, so a set of survivors is not a typical set.
Self-selection
Companies that create their own profiles may differ from those that do not. They tend to be more active, more promotional and more confident. Do not assume that a database represents the whole market.
Over-narrowing
Adding filter after filter can leave you with a set so small that it reflects chance. If you end up with five companies, ask whether you have enough to draw any conclusion.
Confirmation bias
It is tempting to adjust filters until the set confirms what you already believe. Fix your criteria in advance, and write them down before looking at results.
Goodhart in the data
Goodhart's observation, that a measure becomes a poor measure when it becomes a target, applies to profiles. If companies know that a certain field is used for ranking or filtering, they may report it generously. Be wary of fields that are easy to inflate, such as team size, and prefer verified ones.
Combine filters thoughtfully
Filters can be combined in two ways.
- AND: the company must meet all criteria. This narrows the set.
- OR: the company must meet any criterion. This widens it.
Most filter interfaces use AND across different fields and OR within a field, such as "category is A or B, and country is X". Check how your database behaves.
Consider using multiple passes: a broad set for context, a narrower core set for detailed comparison. You might present the core as "closest comparables" and the broad as "wider field".
Use the comparison
Once you have a set, put it to work.
- Tabulate key facts: founding date, location, team size, stage, category, funding events, with dates and sources.
- Read descriptions for positioning and language.
- Look at product details: features, pricing model, target customers.
- Note patterns and outliers. Are most companies in the set similar in age? Which are exceptions and why?
- Record gaps. What could you not find?
- Share your method along with the results, so others can judge them.
A worked example
A founder of a scheduling tool for physiotherapists wants to benchmark team size and funding among comparable companies. Her question: "At about a year old, how big are companies like ours?"
She filters by category (healthcare scheduling), then region (United Kingdom and Ireland), then founding year (the last three years). The database returns forty-two companies. She adds team size (2 to 10) and finds twenty-six. She runs the same search without the team size filter and finds that sixteen companies had no team size listed. She opens those sixteen and finds that eight have a public team page she can use.
She reads the profiles of the remaining set and removes four that are consultancies, not product companies. She tabulates the rest, with dates. The typical team size in her set is between four and eight, and about a third have announced seed funding. She notes that the data is mostly company-reported and that the set probably over-represents active, promotional companies. She records her filters, the date and her caveats, and saves them to update in six months.
A checklist
- Question written first
- Core dimension, then second, then size or stage
- Count watched at each step
- Blanks checked with and without each filter
- Profiles read, misfits removed
- Data quality assessed
- Selection traps considered
- Criteria recorded
The categories and search pages here let you filter and browse, the leaderboards page shows how rankings are labelled and the submit page is where a company adds its own profile.
Sharing and updating your set
A comparison set is most useful when others can understand and reuse it. Save your filters as a short note: the question, the criteria, the date, the number of companies and the caveats. If you share the set with colleagues, include the list with each company's key facts and the dates they were checked. Plan to refresh it when the database updates, because new companies appear and old ones close. Over time, a maintained comparison set becomes a small, trusted asset for the team, rather than a one-off spreadsheet that nobody dares to rely on.
Frequently asked questions
Should I rely on one database? Use more than one if you can, and compare.
Can I export results? That depends on the database. Check its terms.
How often should I rerun the filters? When your question changes, or every six months for an ongoing comparison.
Questions and answers
- What makes two companies comparable?
- Similar products and customers, a similar stage and size and a similar market. Which factors matter most depends on your question.
- How many filters should I use?
- Enough to narrow to a meaningful set, but not so many that only a handful of companies remain.
- Why do results change when I add a filter?
- Missing or unstated values may exclude companies. Check how the database treats blanks.
- How big should a comparison set be?
- Large enough that no single company dominates, and small enough that you can read the profiles. Ten to thirty is common.