A select-all-that-apply question has come back, and its percentages add up to well over 100. Before you filter or compare anything, it helps to know what the numbers mean. This article covers the arithmetic, the filtering logic, the base, and the chart, as part of our method for how to analyze survey data.
Key takeaways
- A multiple-response question lets respondents pick several answers, so percentages of respondents add up to more than 100% and should describe people, not responses.
- In the example, 200 respondents made 400 selections, so channels sum to 200% of respondents and 100% of responses; use percent of respondents for almost everything.
- Any-of logic keeps respondents who chose at least one named option, while all-of keeps only those who chose every option, so all-of filters shrink groups quickly.
- Name the base in every table, because the same 140 search-engine users are 70.0% of 200 respondents but 87.5% of the 160 who saw the question.
- Chart multiple-response results as a sorted horizontal bar chart rather than a pie chart, because the same person sits in several slices.
What is a multiple response question?
A multiple response question lets a respondent pick as many answers as apply. “Which of these channels do you use when researching a purchase?” with a list of checkboxes is the standard case. A single-choice question stores one answer per person. A multiple response question stores a set of answers per person.
In the data file this usually appears as one 0/1 column per option: 1 if the respondent ticked it, 0 if not. Analysis software groups those columns into one “multiple response set” so the question reads as a single table. Analyze each column alone and you lose the link between them, which is what makes filtering possible.
Why do multiple response percentages add up to more than 100%?
Because each respondent can be counted once per option they chose. There are two ways to turn counts into percentages, and they answer different questions.
Take an illustrative survey of 200 respondents. They chose these channels:
| Channel | Respondents who chose it | % of respondents (count / 200) | % of responses (count / 400) |
|---|---|---|---|
| Search engines | 140 | 70.0% | 35.0% |
| Retailer sites | 110 | 55.0% | 27.5% |
| Social media | 90 | 45.0% | 22.5% |
| Friends and family | 60 | 30.0% | 15.0% |
| Total | 400 responses | 200.0% | 100.0% |
The four counts sum to 400 responses from 200 people, an average of 2.0 selections each. Divide by the 200 respondents and the column sums to 200%. Divide by the 400 responses and it sums to 100%.
Use percent of respondents for almost everything. “70% of respondents use search engines” is a statement about people, which is what stakeholders expect. Percent of responses suits share-of-mentions questions and is easy to misread as a share of people.
How does filtering on a multiple response answer work?
Filtering means keeping only the respondents who meet a condition, then analyzing their other answers. With a multiple response question the condition needs a rule, because “chose Social media” and “chose Social media and Retailer sites” are different groups. The two rules are any-of and all-of.
What is any-of (OR) logic?
Any-of keeps a respondent who chose at least one of the options you name. Filtering to “Social media or Retailer sites” includes people who chose only one of them and people who chose both.
Respondents who chose both count once, not twice. Suppose 50 of the 200 chose both. Then 90 + 110 − 50 = 150 respondents meet the any-of condition.
What is all-of (AND) logic?
All-of keeps a respondent only if they chose every option you name. Filtering to “Social media and Retailer sites” gives the 50 who chose both. Each extra option you require can only shrink the group, so all-of filters run out of respondents quickly.
| Filter | Rule | Respondents | % of 200 |
|---|---|---|---|
| Social media or Retailer sites | Any-of | 150 | 75.0% |
| Social media and Retailer sites | All-of | 50 | 25.0% |
| Social media but not Retailer sites | Chose one, excluded the other | 40 | 20.0% |
The last row is 90 − 50 = 40. Many tools offer this “chose A but not B” condition, and it is easy to forget.
How do you filter another question by a multiple response answer?
The most common use is to split a different question by a multiple response answer. Suppose 120 of the 200 respondents said they would recommend the brand, and you want to compare people who use social media for research with those who do not.
| Group | Respondents | Would recommend | Rate |
|---|---|---|---|
| Chose Social media | 90 | 63 | 70.0% |
| Did not choose Social media | 110 | 57 | 51.8% |
| All respondents | 200 | 120 | 60.0% |
The two groups cover everyone, and 63 + 57 = 120, so the arithmetic closes. This is a crosstab with a multiple response option as the column. The numbers are illustrative, and a gap of 70% against 52% would deserve a significance test before you act on it.
One caution. A banner built from the options of a multiple response question has overlapping columns, since one respondent can appear in several of them. Treat tests between those columns with care, because standard tests assume the groups are independent.
What base should you use for percentages?
The base is the group you divide by, and you should name it in every table. For a multiple response question there are three common choices:
- All respondents. Simple, and right when everyone saw the question.
- Respondents who answered the question. Right when some people skipped it and you want to describe the people who responded.
- Respondents who were eligible. Right when skip logic showed the question to only part of the sample.
The choice changes the answer. In a variation of the example, suppose the same 140 chose search engines, but 40 of the 200 were screened out of the question. Divide by 160 and the figure is 87.5%. Divide by 200 and it is 70.0%.
Neither is wrong, but they describe different groups, and the table should say which.
When you filter, the base becomes the filtered group. A percentage “among those who chose Social media” has a base of 90, not 200. State that base beside the result, and be cautious with small ones: a base of 50 gives a wide margin of error.
How should you chart multiple response results?
Use a horizontal bar chart, with one bar per option, sorted from largest to smallest, showing the percent of respondents. Put the base in the chart note: “Base: all respondents (n = 200). Respondents could select more than one.” That note tells the reader why the bars do not sum to 100%.
Avoid a pie chart. A pie asserts that the slices are parts of one whole, and here they are not, because the same person sits in several slices. A stacked bar to 100% has the same problem. A bar chart makes no such claim, and readers can compare lengths along a common baseline.
Show an exclusive option such as “None of the above” apart from the rest. If there is an “Other (please specify)” option, code the write-ins first, or that bar will hide several distinct answers.
How do you do this in practice?
In a spreadsheet, you build 0/1 flags for each option, then use COUNTIFS or a pivot table for each filter, and track the base by hand. That is error-prone across many questions. In Halo Reports, you build crosstabs and banner tables with weighting and significance testing, then turn them into charts and interactive report pages, so the filter, the base, and the chart stay tied to the same data. If you would rather ask a question in plain language, you get the answer from your own research, with the source shown.
What are the common mistakes?
- Reading column totals as a check. A multiple response table that sums to 200% is not broken. A single-choice table that does is.
- Mixing up any-of and all-of. “Social media or Retailer sites” and “Social media and Retailer sites” differ by 100 respondents in the example above. Write the rule next to the filter.
- Dropping the base. A percentage with no base cannot be reproduced and is easy to misquote.
- Using a pie chart. It implies a whole that does not exist.
- Treating overlapping groups as independent. Check whether respondents sit in both groups before testing the difference.
- Ignoring the exclusive option. Someone who picked “None of the above” and another option is a data-quality flag.
What should you do next?
Pick one multiple response question in your current study. Compute percent of respondents and the average number of selections, run one any-of and one all-of filter, and write down each base. Chart the result as a sorted bar, then cross the question against your key segments. Insights teams running wide, complex studies can do this in one governed platform for research teams.