Before you chart a question or run a test on it, you need to know what kind of data it produced. A satisfaction rating, a brand choice and a monthly spend figure all arrive as columns in the same file, but they support different arithmetic. Picking the wrong statistic is one of the quietest ways to mislead a stakeholder, because the output still looks like a number.
This guide walks through the data types you meet in survey work, how each is analyzed and charted, and how they are coded for analysis. Two topics have their own deeper articles: what is ordinal data and what are dummy variables.
Key takeaways
- Survey data types come in four measurement levels, nominal, ordinal, interval and ratio, and the level decides which statistics and charts fit.
- Most survey rating scales are ordinal, so report the full distribution and top-box share alongside any mean rather than relying on the mean alone.
- A nominal question with k categories needs k - 1 dummy variables, because the last category is fully determined by the others.
- Multiple-response questions are stored as one 0/1 column per option, and age is ratio when entered as a number but ordinal when collected in bands.
- Qualitative data such as open-ended text becomes quantitative once it is coded into themes, which can then be counted and cross-tabulated by segment.
What is the difference between categorical and numerical data?
Every survey variable falls on one side of a first split.
Categorical data places each respondent in a group: brand owned, region, job function, “agree” or “disagree”. You can count how many fall in each group, but you cannot add two categories together in a meaningful way.
Numerical data records an amount: dollars spent, number of visits, minutes on a task, a 0–10 rating. You can add, subtract and average the values, within the limits set by the measurement level below.
Categorical data splits into nominal and ordinal. Numerical data splits into interval and ratio. That gives the four levels of measurement, first set out by the psychologist S. S. Stevens and still the standard framework for choosing analysis methods.
What are the four levels of measurement?
Each level keeps everything the previous level has and adds one property.
| Level | What it tells you | Survey example | Property added |
|---|---|---|---|
| Nominal | Which category | Brand most often bought, region | Names only, no order |
| Ordinal | Which category, and the order | Very dissatisfied to very satisfied, rank order | Order, but unknown spacing |
| Interval | Order and equal distances | Temperature in Celsius, a scale treated as equal-step | Equal spacing, no true zero |
| Ratio | Order, distance and a true zero | Spend, number of purchases, time | A zero that means “none” |
Nominal data
Nominal values are labels. “Brand A” is not more or less than “Brand B”. Frequencies, percentages and crosstabs are the tools, and the mode (the most common category) is the only meaningful average. Numeric codes such as 1 = North, 2 = South are only tags; their average means nothing.
Ordinal data
Ordinal values have a rank, but the gap between ranks is unknown. The distance from “very dissatisfied” to “dissatisfied” need not match the distance from “satisfied” to “very satisfied”. Most rating scales in surveys are ordinal. We cover how to handle them in detail in the next section and in the article on ordinal data.
Interval data
Interval data has equal steps but no true zero. Temperature in Celsius is the textbook case: the gap from 10 to 20 degrees equals the gap from 20 to 30, but 0 degrees is not “no temperature”, so 40 degrees is not “twice as hot” as 20. In survey work, many rating scales are treated as interval by convention, which is a judgment call rather than a fact about the scale.
Ratio data
Ratio data has equal steps and a true zero, so ratios are meaningful. A customer who spent $40 spent twice what a customer who spent $20 did. Spend, number of purchases, frequency and time are ratio data, and every statistic applies.
How should you handle ordinal data?
Most survey scales are ordinal, and two approaches coexist in practice. Reporting the distribution and top-box percentages respects the data. Reporting means trends well and is widely understood. Do both, state the base, and be cautious when the scale has few points or an uneven layout.
For significance testing, tests on proportions or rank-based tests are safer than t-tests on short scales. The ordinal data article covers the options and when each fits.
Here is a worked example. Ten respondents rate satisfaction from 1 (very dissatisfied) to 5 (very satisfied). The numbers are illustrative.
| Rating | Respondents | Contribution to the sum |
|---|---|---|
| 1 | 1 | 1 x 1 = 1 |
| 2 | 1 | 2 x 1 = 2 |
| 3 | 2 | 3 x 2 = 6 |
| 4 | 4 | 4 x 4 = 16 |
| 5 | 2 | 5 x 2 = 10 |
| Total | 10 | 35 |
- Mode: 4, the most frequent rating.
- Median: sort the ten ratings (1, 2, 3, 3, 4, 4, 4, 4, 5, 5). The 5th and 6th values are both 4, so the median is 4.
- Mean: 35 / 10 = 3.5.
- Top-2 box: ratings of 4 or 5 are 4 + 2 = 6 respondents, so 6 / 10 = 60%.
Each statistic answers a slightly different question. The mean of 3.5 hides the fact that six in ten are satisfied; the top-2 box says so directly. Reporting the distribution and one summary keeps both views visible.
What is the difference between discrete and continuous data?
This is a separate split from the four levels, and it applies to numerical data.
Discrete data takes countable, separate values: number of purchases last month, people in a household, store visits. There is nothing between 2 and 3 visits.
Continuous data can take any value in a range: time spent on a task, weight, or an exact amount of spend. In real surveys, continuous quantities are recorded to a limited precision (age in whole years, spend to the cent), so the distinction is about the thing measured, not the way it is typed in.
The practical consequence is chart choice and summary choice. Discrete counts with few values suit a column chart of frequencies. Continuous measures with many values suit a histogram, where respondents are grouped into bins, or a box plot. Both can be ratio data, so the same statistics apply.
Which survey questions produce which data types?
Mapping your questionnaire to data types before fieldwork saves a lot of rework. This table covers the common question formats.
| Question format | Data type | Typical statistic | Typical chart |
|---|---|---|---|
| Single choice (brand most often bought) | Nominal | Percent per category, mode | Bar chart |
| Multiple response (select all that apply) | Nominal; each option is a yes/no | Percent selecting each option | Sorted bar chart |
| Yes/no | Nominal (two categories) | Percent yes | Bar or single value |
| Rank order | Ordinal | Median rank, percent ranked first | Bar chart of percent ranked first |
| Age band or income band | Ordinal | Distribution, median band | Column chart |
| Likert agreement scale | Ordinal, often treated as interval | Distribution, top-box, mean | Stacked or diverging bar |
| 0–10 likelihood to recommend | Ordinal, often treated as interval | Distribution, mean, Net Promoter Score | Stacked column |
| Numeric entry (spend, number of visits) | Ratio; discrete or continuous | Mean, median, percentiles | Histogram, box plot |
| Open-ended text | Qualitative until coded | Share mentioning each theme | Bar chart of themes |
Two cases trip people up. Multiple-response questions are stored as one 0/1 column per option, not one column, so treat each option as its own binary variable. And age is ratio when entered as a number but ordinal when collected in bands, which means the question format you choose limits the analysis you can do later.
Which statistics and charts suit each data type?
Match the method to the level. The lower levels allow fewer operations, and the higher levels allow everything below them.
| Level | Averages | Spread | Common tests | Chart choices |
|---|---|---|---|---|
| Nominal | Mode | Frequencies | Chi-square test of independence | Bar chart; pie or donut for a few categories |
| Ordinal | Median, top-box | Percentiles, distribution | Rank-based tests, tests on proportions | Stacked or diverging bar |
| Interval | Mean | Standard deviation | t-test, ANOVA, Pearson correlation | Line over time, column, histogram |
| Ratio | Mean, median, geometric mean | Standard deviation, coefficient of variation | As interval, plus ratio comparisons | Histogram, box plot, scatter |
Use the tests as a guide, not a rule. The more your scale departs from equal steps, or the fewer points it has, the more cautious you should be about using interval methods on it.
A useful habit: write the data type next to each question in your analysis plan. When someone later asks for an average of a nominal code or a pie chart of an ordinal scale, the plan answers the question before the chart gets built.
What are dummy variables, and when do you need them?
Regression and driver models need numeric inputs, and a nominal question is not numeric. The fix is to recode the category into 0/1 indicator columns called dummy variables. The full explanation is in what are dummy variables; here is the essentials.
Take a Region question with four answers. It becomes three dummy variables, with the fourth category as the reference group. A respondent in the reference group has 0 in every dummy column.
| Respondent | Region | North | South | East |
|---|---|---|---|---|
| 1 | North | 1 | 0 | 0 |
| 2 | South | 0 | 1 | 0 |
| 3 | East | 0 | 0 | 1 |
| 4 | West (reference) | 0 | 0 | 0 |
A question with k categories needs k - 1 dummies. Including all four would make the columns redundant, because the fourth is fully determined by the other three, and most models cannot be estimated that way.
Each coefficient then reads as the difference from the reference. If the North coefficient in a model of satisfaction on a 0–10 scale were 0.4 (an illustrative figure), North respondents would score about 0.4 points higher than West respondents, with the other variables held constant. Choose a reference that makes the comparison meaningful, such as the largest group or the current product.
How are data types coded for analysis?
Coding turns raw answers into columns a tool can analyze. A sound codebook does the following.
- Assign a type to every variable. Label each as nominal, ordinal, scale (interval or ratio), or text, and keep that label with the data file.
- Store category codes with labels. Code 1 = “Very dissatisfied” through 5 = “Very satisfied” keeps the order of an ordinal scale, and the label travels with the number.
- Handle “don’t know” and “not applicable” separately. Code them outside the scale (for example 98 and 99) and declare them missing, so they do not enter the mean of a 1–5 rating.
- Split multiple-response questions into one 0/1 column per option. The mean of each column is the share who selected that option.
- Create dummy variables for modeling. Do this from the nominal variables you want in a regression, and record which category is the reference.
- Build scales on purpose. If you sum or average several rating items into one score, document which items went in and treat the result as interval only if you can justify it.
- Code open-ended text into themes and add the theme as a new categorical variable (see the next section).
The same variable can be coded more than one way, and that is fine. A 0–10 rating can be kept as 0–10 for a mean and also recoded to detractor, passive and promoter for a categorical view. Keep both versions and name them clearly.
What is the difference between qualitative and quantitative data?
Quantitative data are counts and measurements analyzed statistically. Qualitative data are words, images and observations analyzed for meaning. Surveys produce both: closed questions yield quantitative data, and open-ended questions yield qualitative data that becomes quantitative when coded into themes.
The most useful analysis usually combines them. The number says what moved; the verbatims say why. A satisfaction score that fell four points tells you something changed. Coded comments that show a rise in mentions of delivery delays tell you where to look.
Coded themes are categorical data like any other, so the rules above apply: count them, cross-tabulate them by segment, and be wary of reading an order into themes that have none.
Common mistakes when working with survey data types
- Averaging nominal codes. If region is coded 1 to 4, the mean is meaningless. Report the percentage in each region.
- Treating the mean as the whole story for a rating. A mean of 3.5 can come from a tight cluster around 3 and 4 or from a split between 1 and 5. Show the distribution.
- Calling a scale ratio because it has a zero. A 0–10 rating where 0 means “least likely” is not a true zero, so 8 is not “twice” 4.
- Using all k dummies in a model. This creates redundancy. Drop one as the reference.
- Binning continuous data too early. If you collect age in bands, you cannot go back to exact ages. When the analysis may need a number, ask for the number and band it afterward.
How do you do this in practice?
In a crosstab, the data type decides what each cell should show: percentages for nominal and ordinal questions, and means or medians for numeric ones. Halo Reports produces crosstabs and banner tables with weighting and significance testing, along with charts and interactive report pages, so each question can be shown the way its type allows.
To start, add a data-type column to your questionnaire or codebook this week, then check the statistic and chart assigned to each question against the tables above. Questions where the plan and the type disagree are where errors tend to hide. If you run wide, complex studies or global trackers, one governed platform for insights teams keeps those coding rules consistent across them.