Skip to content

Survey data types: nominal, ordinal, interval, dummy variables, and qualitative vs. quantitative

Survey data comes in four measurement levels: nominal (categories with no order, such as brand chosen), ordinal (ordered categories, such as a satisfaction scale), interval (equal steps without a true zero) and ratio (equal steps with a true zero, such as spend). The level determines which statistics and charts fit. Dummy variables recode categories as 0/1 indicators for models.

Tracey Stuart

Global VP Analysis & Insights · Updated July 10, 2025

Black-and-white photo of monitors showing line charts and candlestick charts

Before you chart a question or run a test on it, you need to know what kind of data it produced. A satisfaction rating, a brand choice and a monthly spend figure all arrive as columns in the same file, but they support different arithmetic. Picking the wrong statistic is one of the quietest ways to mislead a stakeholder, because the output still looks like a number.

This guide walks through the data types you meet in survey work, how each is analyzed and charted, and how they are coded for analysis. Two topics have their own deeper articles: what is ordinal data and what are dummy variables.

Key takeaways

  • Survey data types come in four measurement levels, nominal, ordinal, interval and ratio, and the level decides which statistics and charts fit.
  • Most survey rating scales are ordinal, so report the full distribution and top-box share alongside any mean rather than relying on the mean alone.
  • A nominal question with k categories needs k - 1 dummy variables, because the last category is fully determined by the others.
  • Multiple-response questions are stored as one 0/1 column per option, and age is ratio when entered as a number but ordinal when collected in bands.
  • Qualitative data such as open-ended text becomes quantitative once it is coded into themes, which can then be counted and cross-tabulated by segment.

What is the difference between categorical and numerical data?

Every survey variable falls on one side of a first split.

Categorical data places each respondent in a group: brand owned, region, job function, “agree” or “disagree”. You can count how many fall in each group, but you cannot add two categories together in a meaningful way.

Numerical data records an amount: dollars spent, number of visits, minutes on a task, a 0–10 rating. You can add, subtract and average the values, within the limits set by the measurement level below.

Categorical data splits into nominal and ordinal. Numerical data splits into interval and ratio. That gives the four levels of measurement, first set out by the psychologist S. S. Stevens and still the standard framework for choosing analysis methods.

What are the four levels of measurement?

Each level keeps everything the previous level has and adds one property.

LevelWhat it tells youSurvey exampleProperty added
NominalWhich categoryBrand most often bought, regionNames only, no order
OrdinalWhich category, and the orderVery dissatisfied to very satisfied, rank orderOrder, but unknown spacing
IntervalOrder and equal distancesTemperature in Celsius, a scale treated as equal-stepEqual spacing, no true zero
RatioOrder, distance and a true zeroSpend, number of purchases, timeA zero that means “none”

Nominal data

Nominal values are labels. “Brand A” is not more or less than “Brand B”. Frequencies, percentages and crosstabs are the tools, and the mode (the most common category) is the only meaningful average. Numeric codes such as 1 = North, 2 = South are only tags; their average means nothing.

Ordinal data

Ordinal values have a rank, but the gap between ranks is unknown. The distance from “very dissatisfied” to “dissatisfied” need not match the distance from “satisfied” to “very satisfied”. Most rating scales in surveys are ordinal. We cover how to handle them in detail in the next section and in the article on ordinal data.

Interval data

Interval data has equal steps but no true zero. Temperature in Celsius is the textbook case: the gap from 10 to 20 degrees equals the gap from 20 to 30, but 0 degrees is not “no temperature”, so 40 degrees is not “twice as hot” as 20. In survey work, many rating scales are treated as interval by convention, which is a judgment call rather than a fact about the scale.

Ratio data

Ratio data has equal steps and a true zero, so ratios are meaningful. A customer who spent $40 spent twice what a customer who spent $20 did. Spend, number of purchases, frequency and time are ratio data, and every statistic applies.

How should you handle ordinal data?

Most survey scales are ordinal, and two approaches coexist in practice. Reporting the distribution and top-box percentages respects the data. Reporting means trends well and is widely understood. Do both, state the base, and be cautious when the scale has few points or an uneven layout.

For significance testing, tests on proportions or rank-based tests are safer than t-tests on short scales. The ordinal data article covers the options and when each fits.

Here is a worked example. Ten respondents rate satisfaction from 1 (very dissatisfied) to 5 (very satisfied). The numbers are illustrative.

RatingRespondentsContribution to the sum
111 x 1 = 1
212 x 1 = 2
323 x 2 = 6
444 x 4 = 16
525 x 2 = 10
Total1035
  • Mode: 4, the most frequent rating.
  • Median: sort the ten ratings (1, 2, 3, 3, 4, 4, 4, 4, 5, 5). The 5th and 6th values are both 4, so the median is 4.
  • Mean: 35 / 10 = 3.5.
  • Top-2 box: ratings of 4 or 5 are 4 + 2 = 6 respondents, so 6 / 10 = 60%.

Each statistic answers a slightly different question. The mean of 3.5 hides the fact that six in ten are satisfied; the top-2 box says so directly. Reporting the distribution and one summary keeps both views visible.

What is the difference between discrete and continuous data?

This is a separate split from the four levels, and it applies to numerical data.

Discrete data takes countable, separate values: number of purchases last month, people in a household, store visits. There is nothing between 2 and 3 visits.

Continuous data can take any value in a range: time spent on a task, weight, or an exact amount of spend. In real surveys, continuous quantities are recorded to a limited precision (age in whole years, spend to the cent), so the distinction is about the thing measured, not the way it is typed in.

The practical consequence is chart choice and summary choice. Discrete counts with few values suit a column chart of frequencies. Continuous measures with many values suit a histogram, where respondents are grouped into bins, or a box plot. Both can be ratio data, so the same statistics apply.

Which survey questions produce which data types?

Mapping your questionnaire to data types before fieldwork saves a lot of rework. This table covers the common question formats.

Question formatData typeTypical statisticTypical chart
Single choice (brand most often bought)NominalPercent per category, modeBar chart
Multiple response (select all that apply)Nominal; each option is a yes/noPercent selecting each optionSorted bar chart
Yes/noNominal (two categories)Percent yesBar or single value
Rank orderOrdinalMedian rank, percent ranked firstBar chart of percent ranked first
Age band or income bandOrdinalDistribution, median bandColumn chart
Likert agreement scaleOrdinal, often treated as intervalDistribution, top-box, meanStacked or diverging bar
0–10 likelihood to recommendOrdinal, often treated as intervalDistribution, mean, Net Promoter ScoreStacked column
Numeric entry (spend, number of visits)Ratio; discrete or continuousMean, median, percentilesHistogram, box plot
Open-ended textQualitative until codedShare mentioning each themeBar chart of themes

Two cases trip people up. Multiple-response questions are stored as one 0/1 column per option, not one column, so treat each option as its own binary variable. And age is ratio when entered as a number but ordinal when collected in bands, which means the question format you choose limits the analysis you can do later.

Which statistics and charts suit each data type?

Match the method to the level. The lower levels allow fewer operations, and the higher levels allow everything below them.

LevelAveragesSpreadCommon testsChart choices
NominalModeFrequenciesChi-square test of independenceBar chart; pie or donut for a few categories
OrdinalMedian, top-boxPercentiles, distributionRank-based tests, tests on proportionsStacked or diverging bar
IntervalMeanStandard deviationt-test, ANOVA, Pearson correlationLine over time, column, histogram
RatioMean, median, geometric meanStandard deviation, coefficient of variationAs interval, plus ratio comparisonsHistogram, box plot, scatter

Use the tests as a guide, not a rule. The more your scale departs from equal steps, or the fewer points it has, the more cautious you should be about using interval methods on it.

A useful habit: write the data type next to each question in your analysis plan. When someone later asks for an average of a nominal code or a pie chart of an ordinal scale, the plan answers the question before the chart gets built.

What are dummy variables, and when do you need them?

Regression and driver models need numeric inputs, and a nominal question is not numeric. The fix is to recode the category into 0/1 indicator columns called dummy variables. The full explanation is in what are dummy variables; here is the essentials.

Take a Region question with four answers. It becomes three dummy variables, with the fourth category as the reference group. A respondent in the reference group has 0 in every dummy column.

RespondentRegionNorthSouthEast
1North100
2South010
3East001
4West (reference)000

A question with k categories needs k - 1 dummies. Including all four would make the columns redundant, because the fourth is fully determined by the other three, and most models cannot be estimated that way.

Each coefficient then reads as the difference from the reference. If the North coefficient in a model of satisfaction on a 0–10 scale were 0.4 (an illustrative figure), North respondents would score about 0.4 points higher than West respondents, with the other variables held constant. Choose a reference that makes the comparison meaningful, such as the largest group or the current product.

How are data types coded for analysis?

Coding turns raw answers into columns a tool can analyze. A sound codebook does the following.

  1. Assign a type to every variable. Label each as nominal, ordinal, scale (interval or ratio), or text, and keep that label with the data file.
  2. Store category codes with labels. Code 1 = “Very dissatisfied” through 5 = “Very satisfied” keeps the order of an ordinal scale, and the label travels with the number.
  3. Handle “don’t know” and “not applicable” separately. Code them outside the scale (for example 98 and 99) and declare them missing, so they do not enter the mean of a 1–5 rating.
  4. Split multiple-response questions into one 0/1 column per option. The mean of each column is the share who selected that option.
  5. Create dummy variables for modeling. Do this from the nominal variables you want in a regression, and record which category is the reference.
  6. Build scales on purpose. If you sum or average several rating items into one score, document which items went in and treat the result as interval only if you can justify it.
  7. Code open-ended text into themes and add the theme as a new categorical variable (see the next section).

The same variable can be coded more than one way, and that is fine. A 0–10 rating can be kept as 0–10 for a mean and also recoded to detractor, passive and promoter for a categorical view. Keep both versions and name them clearly.

What is the difference between qualitative and quantitative data?

Quantitative data are counts and measurements analyzed statistically. Qualitative data are words, images and observations analyzed for meaning. Surveys produce both: closed questions yield quantitative data, and open-ended questions yield qualitative data that becomes quantitative when coded into themes.

The most useful analysis usually combines them. The number says what moved; the verbatims say why. A satisfaction score that fell four points tells you something changed. Coded comments that show a rise in mentions of delivery delays tell you where to look.

Coded themes are categorical data like any other, so the rules above apply: count them, cross-tabulate them by segment, and be wary of reading an order into themes that have none.

Common mistakes when working with survey data types

  • Averaging nominal codes. If region is coded 1 to 4, the mean is meaningless. Report the percentage in each region.
  • Treating the mean as the whole story for a rating. A mean of 3.5 can come from a tight cluster around 3 and 4 or from a split between 1 and 5. Show the distribution.
  • Calling a scale ratio because it has a zero. A 0–10 rating where 0 means “least likely” is not a true zero, so 8 is not “twice” 4.
  • Using all k dummies in a model. This creates redundancy. Drop one as the reference.
  • Binning continuous data too early. If you collect age in bands, you cannot go back to exact ages. When the analysis may need a number, ask for the number and band it afterward.

How do you do this in practice?

In a crosstab, the data type decides what each cell should show: percentages for nominal and ordinal questions, and means or medians for numeric ones. Halo Reports produces crosstabs and banner tables with weighting and significance testing, along with charts and interactive report pages, so each question can be shown the way its type allows.

To start, add a data-type column to your questionnaire or codebook this week, then check the statistic and chart assigned to each question against the tables above. Questions where the plan and the type disagree are where errors tend to hide. If you run wide, complex studies or global trackers, one governed platform for insights teams keeps those coding rules consistent across them.

Frequently asked questions

Can you take the mean of an ordinal scale?

It is common practice for rating scales with roughly equal steps, and useful for tracking. Report the distribution alongside it, and use top-box or median measures when the scale is clearly uneven.

What is a dummy variable?

A 0/1 variable that indicates membership in one category, created so a categorical question can be used in a regression or driver model. A question with k categories needs k - 1 dummies.

Is verbatim text qualitative or quantitative data?

Qualitative until it is coded. Once responses are coded into themes and counted, the theme frequencies are quantitative and can be cross-tabulated.

Is a yes/no question nominal or ordinal?

Nominal. Yes and no are two unordered categories. Coded as 1 and 0, the average of the column equals the share who said yes, which is why binary questions work so well in models.

See how insights teams do this in mTab.

Bring one tracker or study. We'll build the tables and the report before we meet.

Book a demo