Survey Design and Analysis
A survey is a method for measuring a population by asking a standard set of questions to a sample of its members and generalizing the answers to the whole group. It exists because you usually cannot ask everyone, and because how you select people and word questions changes the result. Survey design and analysis is the practice of making those choices deliberately so the final numbers carry a defensible margin of error rather than hidden bias.
itData engineering and analytics | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic: Survey Design and Analysis
A survey measures a whole population by asking a fixed set of questions to a sample of its members and generalizing the answers. You reach for one when the fact you need lives only in people's heads: opinions, intentions, what they remember doing, how satisfied they are. If the answer already exists in a log or a billing system, use that. A survey is the instrument for questions only a person can answer, and the catch is that people answer imperfectly.
The work splits into two halves that break separately. Design is who you ask and how you word the questions. Analysis is turning the returned answers into population estimates. A flawless questionnaire sent to the wrong sample gives you precise nonsense; a perfect sample answering a loaded questionnaire gives you representative bias. Neither half rescues the other.
The idea that ties it together is total survey error, the accounting system for every way an estimate can differ from the truth. One branch is representation: does your sampling frame, the list or method you draw people from, actually cover the population, do responders differ from the people who ignored you, does the weighting distort things. The other branch is measurement: does the question capture the concept, does the respondent understand and answer honestly. Design is a series of trades among these, against cost.
Here is the part that surprises people. The margin of error you see quoted, the plus-or-minus figure, covers only sampling error, the fact that you asked a sample and not everyone. Coverage gaps, nonresponse, and question wording are usually bigger, and none of them are in that number. A giant opt-in panel can be more wrong than a small probability sample, one where every person has a known chance of selection, while its tighter margin of error hides the problem.
Two smaller traps. A number for a subgroup carries the margin of that subgroup's size, not the whole sample's. And the gap between two groups has a wider margin than either group alone, so test a difference before you describe it.
On the questionnaire: ask one thing per question, avoid loaded words, and be wary of agree-or-disagree scales, because some people agree with anything, an effect called acquiescence. Expect satisficing, where a respondent who does not want to work picks the first acceptable option, the midpoint, or "don't know".
Read the Intro for the full vocabulary and the sampling designs. Slides show how the error sources connect. The Cheatsheet has the formulas and the sample-size table. The Practice Reference walks through writing questions, sizing a sample, computing a response rate, and weighting. Field Notes is blunter: response rates keep falling, nobody has solved nonresponse, and chasing the rate is optimizing the wrong number.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://www.pewresearch.org/u-s-survey-methodology/
Supports
- Total survey error framework and its representation and measurement branches
- Probability-based panel recruitment and online administration
- Weighting and calibration to population benchmarks
- Method disclosure alongside published estimates
- https://www.pewresearch.org/writing-survey-questions/
Supports
- Open-ended versus closed-ended format effects (economy example, 35 vs 58 percent)
- Question wording effects on responses
- Double-barreled questions and one concept per item
- Acquiescence bias in agree-disagree formats and its education gradient
- Question order and context effects, general before specific
- Randomizing nominal option order to spread primacy and recency
- Split-ballot A/B testing of question versions
- Rewording an item breaks a time series; run versions in parallel
- https://www.pewresearch.org/methods/2018/01/26/variability-of-survey-estimates/
Supports
- Margin of error as the half-width of a 95 percent confidence interval
- Margin of error near 3 points at n=1,000 and near 2 points at n=2,000
- Design effect and effective sample size (nominal n divided by design effect)
- Subgroup estimates carry the subgroup's margin of error
- Modeled margin of error for opt-in samples
- https://www.pewresearch.org/methods/2018/01/26/how-different-weighting-methods-work/
Supports
- Raking (iterative proportional fitting) needs only marginal totals
- Matching discards unmatched cases; propensity weighting keeps all cases
- Post-stratification and sequential weighting approaches
- Choice of calibration variables matters more than the algorithm for opt-in samples
- https://aapor.org/wp-content/uploads/2023/05/Standards-Definitions-10th-edition.pdf
Supports
- Disposition codes and the RR1 through RR6 response-rate formulas
- Cooperation, contact, and refusal rates as separate metrics
- Response rate alone does not quantify nonresponse bias
- Probability sample definition and known selection probabilities
- https://aapor.org/standards-and-ethics/best-practices/
Supports
- Choosing whether a survey is the right method and which mode to use
- Mutually exclusive and exhaustive response options
- Interviewer training and disclosure practices
- https://nces.ed.gov/fcsm/pdf/OMB_Standards_Guidelines_Statistical_Surveys.pdf
Supports
- Federal requirements for frame construction and coverage evaluation
- Pretesting requirements including cognitive interviewing
- Nonresponse bias analysis and documentation standards
- https://www.pewresearch.org/methods/2016/05/02/evaluating-online-nonprobability-surveys/
Supports
- Accuracy of online opt-in samples varies widely by vendor
- Opt-in error is largest for estimates about subgroups
- https://aapor.org/wp-content/uploads/2022/11/AAPOR-2016-Election-Polling-Report.pdf
Supports
- National 2016 popular-vote polls among the most accurate since 1936
- State polls underestimated one candidate, partly from unweighted education
- Late deciding as a contributor to state-level polling error
- https://academic.oup.com/poq/article-abstract/72/2/167/1920564
Supports
- Meta-analysis of 59 studies finds little correlation between response rate and nonresponse bias
- https://legacy.voteview.com/pdf/Likert_1932.pdf
Supports
- Likert's 1932 multi-item attitude measurement technique
- https://doi.org/10.1111/j.2397-2335.1934.tb04184.x
Supports
- Neyman's 1934 formal basis for probability sampling and confidence intervals
- Stratified random sampling versus purposive selection
- https://academic.oup.com/poq/article-abstract/52/1/125/1878544
Supports
- 1936 Literary Digest failure from a biased frame and self-selected response
- https://aapor.org/about-us/history/
Supports
- AAPOR formally organized on September 4, 1947
- 1946 founding conference in Central City, Colorado
- https://link.springer.com/chapter/10.1007/978-0-387-77956-0_1
Supports
- 1948 pollsters predicted the wrong US presidential winner
- Quota sampling and early cessation of interviewing as causes
- Social Science Research Council committee review of the 1948 polls
- https://sesrc.wsu.edu/about/total-design-method/
Supports
- Dillman's 1978 Total Design Method for mail and telephone surveys
- Standardized contact sequence raising mail response rates to 60 to 70 percent
- https://onlinelibrary.wiley.com/doi/book/10.1002/0471725277
Supports
- Groves's 1989 organization of survey error into a mean-squared-error structure
- Trading error sources against survey cost
- https://academic.oup.com/poq/article/74/5/849/1817502
Supports
- History and definition of the total survey error framework
- https://aapor.org/standards-and-ethics/standard-definitions/
Supports
- AAPOR first issued Standard Definitions in 1998
- https://aapor.org/wp-content/uploads/2023/01/2010AAPORCellPhoneTFReport.pdf
Supports
- 2010 guidance on adding cell phone frames to RDD telephone surveys
- Dual-frame overlap and higher cost of mobile interviewing
- https://www.pewresearch.org/methods/2015/04/08/building-pew-research-centers-american-trends-panel/
Supports
- Pew recruited the American Trends Panel via RDD in 2014, then moved to web
- https://aapor.org/wp-content/uploads/2022/11/AAPOR-Task-Force-on-2020-Pre-Election-Polling_Report-FNL.pdf
Supports
- Review of more than 2,800 polls from the 2020 US general election
- No single identified cause for the 2020 polling error
- Electorate composition assumptions ruled out as the primary cause
- https://onlinelibrary.wiley.com/doi/abs/10.1002/acp.2350050305
Supports
- Satisficing response strategies under cognitive burden
- Choosing the first acceptable option, agreeing, midpoint selection, and "don't know"
- Straight-lining on grids as an observable of satisficing
- https://www.qualtrics.com/
Supports
- General-purpose survey research platform landscape entry
- https://www.surveymonkey.com/
Supports
- Self-serve questionnaire builder with an opt-in audience panel landscape entry
- https://www.limesurvey.org/
Supports
- Open-source self-hosted survey platform landscape entry
- https://projectredcap.org/
Supports
- Consortium-licensed research data collection platform landscape entry
- https://www.surveycto.com/
Supports
- Enumerator-administered offline field data collection landscape entry
- https://sawtoothsoftware.com/
Supports
- Choice-based conjoint and MaxDiff instrument landscape entry
- https://amerispeak.norc.org/
Supports
- Probability-based US household panel landscape entry
- https://www.prolific.com/
Supports
- Pay-per-response nonprobability participant pool landscape entry
