Data Wrangling
Data wrangling is the process of cleaning, reshaping, and enriching raw data into a form suitable for analysis or modeling. It handles missing values, inconsistent formats, duplicates, and structural mismatches that prevent data from being used directly.
itData engineering and analytics | OpenSkills.info
Course pathWalk it in order
Look it upDip in anytime
Go furtherLeaves this page
Don't Panic
Don't Panic - Data Wrangling
Data Wrangling is the subject of this course. Data rarely arrives in the shape that analysis or a downstream system needs. Names vary.
The useful unit of work is a closed loop: clarify the goal and boundaries, gather the inputs the practice requires, make the decision or change, record evidence, and return with owners for the next cycle. Skipping any link leaves teams busy without durable results.
Tooling supports the loop; it does not replace it. Choose tools after the boundary and evidence model are clear. Comparing products without that model produces feature matrices that do not change how the work runs.
Common failure modes include undefined ownership, metrics that count activity instead of outcomes, and irreversible steps taken without a review path. Treat those as design defects in the practice, not as individual heroics to compensate later.
Operators should be able to explain which signals would change a decision this week. If no signal can change the plan, the practice has become ritual. Keep the feedback path short enough that evidence still influences the next cycle.
Name the owners for each stage of the loop before the work scales. Unowned stages become permanent exceptions. Record decisions with enough context that a future operator can tell why a tradeoff was accepted. Prefer fewer, sharper metrics that change behavior over broad dashboards that only describe activity after the fact.
Read the Intro for the core model. Use the Cheatsheet when you need the operating map. Updates tracks official guidance when this course configures an update source; otherwise the practice is settled without a live feed.
Where this skill leads
Relevant careers
See how this topic contributes to broader role-level skill maps.
Sources
- https://www.gov.uk/government/publications/the-government-data-quality-framework/the-government-data-quality-framework
Supports
- Data quality as fitness for purpose
- Completeness, uniqueness, consistency, timeliness, validity, and accuracy as distinct dimensions
- Quality trade-offs based on user needs
- Assessment throughout the data lifecycle
- Treating recurring quality problems at their source
- https://www.gov.uk/government/publications/the-government-data-quality-framework/the-government-data-quality-framework-guidance
Supports
- Quality action plans tied to user needs
- Baseline assessment, metrics, and automated checks
- Root-cause analysis and metadata practices
- Communication of cleaning, missingness, deduplication, and known limitations
- https://tidyr.tidyverse.org/articles/tidy-data.html
Supports
- Tidy data structure with variables in columns and observations in rows
- Reshaping data between wide and long representations
- Separating and combining variables during structural normalization
- https://pandas.pydata.org/docs/user_guide/missing_data.html
Supports
- Multiple missing-value representations across data types
- Detection, dropping, and filling of missing data
- The behavior and consequences of missing values in calculations
- https://pandas.pydata.org/docs/user_guide/indexing.html#duplicate-data
Supports
- Identification of duplicate rows with selected fields
- Duplicate removal and explicit first, last, or all survivor behavior
- https://pandas.pydata.org/docs/reference/api/pandas.merge.html
Supports
- Merge indicators for left-only, right-only, and matched rows
- Validation of one-to-one, one-to-many, and many-to-one key relationships
- Join behavior based on declared keys
- https://pandas.pydata.org/docs/user_guide/merging.html
Supports
- Merge, join, concatenate, and comparison operations
- Relational join semantics and key selection
- Cardinality and source-indicator checks for joins
- https://openrefine.org/docs/manual/facets
Supports
- Facets for summarizing values and selecting subsets
- Interactive inspection of textual, numeric, and timeline distributions
- https://openrefine.org/docs/technical-reference/clustering-in-depth
Supports
- Clustering as detection of possible alternate string representations
- Syntactic limits of clustering for semantic reconciliation
- Key-collision and nearest-neighbor clustering approaches
