What Is a Spatial Data Quality Checklist?
A spatial data quality checklist is a repeatable set of checks for structure, CRS, geometry, attributes, identifiers, time, provenance, licensing and fitness for a specific use.
Provide a practical checklist while emphasising that data quality is fitness for purpose rather than a universal pass/fail score.
What Is a Spatial Data Quality Checklist?
A spatial data quality checklist is a repeatable way to inspect whether a geographic dataset is structurally sound, interpretable and fit for the task you intend to perform. It should not reduce quality to “valid/invalid”, because a dataset can pass every structural check and still be too old, too coarse, too incomplete or legally unsuitable for the intended use.
The checklist is most useful when each check is tied to a decision.
1. Identity and scope
Before testing geometry, establish what the dataset claims to contain.
Check:
title and description;
feature type or phenomenon represented;
geographic extent;
expected feature count or coverage;
responsible organisation or source;
date or period represented;
version or release identifier.
A dataset called “roads” may include only major roads, exclude private roads or represent a historical network, and its quality cannot be judged against assumptions it never claimed to satisfy.
2. Coordinate reference system and extent
Confirm the CRS is known and plausible for the coordinate values.
Inspect:
CRS identifier and datum/reference frame;
units;
axis interpretation where relevant;
bounding box;
outlier coordinates;
whether transformation is required for the intended analysis.
Do not treat a layer that happens to plot near the expected location as proof that the CRS is correct.
3. Geometry structure and validity
Check that geometry types match the dataset's purpose and that invalid structures are identified.
Useful checks include:
null and empty geometry count;
expected Point/LineString/Polygon/Multi* types;
self-intersections;
ring or topology errors where relevant;
extreme vertex counts;
duplicate geometries;
multipart features that may need explanation.
Structural validity is not positional accuracy, so a perfectly valid polygon can still represent the wrong boundary.
4. Attributes and domains
Inspect whether the fields are usable and internally consistent.
Check:
expected columns;
data types;
null rates;
allowed categorical values;
impossible numeric ranges;
date formats;
units;
character encoding;
field descriptions.
For example, a temperature field without a unit or a code field whose leading zeros were stripped can undermine downstream analysis even when every row is present.
5. Identifiers and joins
Determine what uniquely identifies a feature or observation.
Check:
uniqueness of the stated primary identifier;
missing identifiers;
code-system namespace and version;
leading zeros and text/numeric coercion;
one-to-many relationships;
unmatched records in important joins.
Stable identifiers are especially important when geography changes over time.
6. Temporal quality
Ask what date the dataset represents rather than looking only at when the file was downloaded.
Record:
observation date or validity period;
publication date;
last update date;
retrieval date;
boundary or code-system version;
refresh frequency where relevant.
A recently downloaded dataset may still describe conditions from years earlier.
7. Provenance and transformation history
Document where the data came from and what has happened to it.
At minimum, preserve:
original source;
source URL or identifier;
source version/date;
transformations performed;
joins or enrichments;
geometry repairs;
filters or exclusions;
derived fields;
software or scripts when reproducibility matters.
What Is Spatial Data Provenance? explains the distinction between source, process and claim-level provenance.
8. Licence, terms and publication constraints
Before reuse or publication, record the licence and any relevant attribution or share-alike obligations.
Also check terms of access, privacy, confidentiality and contractual restrictions, because a dataset being publicly downloadable does not automatically make every republication safe or permitted.
9. Fitness for the intended use
This is the most important check.
State the operation or decision the dataset will support, and then ask whether its quality is sufficient for that purpose.
A generalised national boundary may be excellent for a world map and unsuitable for parcel-level area calculations. A geocoded point with 500-metre uncertainty may be good enough for regional aggregation and inadequate for assigning a household to one side of a small administrative boundary.
Quality is therefore relative to the required resolution, accuracy, completeness, time and legal context.
A compact checklist
Dimension · Core question
Identity — What exactly does this dataset represent?
CRS — Are coordinates correctly interpreted?
Geometry — Are features structurally usable?
Attributes — Are values typed, encoded and documented correctly?
Identifiers — Can records be joined and tracked reliably?
Time — Which real-world period/version does the data represent?
Provenance — Can the source and transformations be traced?
Licence/privacy — Can the data be used and published as intended?
Fitness — Is the quality adequate for this decision?
A checklist should create evidence, not ceremony
The purpose is not to tick nine boxes and declare the dataset “high quality”; the checks should expose assumptions and produce evidence for a specific use.
A good spatial-data checklist therefore ends with a short statement such as: suitable for national thematic mapping; not suitable for property-level measurement because boundary accuracy and update date are insufficiently documented.
That conclusion is more useful than a generic quality score because it connects the inspection directly to the work the data is expected to support.
Related content
What Is Spatial Metadata? — documenting the dataset
What Is Spatial Data Provenance? — source and transformation lineage
Public Data vs Safe-to-Publish Data — publication risk