Journal

Why Evidence Should Travel with Extracted Geography

When a report says that a bridge lies near a village, extracting a point or line is only half the task.

Geobble

Develop the argument implied by “Why Evidence Should Travel with Extracted Geography” as a distinct Geobble perspective, while delegating neutral definitions and procedures to canonical Learn and Help pages.

Why Evidence Should Travel with Extracted Geography

A report says a bridge lies two kilometres north of a village, a field note names a wetland but gives no coordinates, and a PDF table contains latitude and longitude, except one row appears to have them reversed. Turning those statements into points and polygons can make the information far more useful.

It can also make the interpretation look more certain than the source ever was.

Once the geometry leaves the document and enters a dataset, it acquires the visual authority of a feature on a map. Unless the evidence travels with it, later users may have no way to tell an observed coordinate from an inferred location.

Extraction is not copying

Some geographic extraction is literal: a document contains explicit coordinates and the workflow converts them into a point. Much of it, however, is not.

A place name may need to be matched to a gazetteer, a relative phrase such as “near the eastern border” interpreted, and a sketch related to a known landscape. A facility name may also refer to several locations, while even explicit coordinates can have an uncertain reference system or order.

The resulting feature therefore contains two things: information from the source and decisions made during transformation.

If only the final geometry is retained, those categories collapse. The output looks observational even when it is partly inferential.

A document-level citation is often too coarse

It is useful to know that a dataset came from a particular report. That does not necessarily help a reviewer understand one disputed feature.

Imagine a report with 200 pages and 80 extracted locations. When one point is challenged, a citation to the report's landing page forces the reviewer to reconstruct the extraction from scratch. Which paragraph mentioned the place? Did a table contain coordinates? Was the geometry inferred from surrounding prose? Were there several candidate locations?

Feature-level evidence changes the situation. The reviewer can return to the relevant page, passage, table or URL and inspect the basis of the claim.

The same principle applies to attributes. If a status field was derived from a sentence in a document, the supporting passage is more useful than a generic citation to the whole publication.

Evidence should describe the source, not imitate reasoning

There is an important distinction between evidence and an explanation generated by the system.

A model may produce a fluent rationale for why a point belongs in one village rather than another. Although that rationale can help a reviewer understand the proposal, it is not primary evidence; the evidence is the source material, a reference dataset, the geocoding result, the recorded transformation or another inspectable input.

This matters because generated explanations can be persuasive even when they are wrong. Reviewable extraction should therefore preserve verifiable descriptors: document, page, passage, source URL, candidate match, coordinates, or other concrete support.

Provenance standards make the broader idea explicit: useful provenance represents relationships between entities, activities and agents rather than treating an output as context-free. The vocabulary can be lightweight in an application, but the principle is valuable. A derived feature should not lose the chain that explains how it came to exist.

Uncertainty should survive acceptance

A reviewer may accept a feature even when its location is approximate.

Perhaps the source identifies a village but not the exact facility, or the best available evidence places an event within a district rather than at a precise point. Acceptance should not require pretending the ambiguity has vanished.

That means evidence and uncertainty need to survive the transition from proposal to reusable data. A later analyst should be able to distinguish “exact coordinates reported by the source” from “location geocoded from a place name” or “representative point chosen for an area-level reference”.

Without that distinction, downstream analysis can become more precise than the evidence. A nearest-neighbour calculation may use metre-level distances between points whose original locations were only known to the village. A polished map can then conceal the mismatch.

Evidence makes correction local

Good provenance reduces the cost of being wrong.

If a feature turns out to be misplaced, the reviewer can revisit the evidence for that feature, choose another candidate or mark it unresolved. The rest of the dataset does not need to be distrusted simply because one interpretation changed.

This also matters when tools improve. A better gazetteer or reference layer may make a previously ambiguous place resolvable. If the original evidence is still attached, the feature can be reprocessed deliberately rather than treated as an unexplained coordinate inherited from the past.

Evidence therefore increases reusability. It lets a future user judge whether the original interpretation is suitable for a new purpose.

Geography becomes more reusable when its claims remain inspectable

The point of extraction is not to preserve documents in amber. It is to turn useful information into geography that can participate in mapping and analysis.

But reuse should not require context loss.

A point is more valuable when a user can see why it is there, an attribute is more trustworthy when its support can be inspected, and an unresolved feature is safer when it remains visibly unresolved. These properties do not guarantee correctness; they preserve the possibility of correction.

That is the difference between a reusable geographic claim and an unexplained mark on a map.

References

Related content