← All work

Product / Research / Artistic practice

Read asIndustryFounder / CTOAcademicArtist

Terra Cognita

Global · Equal EarthAn equal-area world plate: the atlas covers the whole globe, and an equal-area projection keeps no region visually inflated. It shows extent, not article density.

Fifty thousand English Wikipedia articles kept at their real coordinates, coloured and searched by meaning, so readers can see where written knowledge gathers and where it thins.

What existsA working atlas of 50,000 geotagged English articles and a separate exploration of 23 Wikipedia language extracts.

My role
Independent researcher and developer
Period
2026–present
Status
Research prototype
Focus
Geotagged knowledge and its geographic coverage
Terra Cognita showing a search for volcano, with matching Wikipedia articles on their geographic locations.
Search for “volcano”: related articles stay in their real locations. English development subset. Wikipedia text: CC BY-SA 4.0; Wikidata: CC0. Basemap: CARTO and OpenStreetMap contributors.View full image (opens in a new tab)

The question

Most semantic atlases move places into an abstract model space. I wanted to keep the real geography and make meaning another way to read it, while showing where Wikipedia coverage is uneven.

My contribution

I built the data pipeline, mapping interface and browser search. Geographic coordinates determine position; language-model embeddings supply colour and similarity. Precomputed vectors, Arrow metadata and a web worker support search across the development subset. I also developed a separate study of geographic coverage across selected language editions.

How I approached it

  1. Prepare the articles

    I combined geographic coordinates, article metadata and semantic embeddings into files the browser can load.

  2. Keep meaning on the map

    I used real coordinates for position and embeddings for colour and query similarity.

  3. Examine the gaps

    I compared the uneven geography of selected language editions, keeping that study separate from the served English subset.

What exists

  • A working development subset containing 50,000 geotagged English Wikipedia articles, verified against the shipped Arrow metadata and vector index.
  • An implemented semantic search interface that keeps results at their geographic coordinates.
  • A separate multilingual coverage study using 23 selected Wikipedia extracts.

Where the work stands

The served map is an English development subset, and the language study samples selected editions rather than all of Wikipedia. Semantic similarity is a model-derived association, not proof of a factual connection between places.

Design choices and constraints
  • Keep geographic position separate from model-derived similarity.
  • Run development-subset search in a browser worker using precomputed data.
  • Present article density as knowledge coverage, not the importance of a place.

Language, editorial attention and geotagging all shape the source. Embedding similarity also depends on the model and the text available for each article.

Explore the work