Sources
last updated 6 September 2026
This page credits the sources used in Apera's word lists, coverage figures, maps, and cultural catalogues. Each section names the source and links to its licence or publisher.
The order the word lists are in
The Italian, French, German and Spanish lists are sorted by a published table of word counts — the 2018 release of Hermit Dave’s FrequencyWords, taken at a pinned commit. Its author licenses the contents under CC BY-SA 3.0. We read that table and use it to sort words we already carry; we count nothing ourselves.
Where those words were counted
The counts behind that table — and which words are on our lists at all — come from subtitles: the OpenSubtitles collection as distributed by the OPUS project at the University of Helsinki. The subtitles are written by volunteers rather than by the programmes they transcribe, which is why the counts describe how people speak on screen rather than how anyone writes.
OPUS distributes the collection under the Open Data Commons Attribution License (ODC-BY), which covers its own work of compiling and aligning it — not the subtitle text itself, which OPUS does not claim to own. For the 2018 release its maintainers ask for a citation of P. Lison and J. Tiedemann (2016), “OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles”, and a link to opensubtitles.org. We honour both.
The coverage figures
Where a word list is described as covering some share of ordinary writing, and where a story on the placement page has to place a word rarer than our own list reaches, those figures are read from wordfreq rather than from our own ranking — deliberately, so that a figure describing how far our list gets you is not measured with the list itself. The word shading you see after the quiz reads our own list, not wordfreq. The wordfreq package is released under the Apache License 2.0, and the data it ships is published under the Creative Commons Attribution-ShareAlike 4.0 licence.
The relief maps on the region pages
The shaded land behind each region map is generated once from public elevation data and served as a picture; nothing is fetched while you read it. The heights over Italy come from EU-DEM, produced using Copernicus data and information funded by the European Union, reached through the Terrain Tiles dataset on the Registry of Open Data on AWS, which also contributes ETOPO1 from the U.S. National Oceanic and Atmospheric Administration over the sea. The relief is derived from those elevations rather than reproduced from them — we have modified the data — and it is not endorsed by the European Union. The region outlines and the small shapes on the index are Natural Earth, which asks for no credit and is given one anyway.
The relief plates on the Spanish-language film region pages
The relief behind each Spanish-language film region page is generated once from public elevation data and served as a picture, the same way as the Italian region maps above, from a different mosaic: heights over Spain come from EU-DEM, produced using Copernicus data and information funded by the European Union; heights over the Americas come from the Shuttle Radar Topography Mission (SRTM), GMTED2010 and ETOPO1, all from U.S. federal sources. The relief is derived from those elevations rather than reproduced from them — we have modified the data — and none of it is endorsed by the European Union. The region outlines are Natural Earth, which asks for no credit and is given one anyway.
The relief plates on the French-language film region pages
The relief behind each French-language film region page is generated once from public elevation data and served as a picture, the same way as the maps above, from a third mosaic covering metropolitan France: heights come from EU-DEM, produced using Copernicus data and information funded by the European Union, with the Shuttle Radar Topography Mission (SRTM), GMTED2010 and ETOPO1 — all from U.S. federal sources — where EU-DEM does not reach. The relief is derived from those elevations rather than reproduced from them — we have modified the data — and none of it is endorsed by the European Union. The region outlines are Natural Earth, which asks for no credit and is given one anyway.
Who wrote a book
A book’s entry names its author, and that name comes from Wikidata — the same CC0 record the title, the year and the identifiers on that row come from, published under a waiver that asks for no credit. A row lists at most three names and then says and others. Where a row names nobody, Wikidata records no author for the work: that is a gap in the catalogue rather than a work published anonymously, and each book list says how many of its own titles are in it. We publish nothing else about the person — no dates, no nationality, no link.
Where a title is set, and what form a book takes
The region pages say a title — a film, a series or a book — is set somewhere on three authorities. The first is Wikidata, whose statement of where a work is set is published CC0 and asks for nothing. The other two are Italian Wikipedia and English Wikipedia, whose editors file a title under categories like Film ambientati a Roma and Television shows set in Naples — set there, which is a different claim from filmed there, and only the first puts a title on one of these pages. Those categories are part of each encyclopedia’s content and are published under the Creative Commons Attribution-ShareAlike 4.0 licence. We read the category tags and the tree above them, from each Wikipedia’s own published database exports, and keep no article text.
Italian Wikipedia’s categories answer a second question about a book: whether it is prose or verse and drama, filed under categories like Romanzi and Poemi. We read that only where Wikidata’s own record of a work leaves the question open, and never to override what Wikidata already says. Same licence, same condition, same credit.
If something here is wrong
These are claims about other people’s work, and we would rather be corrected than right by default. Write to [email protected] and say which entry and what is wrong with it.
The film pages in particular show what we can see in the data. The records have gaps and the occasional plain error, and every description of how a film sounds is a judgement rather than a measurement — nobody here has a stopwatch on the dialogue. So if a page disagrees with what you have heard, believe your ears. The best correction to any of it is a trip, and a question to somebody who lives there; the second best is the address above.