Showing posts with label Bad Data. Show all posts
Showing posts with label Bad Data. Show all posts

Monday, March 28, 2022

Do Not Use the Pleiades Data Set

 

The river Quartaccio is a brook (described as 'fossa') just near Marina di San Nicola on the Lazio coast of Italy. This brook diverges from the Statua (another brook) at 41.9298 N, 12.135 E and flows N towards the town of Quartaccio and Valcanneto. Along this brook at about 41.9356 N, 12.1381 E a modern road crosses it. This is approximately the position of a Roman bridge which is given no name. This bridge is marked at this place in the Barrington Atlas (Map 43, A2, no. 2).   Here's what it looks like in Google Earth:



All in all it took me about half an hour to trace down the right Quartaccio (there are several) and take a stab at the position of the bridge based on the Barrington Atlas.


In Pleiades this bridge is no. 426578 and its position is given by Pleiades as 41.875 N, 12.125 E which is in the Tyrrhenian some 3.69 km. from the nearest land and 6.83 km from the putative bridge.  See it at the bottom of the next photo:


Now these gross errors of placement in the Pleiades data are nothing new. I looked at five examples last time and I attributed them to careless digitization and the complete lack of any quality control.   I did notice that Pleiades puts their accuracy estimate for the Quartaccio Bridge as 'Rough'. I thought, 'Aha!',  if I can see what 'Rough' means to the Pleiades group then maybe I can see  how far the  problem goes.   I got the 'Rough' points out of the database that I made from Pleiades' data and plotted that on Google Earth. There are about 5000 points labelled 'Rough' in the Pleiades DB. I used a select to retrieve the first 3000 of these (about 1/10 of the entire Pleiades DB). Here's the result of plotting these 3000:



This is shocking.  

Even if the location of these sites was not known (most of them are known as I will show) Pleiades has made no effort to place them even approximately.  These points are simply dumped off at the nearest half-degree or quarter-degree vertex.  Nothing else could make such a pattern.

And most of these points are not singles.  They are multiples of 5 to 10 distinct locations which are simply dumped on top of each other. All these sites have simply been rounded to the nearest degree or half-degree. And most of these points are easily locatable. When I was criticizing P. for errors I truly believed that the number was, perhaps, about 100 points. Now I see that it is thousands of points that could be known but which are simply dumped off at the nearest vertex.

Greece looks like this:




I've said that many of these points are just carelessly dumped in a pile even though the locations they denote are easily locatable.  Here's an example.  In the Corinthian Gulf Pleiades has overlain six points at 38.25 N, 22.25 E.



I've drawn red lines from the vertex where they were dumped to their real locations.  The line drawn to Phocis is drawn to the nearest point of land in Phocis.  The lines from the  rivers are drawn from the mouths of the respective rivers.  The Styx is the modern Mavronero above Solos.  The village of Solos is marked.  Every point was easily locatable and the average error is 13.9 km.  

In Pleiades-world four rivers, an entire province, and a gulf are all in exactly the same location.   All of their 'rough' points are exactly the same.

I can think of no toponymical practice that can explain this pattern.  What's absolutely clear is that the Pleiades data set has no conceivable use for any scholarly purpose.   

The Pleiades data set is approaching twenty years of existence.  In that time the P. team had plenty of time to fix these gross errors.  I concluded long ago that Pleiades never had any interest in creating a scholarly (or even a reliable) map of the ancient world.  Soon I will dedicate one of these posts to discussing what I think their real interest is.

The Pleiades data is unsuitable as a reference.  

It is unsuitable for study of the Ancient World.  

It is unusable as a basis for developing further geographic tools.

Pleiades owes the scholarly community an explanation of why their data so bad and what they intend to do to rectify this very poor database.









Sunday, March 11, 2018

Correspondent Comments on the Suitability of Pleiades Data for Scholars


A friend of mine replied to my post on the inaccuracy and unsuitability of Pleiades data for scholarly work.  (http://mycenaeanatlasproject.blogspot.com/2018/02/pleiades-data-does-crowd-sourcing-for.html)

I reproduce his letter here:

"Pleiades isn't structured to provide a single, accurate set of coordinates, though I think it hopes to evolve in that direction. Its most useful role currently is as a set of identifiers that allow links to superior gazetteers. For example, the huge error for ancient Messene is the result of displaying a calculated representative point that includes one spurious DARMC location (from the modern village of Messene), in addition to a mildly inaccurate DARMC location plus a very accurate DARE location. (DARE has assimilated a bunch of Google Earth-validated ToposText points for Greece, and from other sources as well, but uses Pleiades IDs as an easy pivot to other resources). Some of the tools Pleiades funding has produced for the purpose of improving its data are not being used very much -- one problem being a technological gap between laborious on-the-ground collectors ... and people who automate things."

Now I look at it piece by piece (original letter in red, my replies in black)

"Pleiades isn't structured to provide a single, accurate set of coordinates,"

So then where do we go from here?

" ..., though I think it hopes to evolve in that direction."

Spoiler alert: they're not going to. This would involve an enormous amount of work - actual scholarship. They're not going to commit to this because they think that this can be done on the cheap - through copying other data sets or through crowd sourcing. That's not the way that any of this works. My experience with them is that they will correct an error if you bring it forcefully to their attention but not otherwise.

"For example, the huge error for ancient Messene is the result of displaying a calculated representative point that includes one spurious DARMC location (from the modern village of Messene), in addition to a mildly inaccurate DARMC location plus a very accurate DARE location. (DARE has assimilated a bunch of Google Earth-validated ToposText points for Greece, and from other sources as well, but uses Pleiades IDs as an easy pivot to other resources). "

You've explained Messene but what about all the other errors? Nor have you questioned my estimate that approx. 1/3 of Pleiades has serious errors. What you're describing sounds like a real incestuous tangle. I don't even want to get into unpacking this beyond saying that topographical accuracy does not come from copying other data sets. It's like the old saw about buying a used car: you're just buying someone else's problems.

" Some of the tools Pleiades funding has produced for the purpose of improving its data are not being used very much -- one problem being a technological gap between laborious on-the-ground collectors (like me) and people who automate things. "

Sounds like you're describing Recogito. Is that what you mean? Are there other tools that they support? I tried out their conversion tool Geocollider. It failed miserably.

Everything about Pleiades/Pelagios/Peripleo is sham.  The Barrington Atlas data was useful for its printed purpose but now they're trying to roll that data over into the digital world where its approximative nature makes it unfit for use. And they've wrapped the whole thing up with bad and outmoded ideas - not from scholarly practice, from anthropology or toponymy or history or classical studies or any other relevant discipline - but from computer science. None of what they're doing (crowd sourcing and linked data) has anything to do with any scholarly practice or purpose but this is what they're selling and they're getting pots of money for it. In the end actual scholars wind up exactly where they started - having to do the topography of the Mediterranean from scratch. I actually know a fellow (from a very prestigious school) who's preparing a study of Mediterranean habitations. I was shown one of his spreadsheets and it was stuffed with errors since he had relied on Pleiades. In fact that's where my blog post came from.

I've asked myself what their game is. I suspect that what they want is to license their data (or perhaps their follow-on project Pelagios/Peripleo) to schools for so much per seat and deny access to non-customers. That's a time-honored approach in Computer Science. First get a lot of contributors to fork over their work for nothing under the name of something noble-sounding like 'Open Data' or the 'Semantic Web'. Second, license the whole to third parties and keep all the money.

Although I'm not sure that they can really carry this out successfully because it's the Underpants Gnomes business model.


Bibliography

DARE: The Digital Atlas of the Roman Empire.  https://dare.ht.lu.se/

DARMC: The Digital Atlas of Roman and Medieval Civilizations. https://darmc.harvard.edu/

Geocollider: https://pleiades.stoa.org/news/blog/introducing-geocollider

Recogito:  https://recogito.pelagios.org/

Underpants Gnomes: https://vimeo.com/79954057

Blog Posts Concerning the Isthmian Wall

Since 2023 a number of posts concerning the Isthmian Wall and how we located its remaining segments, have appeared on this blog.  This post ...