Showing posts with label Pelagios. Show all posts
Showing posts with label Pelagios. Show all posts
Sunday, March 11, 2018
Correspondent Comments on the Suitability of Pleiades Data for Scholars
A friend of mine replied to my post on the inaccuracy and unsuitability of Pleiades data for scholarly work. (http://mycenaeanatlasproject.blogspot.com/2018/02/pleiades-data-does-crowd-sourcing-for.html)
I reproduce his letter here:
"Pleiades isn't structured to provide a single, accurate set of coordinates, though I think it hopes to evolve in that direction. Its most useful role currently is as a set of identifiers that allow links to superior gazetteers. For example, the huge error for ancient Messene is the result of displaying a calculated representative point that includes one spurious DARMC location (from the modern village of Messene), in addition to a mildly inaccurate DARMC location plus a very accurate DARE location. (DARE has assimilated a bunch of Google Earth-validated ToposText points for Greece, and from other sources as well, but uses Pleiades IDs as an easy pivot to other resources). Some of the tools Pleiades funding has produced for the purpose of improving its data are not being used very much -- one problem being a technological gap between laborious on-the-ground collectors ... and people who automate things."
Now I look at it piece by piece (original letter in red, my replies in black)
"Pleiades isn't structured to provide a single, accurate set of coordinates,"
So then where do we go from here?
" ..., though I think it hopes to evolve in that direction."
Spoiler alert: they're not going to. This would involve an enormous amount of work - actual scholarship. They're not going to commit to this because they think that this can be done on the cheap - through copying other data sets or through crowd sourcing. That's not the way that any of this works. My experience with them is that they will correct an error if you bring it forcefully to their attention but not otherwise.
"For example, the huge error for ancient Messene is the result of displaying a calculated representative point that includes one spurious DARMC location (from the modern village of Messene), in addition to a mildly inaccurate DARMC location plus a very accurate DARE location. (DARE has assimilated a bunch of Google Earth-validated ToposText points for Greece, and from other sources as well, but uses Pleiades IDs as an easy pivot to other resources). "
You've explained Messene but what about all the other errors? Nor have you questioned my estimate that approx. 1/3 of Pleiades has serious errors. What you're describing sounds like a real incestuous tangle. I don't even want to get into unpacking this beyond saying that topographical accuracy does not come from copying other data sets. It's like the old saw about buying a used car: you're just buying someone else's problems.
" Some of the tools Pleiades funding has produced for the purpose of improving its data are not being used very much -- one problem being a technological gap between laborious on-the-ground collectors (like me) and people who automate things. "
Sounds like you're describing Recogito. Is that what you mean? Are there other tools that they support? I tried out their conversion tool Geocollider. It failed miserably.
Everything about Pleiades/Pelagios/Peripleo is sham. The Barrington Atlas data was useful for its printed purpose but now they're trying to roll that data over into the digital world where its approximative nature makes it unfit for use. And they've wrapped the whole thing up with bad and outmoded ideas - not from scholarly practice, from anthropology or toponymy or history or classical studies or any other relevant discipline - but from computer science. None of what they're doing (crowd sourcing and linked data) has anything to do with any scholarly practice or purpose but this is what they're selling and they're getting pots of money for it. In the end actual scholars wind up exactly where they started - having to do the topography of the Mediterranean from scratch. I actually know a fellow (from a very prestigious school) who's preparing a study of Mediterranean habitations. I was shown one of his spreadsheets and it was stuffed with errors since he had relied on Pleiades. In fact that's where my blog post came from.
I've asked myself what their game is. I suspect that what they want is to license their data (or perhaps their follow-on project Pelagios/Peripleo) to schools for so much per seat and deny access to non-customers. That's a time-honored approach in Computer Science. First get a lot of contributors to fork over their work for nothing under the name of something noble-sounding like 'Open Data' or the 'Semantic Web'. Second, license the whole to third parties and keep all the money.
Although I'm not sure that they can really carry this out successfully because it's the Underpants Gnomes business model.
Bibliography
DARE: The Digital Atlas of the Roman Empire. https://dare.ht.lu.se/
DARMC: The Digital Atlas of Roman and Medieval Civilizations. https://darmc.harvard.edu/
Geocollider: https://pleiades.stoa.org/news/blog/introducing-geocollider
Recogito: https://recogito.pelagios.org/
Underpants Gnomes: https://vimeo.com/79954057
Sunday, February 25, 2018
Pleiades Data - Does Crowd Sourcing for Toponymy Actually Work?
In a previous post I discussed the magnitude of the errors
we could expect in a positioning system limited to two-place decimal
fractions. I suggested that the average
error that we could expect in such a system at latitudes approximately 35° N
would be about 400 m. I thought that
this imprecision would bar such a system from serious work in toponymy and I
asked what product would use a system so limited in mathematical range and so obviously unsuitable for the purpose.
The Pleiades Project is an attempt to adapt crowd-sourcing
to the field of toponymy. They have
received generous financial support from the National Endowment for the
Humanities to the amount of $1,140,780.
You can read about their grants here, here, here, and here.
What has the world of scholarship received for all this
money?
Recently I was checking the positions for a random sample of
geographic locations of classical and hellenistic locations in Greece. The locations were derived from Pleiades. This list was not selected by me but by a
colleague. Out of 45 locations 17 (36%)
had significant errors.
The smallest error
was 357 m. and the largest was 3775 m.
The average error or arithmetic mean (in the erroneous part of the
sample) was 1,514.7 m. The median error
was 1200 m. I present no standard deviation because the
errors are not normally distributed. In fact it appears as though the error
distribution in Pleiades data might be bimodal. This suggests that there is more than
one underlying cause for the Pleiades system’s inaccuracies.
![]() |
| From my error worksheet. Y-dimension is error in m. |
The 13 less erroneous locations may derive from crowd
contributions. The uppermost 4 positions (Passaron, Skotoussa, Messene, Antigonia) may be remnants of the original digitization of the Barrington Atlas data –
however that was accomplished. In other
words I am suggesting that it appears
as though crowd-sourcing tends to smooth out but not eliminate the original
digitizing errors. I emphasize that these are suppositions on my part. But, clearly, the complicated history of Pleiades' data generation has left a signature in the error results.
Here is a link to my worksheet. Occasional references in that worksheet to sites as 'Fnnn' or 'Cnnn' may be resolved at the site http://helladic.info/.
If these results are upheld by others then I would suggest that Pleiades is not an appropriate component of any scholarly work. If a system with a two-place fractional component has an average error of nearly 400 m. then the average error of Pleiades data of more than 1500 m. suggests that Pleiades data - at some point in its generation - never had an accuracy better than 0.5 to 1.5 fractional places (10^-0.5 to 10^-1.5).
I estimate that it takes at least 2 hours of research to reliably establish a location from Bronze Age or later sites in Greece. I do not know how many data points Pleiades claims but if it is, for example, 10000 points then it would require an effort of about 20000 man hours to complete a reliability review for Pleiades. At 2200 man hours in a man year that would require about 9 man years to complete. This is an order of magnitude estimate.
Crowd-sourcing in toponymy studies does not appear to work.
Crowd-sourcing in toponymy studies does not appear to work.
This defective data of Pleiades casts a shadow downstream - for example in such derivative products as Pelagios/Peripleo.
If Pleiades cannot undertake a good faith reliability study it should be rejected by the scholarly community.
Sunday, February 11, 2018
Fact Computing, Part 2
In a previous post I said that Pelagios Commons/Peripleo does not have its roots in the world of scholarship but in the world of computer science - specifically the ideas of Tim Berners-Lee.(1) Now let's concentrate more specifically on what Pelagios Commons/Peripleo really does.
The first
thing that must be clearly understood is that Pelagios Commons/Peripleo creates
no scholarly content and has nothing whatsoever to do with any Classical scholarship. It is strictly a computer science construct and, with a different database, would be perfectly at home in the world of migration tracking, chemical research, or anything else.
Peripleo is simply a
front-end site or data aggregator of a very common type.
Pelagios Commons links
large amounts of data produced by other non-related sites and entities and
subsumes them under a common format. It then
exploits this umbrella format in order to write its own front-end viewing tools
(Peripleo).
Its business
model is exactly like that of Huffington Post and any one of hundreds of
similar sites. Through an agreement with
providers it reproduces their work tout
court. They say that these unpaid
contributors are members of a ‘Community’ but this ‘Community’ is nothing more
than the stable of content providers who give away to Pelagios the fruits of their
labors. The most amusing statement on
the Pelagios website strenuously denies this plainly obvious fact:
Well,
Pelagios provides no original scholarly content.
Pelagios exclusively displays content provided
by others.
Pelagios forces their providers to reduce their own work into a Pelagios format in order
for Pelagios' software to display it.
Peripleo implements numerous search options.
What else
can Pelagios/Peripleo be but an
aggregator/search portal? In fact, if you go to their Peripleo splash page they clearly say 'Peripleo is a search engine ...'. The fact that
their content providers cooperate in the theft of their own labors does not
change the essential nature of the arrangement (This was true for Huffington Post which disguised its essential nature until the moment it went public).
The content providers are said by Pelagios to be members of a ‘Community’. From my many years as a professional computer
scientist I can assure my readers that this type of dishonest rebranding is
quite common everywhere in the online world.
The first step in any internet grift is to give it a
name that expresses the opposite of what it really is.
That it is the contributors who are to do all the work is also obvious from the tools that Pelagios Commons provides:
Recogito This
is an ‘online platform for collaborative document annotation’. But it is not the staff of Pelagios Commons
that’s going to do this annotation (how could they?). It is the contributor, the member of the ‘Community’
who creates this content.
Their Cookbook makes it easy to see who it is who does all the work for Pelagios (hint: not the
Pelagios staff themselves). In every
case the contributors are
responsible for massaging all their data into a form that Pelagios can accept. This is a cost to the contributor of many
hours of uncompensated labor. Pelagios
should disguise this aspect better than they do.
The following picture should make these several relationships clearer.
I have claimed that the Pelagios
Commons enterprise creates no content.
Strictly speaking that is not quite true. In fact, Pelagios Commons has achieved the Holy
Grail of academia: it is a perpetual motion machine for producing conference papers and web presentations. If you inspect the list to which I’ve linked
you will quickly see who it is who specifically benefits from the Pelagios Commons
enterprise.
~~~~~~~~~~~~~~~~~~~~~~~~~
Casting doubt on the Pelagios enterprise is
not to deny that some sort of digital structuring of the data that we have from
Mediterranean societies of antiquity would be useful. It would
be useful. But how is that goal to be
attained?
The data that comes to us (or
generated by us) relative to antiquity is of the most heterogeneous forms. Locations, building plans, daily customs,
food stuffs and their hypothesized yields, customs, clothing, trade, etc. Everything of human interest falls within the
purview of scholars of antiquity. This
is a classic data fusion problem. Data
fusion problems arise in environments where a number of sensors of different
types provide data of interest that is to be presented in a uniform view. Such problems arise in the cockpits of
fighter pilots and in very many environmental studies where, again, different
sensors (or the same types of sensors with different capabilities) are used to
gather data which is then to be united, combined or fused into a single point
of view.
Pelagios Commons dimly recognizes that this is the real problem. But they have performed this task
backwards. They start from the
assumption that Linked Data is the solution to everything. Upon that ideology they built a product which is useful
for no one. That’s the essential
problem. The site really isn’t good for
anything because it started ideologically. It did not start by asking what it is that scholars of ancient
societies really need in the form of digital support.
How should the social data from ancient
Mediterranean societies be fused? But, before
that, what does it mean, from the digital point of view, to support such scholars? Particularly in view of the fact that the
scholars in such fields have radically differing interests.
Notes
1) Pelagios Commons here links directly to a discussion of Tim Berners-Lee idea of Linked Data here.
Subscribe to:
Posts (Atom)
Blog Posts Concerning the Isthmian Wall
Since 2023 a number of posts concerning the Isthmian Wall and how we located its remaining segments, have appeared on this blog. This post ...
-
Users of Chapter 23 of the Laconia Survey, Cavanagh et al. [1996], etc. will probably be frustrated by the unnecessarily arcane Transverse M...
-
I learned from the Aegeanet bulletin board that the Arachne CMS databases are online. This includes a very large engraved seals data...

