TY - JOUR AU - Wallgrün, Jan O. AB - In this article we present GeoTxt, a scalable geoparsing system for the recognition and geolocation of place names in unstructured text. GeoTxt offers six named entity recognition (NER) algorithms for place name recognition, and utilizes an enterprise search engine for the indexing, ranking, and retrieval of toponyms, enabling scalable geoparsing for streaming text. GeoTxt offers a flexible application programming interface (API), allowing for customized attribute and/or spatial ranking of retrieved toponyms. We evaluate the system on a corpus of manually geo‐annotated tweets. First, we benchmark the performance of the six NERs that GeoTxt provides access to. Second, we assess GeoTxt toponym resolution accuracy incrementally, demonstrating improvements in toponym resolution achieved (or not achieved) by adding specific heuristics and disambiguation methods. Compared to using the GeoNames web service, GeoTxt's toponym resolution demonstrates a 20% accuracy gain. Our results show that places mentioned in the same tweet do not tend to be geographically proximate. TI - GeoTxt: A scalable geoparsing system for unstructured text geolocation JO - Transactions in Gis DO - 10.1111/tgis.12510 DA - 2019-02-01 UR - https://www.deepdyve.com/lp/wiley/geotxt-a-scalable-geoparsing-system-for-unstructured-text-geolocation-ZXy8b4Osdq SP - 118 EP - 136 VL - 23 IS - 1 DP - DeepDyve ER -