Digital Geo Media has been integrated into the curriculum, and students are already benefiting from it. This enables a previously untapped connection between spatial context and an understanding of how to use digital media at our university. Students gain insights into methodologies and data foundations that were not previously included in the existing modules of the two degree programs but are becoming increasingly important in the context of data literacy and digital sovereignty in today’s world—and for our students as well.
Future results will be continuously added here throughout the project’s duration.
1. Automated Georeferencing of German News Articles – Using *tagesschau* as an Example
How can German news articles be georeferenced, and what is the quality of the results? With this question in mind, students Lukas Kauf and Fabian Püschel began their project titled “Automated Georeferencing of German News Articles – Using the Tagesschau as an Example” as part of the Geogovernment 2 master’s module in the Geoinformatics and Surveying program. They developed a methodology for evaluating various geotagging systems (see Figure 1). In this process, a manually created reference dataset consisting of 120 news articles was compared with the results from specialized English- and German-speaking taggers.
The results in Figure 2 show that, in this study, the Stanford University geotagger—which was trained on German texts—performed best with an F1 score of 0.77. While it produced a comparable number of correct toponyms, but with a higher number of “false positive” classifications, the complete system from the University of Edinburgh, the English-language model from Stanford University, and the multilingual model from Spacy performed slightly worse.
3. Visual Analytics & Data Visualization with a Hands-On Session
An essential component of data journalism is the visual analysis of data and information to make it accessible to a broad audience. In the “Data Journalism and Visualization” elective, students were introduced to approaches to visual analysis and data visualization. Afterward, students used the web-based application Datawrapper (https://www.datawrapper.de) to analyze and visualize their own data (e.g., in the form of bar and column charts).
4. Visual Communication with Geodata—An Introduction to Visual Analytics in Deep Learning and Data Visualization with a Hands-On Session
In the “Digital Communication” module, various approaches to visual analysis and communication using spatial data were presented (e.g., communication via maps). Students then applied this theoretical knowledge in practice by analyzing sample datasets and visualizing interesting aspects on maps.
5. Untapped Information in Social Media Posts—Geolocation and Sentiment Analysis
The goal of this study was to determine how sentiments in social media posts about tourist attractions in Mainz can be captured and analyzed using untapped data, as well as what added value this data could offer in the future. To this end, two fundamental steps were carried out: data collection and data analysis. As part of the data collection process, suitable Instagram posts were selected and assigned a location (geotagging). The posts were divided into 10 categories of tourist attractions. During the data analysis phase, a sentiment analysis was performed for each tourist attraction. Investigations were conducted into the correlation between “likes” and sentiment, as well as between the number of followers and “likes,” and the relationship between “likes,” days of the week, and sentiment. The following section describes the results of the study.
The study focused on determining the prevailing sentiment toward various landmarks in Mainz. An analysis was conducted in which all landmarks were compared to identify which sentiment occurred most frequently in the selected posts. Mainz’s Old Town and Mainz Volkspark received the most positive ratings, while Mainz Central Station was the least popular attraction (see Figure 1). It was also found that the total number of “likes” is not an indicator of a positive rating for a tourist attraction, as was the case with Mainz’s Neustadt (see Figure 2).
The students also conducted an analysis of the days of the week (see Figure 3) and found that Thursday had by far the fewest posts. Most posts were published on Sundays and Mondays, but the day with the most likes was Tuesday. On Thursdays, most “likes” are given with a neutral sentiment. Influencers could benefit from these results by posting on Thursdays to potentially receive more “likes.” On Sundays, positive “likes” outnumber those with other sentiments.
An analysis of follower counts shows that the mean is 3,703, but the median—at 667—provides a more accurate picture of the average Instagram user. This discrepancy is due to influencer profiles with over 10,000 followers. The highest number of followers is 177,000, but this account has only 343 likes. The mean number of likes is 183.74, and there is a moderately strong positive relationship between likes and followers, with a correlation of 0.41.
This project was conducted by Malte Schildmann, Luca Fahl, Sophie Schraml, Sarah Schnabel, and Delphine Pourikas.