Showing posts with label GeoSocial. Show all posts
Showing posts with label GeoSocial. Show all posts

Wednesday, May 27, 2026

New Paper: Exploring Fear in Urban Environments

In the past we have written about how we have used social media to study a plethora of topics with respect the the form and function of cities among many other things. But one thing we have not explored is fear and more specifically fear of crime and how this can be mined through geosocial media. 

This has now changed with a new paper entitled "Exploring Fear in Urban Environments: Place and Space Analysis of Social Media Data" which has recently been published in Applied Geography.  In this paper, Ying Zhou and myself extract fear related posts from social media and examine the places and spaces where people experience fear, as well as the factors that contribute to it in New York City. 

We do this by utilizing Natural Language Processing (NLP) techniques for sentiment and text analysis, including a RoBERTa-based emotion classification model and the BERTopic model for topic modeling. The former model narrowed the raw data to those with the dominant emotion of fear, and the latter analyzed space- and place-related features that contribute to the fear sentiment. Then, the selected social media data were analyzed using spatial clustering methods (i.e., Hotspot Analysis (Getis-Ord Gi*) and Local Moran’s I) and compared with urban crime data for weekly trends and spatial patterns. As such the paper has the following research objectives:
  1. exploring places where people expressed fear through social media; 
  2. making comparisons between safety-related fear and crime from the perspective of both time and space; 
  3. extracting urban environmental and social features that lead to fear.

If this sounds of interest, and you wish to find out more with respect to our findings, below you can read the abstract to the paper, see some of the figures which describe our research methodology and results while at the bottom of the post you can find a link to the paper itself. Finally the code we utilized in the paper can be found at https://osf.io/y7xfc/overview.

Abstract:

One goal of creating livable cities is to enhance public safety. While previous research in urban studies has focused on correlations between physical environments and crime, it has typically relied on criminal statistics. However, fear of crime is an emotional response to perceived risks rather than a direct reflection of crime levels, so it cannot be analyzed solely by crime data. Additionally, urban planning today has gradually shifted its focus from a top-down to a bottom-up approach, making it essential to understand and foster spaces where residents feel safe. This research examines the spaces and places where people experience fear, as well as the factors that contribute to it, in New York City. We utilized social media data to gather people’s expressions of the city and identified posts expressing fear emotion using the RoBERTa-based model and a rule-based classifier. Then, the selected social media data and crime were compared temporally by weekly trends and spatially by clustering methods (i.e., Hotspot Analysis (Getis-Ord Gi*) and Local Moran’s I). The results show that their temporal and spatial patterns partially have limited alignment. To delve into the origins of fear, we extend our analysis by adopting BERTopic to identify topics and summarize them into themes (e.g., places, transportation, people, others) to understand the bottom-up emergence of fear, thereby informing a people-centered approach to research on urban issues. 

Keywords: Social media; Natural language processing; Sentiment analysis; Urban environment.

Methodology framework.

An example of textual analysis on fear-related tweets: from machine-generated topics to human-interpreted themes describing fear in NYC.

Weekly trends comparison between safety-related fear and violent crime.

Clustering features analysis by the method of hotspot analysis (Getis-Ord Gi∗).

Full Reference: 

Zhou, Y. and Crooks, A.T. (2026), Exploring Fear in Urban Environments: Place and Space Analysis of Social Media Data, Applied Geography, 192: 104051 (pdf)

Tuesday, April 22, 2025

Mapping the Invisible

Readers might of noticed that recently we have been exploring the use of street view images to explore cities or how we can utilize geosocial media to understand the form of function of cities, but one thing we have not explored is the role of smell and how it shapes peoples perceptions of urban spaces. However, in a new paper recently published in the Annals of the American Association of Geographers with Qingqing Chen, Ate Poorthuis we do just that. The paper is entitled "Mapping the Invisible: Decoding Perceived Urban Smells Through Geosocial Media in New York City" In the paper we use text mining techniques to tease out smell related information from over 56 million geolocated tweets which are then assigned to specific small categories (e.g., nature, food, waste) resulting in a new smellscape map for New York city. 

If this sounds of interest, below you can read the abstract to our paper, see our workflow and resulting smellscape map. While the the analysis steps, along with the smell dictionary used, are documented in the research code compendium at  https://figshare.com/s/8418d47cdc5c539b78ab. Finally at the bottom of the page, you can find the full reference and a link to the paper. 

Abstract:

Smells can shape people’s perceptions of urban spaces, influencing how individuals relate themselves to the environment both physically and emotionally. Although the urban environment has long been conceived as a multisensory experience, research has mainly focused on the visual dimension, leaving smell largely understudied. This article aims to construct a flexible and efficient bottom-up framework for capturing and classifying perceived urban smells from individuals based on geosocial media data, thus, increasing our understanding of this relatively neglected sensory dimension in urban studies. We take New York City as a case study and decode perceived smells by teasing out specific smell-related indicator words through text mining techniques from a historical set of geosocial media data (i.e., Twitter/X). The data set consists of more than 56 million data points sent by more than 3.2 million users. The results demonstrate that this approach, which combines quantitative analysis with qualitative insights, can not only reveal “hidden” places with clear spatial smell patterns, but also capture elusive smells that might otherwise be overlooked. By making perceived smells measurable and visible, we can gain a more nuanced understanding of smellscapes and people’s sensory experiences within the urban environment. Overall, we hope our study opens up new possibilities for understanding urban spaces through an olfactory lens and, more broadly, multisensory urban experience research. 

Key Words: geosocial media, multisensory urban experiences, network analysis, New York City, smellscape, text mining, urban smells.

A framework of deriving perceived smells.

An overview of research workflow.

An overview of the six dominant overlapping smells across New York City using the weaving mapping method. The weaving map uses the concept of strands to represent attributes. Each strand here represents one specific smell category, with the intensity of the color changing based on the density of that smell category within each neighborhood (i.e., grid cells).

Full Reference: 

Chen, Q., Poorthuis A. and Crooks, A.T., (2025), Mapping the Invisible: Decoding Perceived Urban Smells through Geosocial Media in New York City, Annals of the American Association of Geographers, 115(6), 1444-1464. Available at https://doi.org/10.1080/24694452.2025.2485233. (pdf)

Monday, November 13, 2023

Synthetic Geosocial Network Generation

In the past the blog has explored the creation of social networks for models. Keeping with this vain of research, I was fortunate to work with Ketevan Gallagher, Taylor Anderson and Andreas Züfle to consider the role of location of individuals when generating social networks. This work has resulted in a new paper entitled "Synthetic Geosocial Network Data Generation"  which was presented at the 7th ACM SIGSPATIAL Workshop on Location-based Recommendations, Geosocial Networks and Geoadvertising (LocalRec 2023). If this sounds of interest, below you can read the abstract to the paper, see some the generated geosoical networks and find the full reference and link to the paper. In addition to this, the Python code and data used to generate the networks is available at https://github.com/KetevanGallagher/Synthetic-Geosocial-Networks.

Abstract: Generating synthetic social networks is an important task for many problems that study humans, their behavior, and their interactions. Geosocial networks enrich social networks with location information. Commonly used models to generate synthetic social networks include the classical Erdos-Renyi, Barabasi-Albert, and Watts-Strogatz models. However, these classic social network models do not consider the location of individuals. Real-world geosocial networks do exhibit a strong spatial autocorrelation, thus having a higher likelihood of a social connection between agents that are spatially close. As such, recent variants of the three classical models have been proposed to consider location information. Yet, these existing solutions assume that individuals are located on a uniform lattice and exhibit certain limitations when applied to real-world data that exhibits clusters. In this work, we discuss these limitations and propose new approaches to extend the three classic social network generation models to geosocial networks. Our experiments show that our generated synthetic geosocial networks address the shortcomings of the state-of-the-art models and generate realistic geosocial networks that exhibit high similarity to real-world geosocial networks. 
Keywords: Geosocial Networks, Network Generation, Synthetic Social Networks, Erdos-Renyi, Watts-Strogatz, Barabasi-Albert.


Real- World Geosocial Network using Facebook Social Connectedness Data between Zone Improvement Plan (ZIP) Region Centroids for the State of Virginia, USA.
Geosocial graphs using Virginia ZIP code data.
Graphs using Fairfax Census Tract data.


Full Referece:
Gallagher, K., Anderson, T., Crooks, A.T. and Züfle, A. (2023), Synthetic Geosocial Network Data Generation, Proceedings of the 7th ACM SIGSPATIAL Workshop on Location-based Recommendations, Geosocial Networks and Geoadvertising (LocalRec 2023), Hamburg, Germany. (pdf) (presentation)

Monday, September 19, 2022

Information propagation on cyber, relational and physical spaces about covid-19 vaccine

It seems that its been a quite some time that we posted about geosocial analysis but in a recent paper with  Fuzhen Yin  and Li Yin entitled "Information Propagation on Cyber, Relational and Physical Spaces about Covid-19 Vaccine: Using Social Media and the Splatial Framework" published in Computers, Environment and Urban Systems we revisit this line of work while at the same time linking it to Covid and vaccination debates. 

Specifically we examine the interaction between cyber, relational (i.e, networks between objects), and physical spaces using the Splatial framework. Through our analysis focused on New York State, we find that non-polarized vaccination debates were observed in cyber, relational, and physical spaces. Furthermore,  we found that while physical space users had less anti-vaccine stance than relational and cyber space users there were strong interactions are observed between physical–relational, and relational-cyber spaces.If this sort of thing interests you. Below we provide the abstract to the paper along with some figures which show the study area, our methodology and some of the results. While at the bottom of the post we provide the full reference and the link to the paper.

Abstract:

With the advent of social media, human dynamics studied in purely physical space have been extended to that of a cyber and relational context. However, connections and interactions between these hybrid spaces have not been sufficiently investigated. The “space-place (Splatial)” framework proposed in recent years allows capturing human activities in the hybrid of spaces. This study applies the Splatial framework to examine the information propagation between cyber, relational, and physical spaces through a case study of Covid-19 vaccine debates in New York State (NYS). Whereby the physical space represents the regional boundaries and locations of social media (i.e., Twitter) users in NYS, the relational space indicates the social networks of these NYS users, and the cyber space captures the larger conversational context of the vaccination debate. Our results suggest that the Covid-19 vaccine debate is not polarized across all three spaces as compared to that of other vaccines. However, the rate of users with a pro-vaccine stance decreases from physical to relational and cyber spaces. We also found that while users from different spaces interact with each other, they also engage in local communications with users from the same region or same space, and distance-based and boundary-confined clusters exist in cyber and relational space communities. These results based on the Splatial framework not only shed light on the vaccination debates but also help to define and elucidate the relationships between the three spaces. The intense interactions between spaces suggest incorporating people’s relational network and cyber presence in physical place-making.

Keywords: Covid-19, Vaccination, Social media, Social network analysis, Community detection, Urban informatics
Schematic representation of the three spaces: cyber, relational and physical spaces.

Map of study area (NYS) with the primary road system. Red dots denote collected vaccine-related tweets in NYS.

Research workflow to investigate the propagation of different opinions between three spaces: cyber, relational and physical spaces.

Network visualization of the eight top large communities in relational space. (A) Visualization of communities using ForceAtlas layout. (B) Project communities into physical space. Nodes without location information are placed outside of NYS.

The hybrid space network shows the information propagation between physical and relational spaces. (A) shows the network of all tweets, (B) shows the pro-vaccine tweets, and (C) shows the anti-vaccine tweets.
 
Full Reference:

Yin, F., Crooks, A.T. and Yin, L. (2022), Information Propagation on Cyber, Relational and Physical Spaces about Covid-19 Vaccine: Using Social Media and the Splatial Framework, Computers, Environment and Urban Systems. Available at: https://doi.org/10.1016/j.compenvurbsys.2022.101887.  (pdf)

Tuesday, May 25, 2021

Achieving Situational Awareness with Geolocated Social Media

Tuning back to our work on geosocial analysis we (Xiaoyi Yuan, Ron Mahabir, Arie Croitoru and myself) recently had a paper published in GeoJournal entitled "Achieving Situational Awareness of Drug Cartels with Geolocated Social Media." 
 
The overarching objective of this paper is to develop an approach that would enable the extraction of potentially relevant situational awareness-related information from geolocated raw data streams (in this example we use Twitter). We accomplish this goal by focusing on Named Entities (NEs) related to drug cartels rather than the raw text as a whole. Specifically, our analysis is performed on the NEs by first extracting them and then clustering them to identify relevant concepts/themes (using TextRazor). This approach gives rise to themes that can then be assessed for temporal and spatial patterns based on frequency in order to gain underlying insights into drug cartels. If is of interest to you below we provide the abstract to the paper, a diagram of our workflow and a sample of our results along with the link to the paper. Also the complete code for the analysis and results is available at https://bitbucket.org/xiaoyiyuan/cartel.
 

Abstract: Using geolocated tweets to achieve situational awareness is an often researched topic in disaster and emergency management. However, little has been done in the area of drug cartels, which, as transnational crime organizations, continue to pose great risk to the stability and safety of our communities. This paper made an initial effort in using geolocated social media (specifically Twitter) to achieve situational awareness of drug cartels through temporal and spatial analysis of derived named entity clusters. The results show that detecting peaks in the time series of frequently occurring entity clusters enabled the tracking of important events in public discourse surrounding drug cartels. Correlations between time series also provided valuable insights into the synchronicity between different events. Further examining the spatial distribution of key events for different countries, we identify thematic hotpots of public discourse on cartel activity. Our methodology also addresses issues of language ambiguity when working with noisy social media data in order to achieve situational awareness on drug cartels.

Keywords: Cartels, Social Media, Situational Awareness and Temporal and Spatial Analysis.

The workflow of achieving situational awareness of drug cartels using geolocated tweets.

Tweet and entity counts by language and geolocation.

An example of tweets of high frequency on peak day in Venezuela

Heat maps of frequencies of a Cluster for Day 14 and Days 18-21.


Full Reference:
Yuan, X., Mahabir, R., Crooks, A.T. and Croitoru, A. (2021), Achieving Situational Awareness of Drug Cartels with Geolocated Social Media, GeoJournal. DOI: https://doi.org/10.1007/s10708-021-10433-2 (pdf)

Friday, January 31, 2020

The Interplay Between the Media and the Public in Mass Shootings

Continuing our work on shootings we recently had a paper published in Criminology and Public Policy entitled: "Responses to Mass Shooting Events: The Interplay Between the Media and the Public." However, here we do not look at bots but instead explore the how the public responds to mass shooting events (e.g. Las Vegas, Sutherland Springs, Marshall County, Parkland, Santa Fe), by seeking additional information or exchanging opinions about them in media coverage (e.g. newspaper articles via LexisNexis) and through online sources of information (e.g. Google Trends, Wikipedia and Online Social Networks (i.e. Twitter)). 

Overall, our results show discernible patterns in both time and space in the public’s online information seeking activities after a mass shooting. In addition we find discernible online information seeking patterns in geographic space, with a focal area of interest in the state in which the shooting event occurs, surrounded by a region of reduced interest. This finding further suggests that online information seeking activities are driven, at least in part, by geographic proximity to mass shooting events.

If you wish to find out more about this research, below we provide the summary and policy implication to the paper along with some figures from our methodology (e.g., how we go about analyzing temporal and geographical trends) and some of the results. Finally at the bottom of the post we provide the full reference and a link to the paper.

Abstract:
Research Summary: Public mass shootings tend to capture the public’s attention and receive substantial coverage in both traditional media and online social networks (OSNs) and have become a salient topic in them. Motivated by this, the overarching objective of this paper is to advance our understanding of how the public responds to mass shooting events in such media outlets. Specifically, it aims to examine whether distinct information seeking patterns emerge over time and space, and whether associations between public mass shooting events emerge in online activities and discourse. Towards this objective, we study a sequence of five public mass shooting events that have occurred in the United States between October 2017 and May 2018 across three major dimensions: the public’s online information seeking activities, the media coverage, and the discourse that emerges in a prominent OSN. To capture these dimensions, respectively, data was collected and analyzed from Google Trends, LexisNexis, Wikipedia Page views, and Twitter. The results of our analysis suggest that distinct temporal patterns emerge in the public’s information seeking activities across different platforms, and that associations between an event and its preceding events emerge both in the media coverage and in OSNs.
Policy Implication: Studying the evolution of discourse in OSNs provides a valuable lens to observe how society’s views on public mass shooting events are formed and evolved over time and space. The ability to analyze such data allows tapping into the dynamics of reshaping and reframing public mass shooting events in the public sphere and enable it to be closely studied and modeled. A deeper understanding of this process, along with the emerging associations drawn between such events, can then provide policy and decision-makers with opportunities to better design policies and communicate the significance of their goals and objectives to the public.
A framework for the analysis of temporal and geographical trends .

The analysis processes of Twitter and LexisNexis data.

Geographic patterns in online search activity in Google Trends for the five events in our study.

Chronologically ordered Google Trends search activity (a, left) and Wikipedia page views (b, right). Each vertical solid black line marks the occurrence of one of four shooting events examined in the analysis (as indicated by the line label).

Mentions of prior events during the first approximately 1-month period following each event in each of the events studied. (a) Sutherland Springs, (b) Marshall County, (c) Parkland, (d) Santa Fe.

Full Reference:
Croitoru, A., Kien, S., Mahabir, R., Radzikowski, J., Crooks, A.T., Schuchard, R., Begay, T., Lee, A., Bettios, A. and Stefanidis, A. (2020), Responses to Mass Shooting Events: The Interplay Between the Media and the Public, Criminology and Public Policy, 19(2): 335–360. (pdf)

Tuesday, November 05, 2019

New Paper: Assessing the Placeness of Locations through User-contributed Content

In the past we have written about how one can use crowdsourced data to gain a collective sense of place from Twitter contributions and also from corresponding Wikipedia entries (e.g. here). In a new paper with Xiaoyi Yuan, we extend this work to explore how user-contributed data can be used to explore if urban places are becoming inauthentic due to urban commodification and standardization by chain stores such as restaurants. To this end, at the at 3rd ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery (GeoAI) we have a paper entitled: "Assessing the Placeness of Locations through User-contributed Content"

In the paper we attempt to understand the relationship between restaurants and urban identities via user-contributed content. We extracted and analyzed information from over 3 million Yelp reviews from 37,000 restaurants using a Convolutional Neural Network (CNN) model in order to study places from the bottom up. Specifically we were interested to what extent cities share similarities or differences in their Yelp restaurant reviews. Furthermore, we wanted to explore how opinion aspects (i.e. what reviewers care about the most) are mentioned differently in urban chain and independent restaurants. Through the analysis of the Yelp reviews we find that online geo-tagged text data is fruitful for understanding places and aspect-based sentiment analysis helps us understand the large volumes of text. Not only did we discover that cities show homogeneity in terms of restaurant reviews, but for chain restaurants, “location” often emphasizes the differences between different stores of the same chain whereas for independent restaurant reviews, the aspect “location” reflects the characteristics of the places the restaurants are situated. If this is of interest to you, below we provide the abstract to the paper, along with some of the key findings and a link to the paper.

Abstract
Previous research has argued that urban places are becoming “placeless” and inauthentic. Many local policies have also proposed to encourage more independent stores in order to restore urban identity. Others argue, however, that chain stores provide affordable merchandise and different locations of the same chain may have different meanings to an individual. The research presented in this paper uses a Convolutional Neural Networks model to extract opinion aspects from more than 3 million user-contributed Yelp restaurant reviews. The results show high homogeneity among cities in terms of the average proportions of aspects in restaurant reviews. In addition, for fast food chains, “location” is the only aspect category reviewed proportionally higher than independent fast food restaurants. An analysis of the co-occurrences of “location” indicates that the identity of chain restaurants stems from the comparison between the same chain of different locations whereas the identity of the independent restaurants is more diverse, implying the intricacies of placeness of urban stores. This research demonstrates that fine-grained sentiment analysis (i.e., opinion aspect extraction and analysis) with geo-tagged text data is fruitful for studying nuanced place perceptions on a large scale.
KEYWORDS: Urban Places, Convolutional Neural Networks, Aspect-based Sentiment Analysis
Figure 1: Illustration of an example of a CNN layer.
Figure 3: Mapping restaurants in NV, AZ, PA, NC, WI, IL. Not all cities are shown in each state. Only cities have data that accounts for the majority of the restaurants in that state are mapped, for the sake of visual clarity.
Figure 6: Average proportions of aspect categories for chain and independent fast food restaurants for two kinds of cuisine (American, Mexican) in Las Vegas, Phoenix, and Charlotte, normalized by dividing the mean for comparison.
Reference:
Yuan X. and Crooks A.T. (2019), Assessing the Placeness of Locations through User-contributed Content, in Gao, S., Newsam, S., Zhao, L., Lunga, D., Hu, Y., Martins, B., Zhou, X. and Chen, F. (eds.), Proceedings of the 3rd ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery (GeoAI), Chicago, IL. pp. 15-23. (pdf)

Friday, July 26, 2019

Location-Based Social Simulation

At the upcoming 16th International Symposium on Spatial and Temporal Databases (SSTD) we have vision paper entitled "Location-Based Social Simulation" accepted. In the paper we discuss issues such as data sparsity and privacy concerns with using real world location-based social networks (LBSNs) like Foursquare and Yelp. To overcomes these issues, we describe how one can employ geospatial simulation (i.e. an agent-based model) to create artificial, but socially plausible LBSN data sets which overcomes some of the limitations with respect to LBSNs.

ABSTRACT:
Location-based social networks (LBSNs) have been studied extensively in recent years. However, utilizing real-world LBSN datasets in such studies has severe weaknesses: sparse and small datasets, privacy concerns, and a lack of authoritative ground-truth. Our vision is to create a large scale geosimulation framework to simulate human behavior and to create synthetic but realistic LBSN data that captures the location of users over time as well as social interactions of users in a social network. While existing LBSN datasets are trivially small, such a framework would provide the first source of very large LBSN benchmark data which would closely mimic the real world, containing high-fidelity information of location, and social connections of millions of simulated agents over several years of simulated time. Therefore, it would serve the research community by revitalizing and reshaping research on LBSNs by allowing researchers to see the (simulated) world through the lens of an omniscient entity having perfect data. These evaluations will guide future research allowing us to develop solutions to improve LBSN applications such as user-location recommendation, friend recommendation, location prediction, and location privacy.

KEYWORDS: Agent-based simulation, location-based social network, data generator, spatial network, human behavior

Full Reference: 
Kavak, H., Kim, J-S., Crooks, A.T., Pfoser, D., Wenk C. and Züfle, A (2019), Location-Based Social Simulation, Proceedings of the 16th International Symposium on Spatial and Temporal Databases, Vienna, Austria, pp 218-221. (pdf)

Update: Our paper was selected as runner-up  for best Vision Paper.






Thursday, November 08, 2018

Refugee Camps and Volunteered Geographical Information

Fig. 7. Stimulus-Awareness-Activism (SA2) framework
Previously we have posted on how one can use new sources of data  (e.g. Volunteered Geographical Information) to explore and understand the world around us, such as mass migration, urban form and function, or be used for the basis of a model. Continuing on with this research theme we recently had a paper published in PLoS ONE entitled: "News Coverage, Digital Activism, and Geographical Saliency: A Case Study of Refugee Camps and Volunteered Geographical Information."

In this paper we explore the relationship between news coverage (via Google news), search trends (via Google trends) and user edit contribution patterns in OpenStreetMap and  Wikipedia for refugee camps from around the world. Specifically we are interested in how news media coverage (and in particular digital media) impacts digital activism (i.e.  volunteers who contribute content to online communities). Based on our analysis we find that digital activism bursts tend to take place during periods of sustained build-up of public awareness deficit or surplus.

These findings are in line with two prominent mass communication theories: agenda setting and corrective action, and suggest the emergence of a novel Stimulus-Awareness-Activism (SA2) framework in today’s participatory digital age. We argue that this paper brings us one step closer to understanding the underlying mechanisms that drive digital activism in particular in the geospatial domain. Below you can read the abstract of the paper, see the refugee camps we studied and some of the results. At the bottom of the post we also provide the full reference and a link to the paper.

Abstract:
The last several decades have witnessed a shift in the way in which news is delivered and consumed by users. With the growth and advancements in mobile technologies, the Internet, and Web 2.0 technologies users are not only consumers of news, but also producers of online content. This has resulted in a novel and highly participatory cyber-physical news awareness ecosystem that fosters digital activism, in which volunteers contribute content to online communities. While studies have examined the various components of this news awareness ecosystem, little is still known about how news media coverage (and in particular digital media) impacts digital activism. In order to address this challenge and develop a greater understanding of it, this paper focuses on a specific form of digital activism, that of the production of digital geographical content through crowdsourcing efforts. Using refugee camps from around the world as a case study, we examine the relationship between news coverage (via Google news), search trends (via Google trends) and user edit contribution patterns in OpenStreetMap, a prominent geospatial data crowdsourcing platform. In addition, we compare and contrast these patterns with user edit patterns in Wikipedia, a well-known non-geospatial crowdsourcing platform. Using Google news and Google trends to derive a measure of thematic public awareness, our findings indicate that digital activism bursts tend to take place during periods of sustained build-up of public awareness deficit or surplus. These findings are in line with two prominent mass communication theories: agenda setting and corrective action, and suggest the emergence of a novel stimulus-awareness-activism framework in today’s participatory digital age. Moreover, these findings further complement existing research examining the motivational factors that drive users to contribute to online collaborative communities. This paper brings us one step closer to understanding the underlying mechanisms that drive digital activism in particular in the geospatial domain.

Figure 1. Study areas  (centroid location of camp).

Figure 5. OSM, Wikipedia, Google News, and Google Trends time series during a -/+4 months period around the strongest extremum point of each camp. The figures show that whereas OSM and Wikipedia entries tend to come in bursts, Google News and Trends display a more sustained type of activity.

Figure 6. The public awareness curve versus the cumulative OSM and Wikipedia edit activity during a -/+4 months period around the strongest extremum point of each camp. For camps such as Nyarugusu, OSM and Wikipedia bursts overlap with public awareness surplus. In other camps, such as Bidibidi, OSM edit activity bursts coincide with public awareness deficit.  

Full Reference: 
Mahabir, R., Croitoru, A., Crooks, A.T., Agouris, P. and Stefanidis, A. (2018), News Coverage, Digital Activism, and Geographical Saliency: A Case Study of Refugee Camps and Volunteered Geographical Information, PLoS ONE, 13(11): e0206825.   https://doi.org/10.1371/journal.pone.0206825 (pdf)

Friday, September 21, 2018

Exodus 2.0: Crowdsourcing Geographical and Social Trails of Mass Migration

Readers of the blog might know we have an interest in volunteered geographic information, social media and Web 2.0 technologies and how they can be used to explore urban systems. Recently however, we turned our focus on how such information and technologies can be used to explore and understand mass migrations.

To this end we recently had a paper published in the Journal of Geographical Systems entitled "Exodus 2.0: Crowdsourcing Geographical and Social Trails of Mass Migration". We adopt the term Exodus 2.0 to refer to this new migration paradigm in the digital age, whereby information is a commodity in the migration process.

Given the nature of migration processes, it is possible to explore them across two key dimensions: geographical and situational. The geographical dimension is associated with the physical migration pathways migrants take from a country of origin to a destination site (often through a number of intermediate “stop” sites). The situational dimension is associated with the social connectivity of moving migrant populations, the conditions on the ground, and the activities that take place as part of migration efforts (including the root conditions, proximate conditions and triggering events).
Factors that potentially cause refugee production and
 mass movement based on identified factors detailed by
Clark (1989) and Zottarelli (1998).
In the paper, we use the ongoing Syrian humanitarian crisis as a case study to to explore how the factors that potentially causes refugee production and mass movement  can be gleamed from new sources of data. Specifically, the potential of crowd-generated data—especially open data, volunteered geographic information and social media content (e.g. OpenStreetMap, Flickr, Twitter and Instagram)  to provide information about migration processes.  Through a series of case studies  we show how such data (when combined with more traditional data sources) offers a new lens to study such the geographical and situational dimensions of mass migration. Finally we discuss  how such data could be used to inform migration modeling. If we have not bored you yet and you are interested in finding out more about this line of inquiry, below we provide the abstract to the paper, some of the figures which go along with our analysis for studying the refugee production and movement. Finally, we also provide the full reference and a link to the paper. 

Abstract:
The exodus of displaced populations is a recurring historical phenomenon, and the ongoing Syrian humanitarian crisis is its latest incarnation. During such mass migration events, information is an essential commodity. Of particular importance is geographical (e.g., pathways and refugee camps) and social (e.g., refugee activities and networking) information. Traditionally, such information had been produced and disseminated by authorities, but a new paradigm is emerging: Web 2.0 and mobile computing technologies enable the involved stakeholder communities to produce, access, and consume migration-related information. The purpose of this article is to put forward a new typology for understanding the factors around migration and to examine the potential of crowd-generated data—especially open data and volunteered geographic information—to study such events. Using the recent wave of migration to Europe from the Middle East and northern Africa as a case study, we examine how migration-related information can be dynamically mined and analyzed to study the migrants’ pathways from their home countries to their destination sites, as well as the conditions and activities that evolve during the migration process. These new data sources can provide a deeper and more fine-grained understanding of the migration process, often in real-time, and often through the eyes of the communities affected by it. Nevertheless, this also raises significant methodological and technical challenges for their future use associated with potential biases, data quality issues, and data processing.

Keywords: Refugees, Forced migration, Humanitarian crisis, Volunteered geographic information, Crowdsourcing, Social media, GIS, Web 2.0.
Cumulative flow (2011–2015) illustrating Syrian forced migration to neighboring countries and other destination countries. Line thickness indicates increasing number of persons migrating.

Retweet network of geolocated Twitter microblogs that are discussing opinions, news and retweeting information related to “refugee” in multiple languages from May to August 2017.

A concept graph illustrating the associations between a keyword related to root factors of mass migration such as poverty (“welfare”) to other keywords, as they appear in our Twitter data corpus. The color of the node refers to specific themes: locations (green), actors (dark red), topics (red), entities and individuals (blue), concepts (white), and events (yellow). Red edges represent active associations between terms; gray edges represent inactive associations between terms.

An agent-based model of migration: top: the spatial environment, where the lines represent migration pathways, and the nodes represent number of migrants. Purple nodes represent final destination sites, red nodes show migrant deaths, and green nodes show migrants en route (source: Hu 2016).

Full Reference: 
Curry, T., Croitoru, A., Crooks, A.T. and Stefanidis, A. (in press), Exodus 2.0: Crowdsourcing Geographical and Social Trails of Mass Migration, Journal of Geographical Systems. DOI: https://doi.org/10.1007/s10109-018-0278-1 (pdf)

Wednesday, July 18, 2018

Online Vaccination Discussion and Communities in Twitter

Continuing on our work of exploring health related issues in social media, Xiaoyi Yuan and myself had a paper accepted at the 9th International Conference on Social Media and Society. In our paper entitled: "Examining Online Vaccination Discussion and Communities in Twitter"  we examined the communication patterns of anti-vaccine and pro-vaccine users on Twitter by studying the retweet network from 660,892 tweets related to the measles, mumps, and rubella (MMR) vaccine published by 269,623 users using supervised learning to identify clusters of users based on their opinions (i.e. a pro-vaccine, anti-vaccine, or neutral user). 

The overall methodology can be seen in Figure 1 and more details can be found in the paper. Our data was collected using the GeoSocial Gauge System, however, since tweets are short and their content diverse, the data corpus needed to be cleaned so that the tweets could then be converted to features (e.g., unigrams or bigrams). After which we were able to use such features for training a variety of classifiers (i.e., logistic regression, support vector machine (linear and non-linear kernel), k-nearest neighbors, nearest centroid, and Naïve Bayes) to identify opinion groups. After this, we moved from on from identifying each user’s opinion to construct a retweet network in order to understand how in-group and cross-group communicate in the committees detected via retweet network. By carrying out this analysis we discovered that pro- and anti-vaccine users retweet predominantly from their own opinion group, while users with neutral opinions are distributed across communities. Below you can read our abstract, see some results from our study and the full reference (and link) to the paper.


Figure1: Steps used in our study to unveil the communication patterns of pro-vaccine and anti-vaccine users on Twitter
 Abstract:
Many states in the US allow a “belief exemption” for measles, mumps, and rubella (MMR) vaccines. People’s opinion on whether or not to take the vaccine could have direct consequences in public health— once the vaccine refusal of a group within a population is higher than what herd immunity can tolerate, a disease can transmit fast causing large scale of disease outbreaks. Social media has been one of the dominant communication channels for people to express their opinions of vaccination. Despite governmental organizations’ effects of disseminating information of vaccination benefits, anti-vaccine sentiment is still gaining its momentum, especially on social media. This research investigates the communicative patterns of anti-vaccine and pro-vaccine users on Twitter by studying the retweet network from 660,892 tweets related to MMR vaccine published by 269,623 users after the 2015 California Disneyland measles outbreak. Using supervised learning, we classified the users into anti-vaccination, neutral to vaccination, and pro-vaccination groups. Using a combination of opinion groups and retweet network structural community detection, we discovered that pro- and anti-vaccine users retweet predominantly from their own opinion group, while users with neutral opinions are distributed across communities. For most cross-group communication, it was found that pro-vaccination users were retweeting anti-vaccination users than vice-versa. The paper concludes that anti-vaccine Twitter users are highly clustered and enclosed communities, and this makes it difficult for health organizations to penetrate and counter opinionated information. We believe that this finding may be useful in developing strategies for health communication of vaccination and overcome some the limits of current strategies.

Key Words: Anti-vaccine movement, Twitter, social media, opinion classification
Figure 2: Network visualizations of the four largest communities. A: is colored by the belonging to a specific structural community and; B: is colored by belonging to opinion groups

Figure 3: Distributions of opinion groups in the four largest structural community

Full Reference:
Yuan, X. and Crooks, A.T. (2018), Examining Online Vaccination Discussion and Communities in Twitter, Proceedings of the 9th International Conference on Social Media and Society, Copenhagen, Denmark, pp 197-206. (pdf)

Update: Our paper was selected as best paper at the conference.