Showing posts with label Twitter. Show all posts
Showing posts with label Twitter. Show all posts

Wednesday, May 27, 2026

New Paper: Exploring Fear in Urban Environments

In the past we have written about how we have used social media to study a plethora of topics with respect the the form and function of cities among many other things. But one thing we have not explored is fear and more specifically fear of crime and how this can be mined through geosocial media

This has now changed with a new paper entitled "Exploring Fear in Urban Environments: Place and Space Analysis of Social Media Data" which has recently been published in Applied Geography.  In this paper, Ying Zhou and myself extract fear related posts from social media and examine the places and spaces where people experience fear, as well as the factors that contribute to it in New York City. 

We do this by utilizing Natural Language Processing (NLP) techniques for sentiment and text analysis, including a RoBERTa-based emotion classification model and the BERTopic model for topic modeling. The former model narrowed the raw data to those with the dominant emotion of fear, and the latter analyzed space- and place-related features that contribute to the fear sentiment. Then, the selected social media data were analyzed using spatial clustering methods (i.e., Hotspot Analysis (Getis-Ord Gi*) and Local Moran’s I) and compared with urban crime data for weekly trends and spatial patterns. As such the paper has the following research objectives:
  1. exploring places where people expressed fear through social media; 
  2. making comparisons between safety-related fear and crime from the perspective of both time and space; 
  3. extracting urban environmental and social features that lead to fear.

If this sounds of interest, and you wish to find out more with respect to our findings, below you can read the abstract to the paper, see some of the figures which describe our research methodology and results while at the bottom of the post you can find a link to the paper itself. Finally the code we utilized in the paper can be found at https://osf.io/y7xfc/overview.

Abstract:

One goal of creating livable cities is to enhance public safety. While previous research in urban studies has focused on correlations between physical environments and crime, it has typically relied on criminal statistics. However, fear of crime is an emotional response to perceived risks rather than a direct reflection of crime levels, so it cannot be analyzed solely by crime data. Additionally, urban planning today has gradually shifted its focus from a top-down to a bottom-up approach, making it essential to understand and foster spaces where residents feel safe. This research examines the spaces and places where people experience fear, as well as the factors that contribute to it, in New York City. We utilized social media data to gather people’s expressions of the city and identified posts expressing fear emotion using the RoBERTa-based model and a rule-based classifier. Then, the selected social media data and crime were compared temporally by weekly trends and spatially by clustering methods (i.e., Hotspot Analysis (Getis-Ord Gi*) and Local Moran’s I). The results show that their temporal and spatial patterns partially have limited alignment. To delve into the origins of fear, we extend our analysis by adopting BERTopic to identify topics and summarize them into themes (e.g., places, transportation, people, others) to understand the bottom-up emergence of fear, thereby informing a people-centered approach to research on urban issues. 

Keywords: Social media; Natural language processing; Sentiment analysis; Urban environment.

Methodology framework.

An example of textual analysis on fear-related tweets: from machine-generated topics to human-interpreted themes describing fear in NYC.

Weekly trends comparison between safety-related fear and violent crime.

Clustering features analysis by the method of hotspot analysis (Getis-Ord Gi∗).

Full Reference: 

Zhou, Y. and Crooks, A.T. (2026), Exploring Fear in Urban Environments: Place and Space Analysis of Social Media Data, Applied Geography, 192: 104051 (pdf)

Tuesday, May 13, 2025

Crowdsourcing dust storms utilizing social media data

In the past we have explored how social media can be used to delineate earthquakes, study human-wildlife interactions, understand urban morphology, urban smells or  locating wildfires among many other things. 

Keeping with the last topic (i.e., locating things), in a new paper published in GeoJournal entitled "Crowdsourcing dust storms in the United States utilizing social media data," Stuart EvansFestus Adegbola and myself explore how we can use X (formerly Twitter) and Flickr  to source observations of windblown dust. 

As such the paper demonstrates how social media data can act as supplementary source for dust events monitoring and captures the seasonal trends of such events. Furthermore, the paper highlights the potential of using crowdsourced data for the often overlooked field of dust monitoring that has substantial health and economic impacts. 

If this sounds of interest, below we provide the abstract to the paper along with some figures which showcase our methodology and comparison with National Weather Service dust advisories and VIIRS satellite data. At the bottom of the post, you can find the full reference to the paper along with a link to it. 

Abstract: 

Dust storms and other dust events are natural phenomena characterized by strong winds carrying large amounts of fine particles which have significant environmental and human impacts. However, capturing the occurrence of such phenomena is a challenge. Previous studies have limitations due to available data, especially regarding short-lived, intense dust storms and events that are not captured by observing stations and satellite instruments. In recent years, the advent of social media platforms has provided a unique opportunity to access vast amounts of crowdsourced data. This paper explores the utilization of Flickr and X (Twitter) data to study dust event occurrences within the United States and their correlation with National Weather Service (NWS) advisories. The work ascertains the reliability of using crowdsourced data as a supplementary source for dust events monitoring. Our analysis of Flickr and X indicates that the Southwest region is most susceptible to dust events, with Arizona leading in the highest number of occurrences. On the other hand, the Great Plains show a scarcity of crowdsourced data related to dust events, which can be attributed to the sparsely populated nature of the region. Furthermore, seasonal analysis reveals that dust events are prevalent during the Summer months followed by Spring. These results are consistent with previous traditional studies that did not use social media of dust occurrences in the U.S., and Flickr-identified images of dust events show substantial co-occurrence with regions of NWS dust warnings. This paper highlights the potential of using crowdsourced data for the often overlooked field of dust monitoring that has substantial health and economic impacts.
Keywords: Dust storms, Crowdsourcing, Social media, Weather. 

 

Flowchart of our workflow
Selected posts retrieved from X showing active dust events.

Selected images retrieved from Flickr showing active dust events.

Map showing the distribution of flickr-identified dust event occurrences, X-identified dust event occurrences, National Weather Service dust advisories, including dust storm (DS) warnings and blowing dust (DU) advisories.

Seasonal cycle of dust events using social media metadata, the National Weather Service advisories, and the VIIRS satellite data.

Examples of social media identified dust events and satellite observations for the same day. Brown shaded pixels indicate locations Suomi-VIIRS observed dust particles. Any VTEC warnings issued by NWS for the location are shown after the date of each dust event, with HWW and DSW indicating High Wind Warning and Dust Storm Warning, respectively.

Full Referece: 
Adegbola, F., Crooks, A.T. and Evans, S.M. (2025). Crowdsourcing dust storms in the United States utilizing social media data. GeoJournal, 90(3), pp.1-18. Available at https://doi.org/10.1007/s10708-025-11359-9 (pdf)

Tuesday, April 22, 2025

Mapping the Invisible

Readers might of noticed that recently we have been exploring the use of street view images to explore cities or how we can utilize geosocial media to understand the form of function of cities, but one thing we have not explored is the role of smell and how it shapes peoples perceptions of urban spaces. However, in a new paper recently published in the Annals of the American Association of Geographers with Qingqing Chen, Ate Poorthuis we do just that. The paper is entitled "Mapping the Invisible: Decoding Perceived Urban Smells Through Geosocial Media in New York City" In the paper we use text mining techniques to tease out smell related information from over 56 million geolocated tweets which are then assigned to specific small categories (e.g., nature, food, waste) resulting in a new smellscape map for New York city. 

If this sounds of interest, below you can read the abstract to our paper, see our workflow and resulting smellscape map. While the the analysis steps, along with the smell dictionary used, are documented in the research code compendium at  https://figshare.com/s/8418d47cdc5c539b78ab. Finally at the bottom of the page, you can find the full reference and a link to the paper. 

Abstract:

Smells can shape people’s perceptions of urban spaces, influencing how individuals relate themselves to the environment both physically and emotionally. Although the urban environment has long been conceived as a multisensory experience, research has mainly focused on the visual dimension, leaving smell largely understudied. This article aims to construct a flexible and efficient bottom-up framework for capturing and classifying perceived urban smells from individuals based on geosocial media data, thus, increasing our understanding of this relatively neglected sensory dimension in urban studies. We take New York City as a case study and decode perceived smells by teasing out specific smell-related indicator words through text mining techniques from a historical set of geosocial media data (i.e., Twitter/X). The data set consists of more than 56 million data points sent by more than 3.2 million users. The results demonstrate that this approach, which combines quantitative analysis with qualitative insights, can not only reveal “hidden” places with clear spatial smell patterns, but also capture elusive smells that might otherwise be overlooked. By making perceived smells measurable and visible, we can gain a more nuanced understanding of smellscapes and people’s sensory experiences within the urban environment. Overall, we hope our study opens up new possibilities for understanding urban spaces through an olfactory lens and, more broadly, multisensory urban experience research. 

Key Words: geosocial media, multisensory urban experiences, network analysis, New York City, smellscape, text mining, urban smells.

A framework of deriving perceived smells.

An overview of research workflow.

An overview of the six dominant overlapping smells across New York City using the weaving mapping method. The weaving map uses the concept of strands to represent attributes. Each strand here represents one specific smell category, with the intensity of the color changing based on the density of that smell category within each neighborhood (i.e., grid cells).

Full Reference: 

Chen, Q., Poorthuis A. and Crooks, A.T., (2025), Mapping the Invisible: Decoding Perceived Urban Smells through Geosocial Media in New York City, Annals of the American Association of Geographers, 115(6), 1444-1464. Available at https://doi.org/10.1080/24694452.2025.2485233. (pdf)

Saturday, December 14, 2024

AGU

This past week we attended the American Geophysical Union (AGU) Fall Meeting in Washington DC. At the AGU we presented two abstracts. 

The first follows on our work with respect to using synthetic populations within agent-based models. This work was with Na Jiang, Fuzhen Yin and Boyu Wang and entitled "A Framework for Populating Urban Digital Twins with Agents." Or more specially why digital twins need agents. Below you can see our abstract and a couple of figures showing our synthetic population workflow and how we integrate these into agent-based models.  

Abstract:

Over the last few years, considerable efforts have been placed in creating digital twins from diverse fields ranging from engineering to urban planning and many things in-between. These digital twins have benefited from the growth and availability of computational power and data. For example, in urban planning the growth of computational resources and the explosion of spatial data sources(e.g. remote sensing) has lead to the creation and widespread adoption of detailed virtual urban environments or urban digital twins. However, we would argue that many of such works emphasize only the physical infrastructure or the built environment of the city instead of considering the key actors of urban systems: the people who live in them. In this work we aim to remedy this by introducing a framework that utilizes agent-based modeling to add humans to such urban digital twins. This framework consists of two components: 1)synthetic populations generated with census data; and 2) pipeline of using the population datasets for agent-based modeling applications within the urban digital twins domain. To demonstrate the utility of this framework, we have representative applications that showcase how digital twins can be created to study various urban phenomena (e.g., evacuation scenarios, traffic congestion and disease transmission). By doing so, we believe this framework will benefit researchers wishing to build urban digital twins and to explore complex urban issues with realistic populations. 


Workflow of utilizing synthetic populations within agent-based models.
Examples of agent-based models utilizing our synthetic popuation.

In a different presentation, we return to how one can use social media to monitor the world around us, in this case dust storms. This work entitled "Mining unconventional data sources: creating a social media-based catalog of dust events in the Western US" is collaboration with Stuart Evans and Festus Adegbola. Generally speaking we explore how social media has the potential for a new unconventional source of observations of windblown dust. If this sounds of interest, below you can read the abstract to the paper and see the visual overlap between social media posts about dust events and official National Weather Service (NWS) dust storm warning coverage. 

Abstract 

Complete observations of dust events are difficult, as dust’s spatial and temporal variability means satellites may miss dust due to overpass time or cloud coverage, while ground stations may miss dust due to not being in the plume. As a result, an unknown number of dust events go unrecorded in traditional datasets. Dust’s importance both for atmospheric processes and as a health and travel hazard makes detecting dust events whenever possible important, and in particular, studies of the health impacts of dust are limited by detailed exposure information, i.e. where is there dust and when. In recent years, social media platforms have provided an opportunity to access vast user-generated data. This research utilizes geotagged Flickr and Twitter posts referencing dust in the western US, and compares it to traditional datasets including blowing dust reports from the National Weather Service and satellite observations from Suomi-VIIRS. Results show that this unconventional dataset broadly recreates the observed spatial and seasonal distributions of dust. Daily analysis of the locations of the social media posts creates a novel catalog of dust events in the western US that can be used for further research. While this catalog is necessarily incomplete, it nonetheless provides a complementary list of events to those detected by traditional means. Analysis of individual events in this catalog shows that social media captures many dust events that previously went undetected by traditional datasets.


References:

Crooks, A.T., Jiang, N., Yin, F. and Wang, B. (2024), A Framework for Populating Urban Digital Twins with Agents, American Geophysical Union (AGU) Fall Meeting, 9th–13th December, Washington, DC. (pdf)

Evans, S., Adegbola, F. and Crooks, A.T. (2024), Mining Unconventional Data Sources: Creating a Social Media-based Catalog of Dust Events in the Western US, American Geophysical Union (AGU) Fall Meeting, 9th–13th December, Washington, DC. (pdf)

Friday, June 07, 2024

A comparison of social surveys and social media for vaccine hesitancy

In the past we have explored various ways to explore vaccine hesitancy and keeping with this theme we have a new paper published in PLOS ONE entitled "Understanding the determinants of vaccine hesitancy in the United States: A comparison of social surveys and social media" with Kuleen Sasse, Ron Mahabir, Olga Gkountouna and Arie Croitoru

In the paper we use social, demographic and economic (e.g., US Censusvariables to predict COVID-19 vaccine hesitancy levels in the ten most populous US metropolitan statistical areas (MSAs). By using  machine learning algorithms (e.g., linear regression, random forest regression, and XGBoost regression) we compare a set of baseline models that contain only these variables with models that incorporate survey data and social media (i.e., Twitter) data separately. 

We find that different algorithms perform differently along with variations in influential variables such as age, ethnicity, occupation, and political inclination across the five hesitancy classes (e.g., “definitely get a vaccine”, “probably get a vaccine”, “unsure”, “probably not get a vaccine”, and “definitely not get a vaccine”).   Further, we find that the application of the models to different MSAs yields mixed results, emphasizing the uniqueness of communities and the need for complementary data approaches. But in summary, this paper shows social media data’s potential for understanding vaccine hesitancy, and tailoring interventions to specific communities. If this sounds of interest, below we provide the abstract to the paper along with our mixed methods matrix, data sources used and the results from the various MSAs. At the bottom of the post, you cans see the full reference and the link to the paper so you can read more if you so desire. 

Abstract:
The COVID-19 pandemic prompted governments worldwide to implement a range of containment measures, including mass gathering restrictions, social distancing, and school closures. Despite these efforts, vaccines continue to be the safest and most effective means of combating such viruses. Yet, vaccine hesitancy persists, posing a significant public health concern, particularly with the emergence of new COVID-19 variants. To effectively address this issue, timely data is crucial for understanding the various factors contributing to vaccine hesitancy. While previous research has largely relied on traditional surveys for this information, recent sources of data, such as social media, have gained attention. However, the potential of social media data as a reliable proxy for information on population hesitancy, especially when compared with survey data, remains underexplored. This paper aims to bridge this gap. Our approach uses social, demographic, and economic data to predict vaccine hesitancy levels in the ten most populous US metropolitan areas. We employ machine learning algorithms to compare a set of baseline models that contain only these variables with models that incorporate survey data and social media data separately. Our results show that XGBoost algorithm consistently outperforms Random Forest and Linear Regression, with marginal differences between Random Forest and XGBoost. This was especially the case with models that incorporate survey or social media data, thus highlighting the promise of the latter data as a complementary information source. Results also reveal variations in influential variables across the five hesitancy classes, such as age, ethnicity, occupation, and political inclination. Further, the application of models to different MSAs yields mixed results, emphasizing the uniqueness of communities and the need for complementary data approaches. In summary, this study underscores social media data’s potential for understanding vaccine hesitancy, emphasizes the importance of tailoring interventions to specific communities, and suggests the value of combining different data sources.
Mixed methods matrix showing the data, processing, and model development steps used in our study.

Data sources used in our study.

MSA model performance (Bolded adjusted R2 values represent the best performing model for each modeling technique and MSA).

Tuesday, May 25, 2021

Achieving Situational Awareness with Geolocated Social Media

Tuning back to our work on geosocial analysis we (Xiaoyi Yuan, Ron Mahabir, Arie Croitoru and myself) recently had a paper published in GeoJournal entitled "Achieving Situational Awareness of Drug Cartels with Geolocated Social Media." 
 
The overarching objective of this paper is to develop an approach that would enable the extraction of potentially relevant situational awareness-related information from geolocated raw data streams (in this example we use Twitter). We accomplish this goal by focusing on Named Entities (NEs) related to drug cartels rather than the raw text as a whole. Specifically, our analysis is performed on the NEs by first extracting them and then clustering them to identify relevant concepts/themes (using TextRazor). This approach gives rise to themes that can then be assessed for temporal and spatial patterns based on frequency in order to gain underlying insights into drug cartels. If is of interest to you below we provide the abstract to the paper, a diagram of our workflow and a sample of our results along with the link to the paper. Also the complete code for the analysis and results is available at https://bitbucket.org/xiaoyiyuan/cartel.
 

Abstract: Using geolocated tweets to achieve situational awareness is an often researched topic in disaster and emergency management. However, little has been done in the area of drug cartels, which, as transnational crime organizations, continue to pose great risk to the stability and safety of our communities. This paper made an initial effort in using geolocated social media (specifically Twitter) to achieve situational awareness of drug cartels through temporal and spatial analysis of derived named entity clusters. The results show that detecting peaks in the time series of frequently occurring entity clusters enabled the tracking of important events in public discourse surrounding drug cartels. Correlations between time series also provided valuable insights into the synchronicity between different events. Further examining the spatial distribution of key events for different countries, we identify thematic hotpots of public discourse on cartel activity. Our methodology also addresses issues of language ambiguity when working with noisy social media data in order to achieve situational awareness on drug cartels.

Keywords: Cartels, Social Media, Situational Awareness and Temporal and Spatial Analysis.

The workflow of achieving situational awareness of drug cartels using geolocated tweets.

Tweet and entity counts by language and geolocation.

An example of tweets of high frequency on peak day in Venezuela

Heat maps of frequencies of a Cluster for Day 14 and Days 18-21.


Full Reference:
Yuan, X., Mahabir, R., Crooks, A.T. and Croitoru, A. (2021), Achieving Situational Awareness of Drug Cartels with Geolocated Social Media, GeoJournal. DOI: https://doi.org/10.1007/s10708-021-10433-2 (pdf)

Tuesday, September 01, 2020

Beyond Words: Comparing Structure, Emoji Use, and Consistency Across Social Media Posts

Continuing our work on Emojis, at the forthcoming International Conference on Social Computing, Behavioral-Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation (or SBP-BRiMS for short) we (Melanie Swartz, Arie Croitoru and myself) have a paper entitled "Beyond Words: Comparing Structure, Emoji Use, and Consistency Across Social Media Posts." In the paper we introduce and demonstrate a language-agnostic methodology to characterize structures of content and emoji use within a document (in this case a tweet), measure consistency of structures across a set of documents, and cluster documents and users with similar patterns and behavior. Using a corpus of 44 million tweets collected in October and November 2018 related to the 2018 U.S. midterm elections based on keywords, hashtags, and user accounts associated with candidates or political parties we were able to gain insights into the unique or shared structures of communication styles and emoji use of over 3.3 million unique users and user roles such as journalists, bots and others. If this sounds of interest to you, below we provide the abstract to the paper, some tables and figures of the our findings along with the full reference and a link to the paper. Furthermore, if you are interested in extending this work to other areas, Melanie has made the code available at https://github.com/msemoji/.

Abstract
Social media content analysis often focuses on just the words used in documents or by users and often overlooks the structural components of document composition and linguistic style. We propose that document structure and emoji use are also important to consider as they are impacted by individual communication style preferences and social norms associated with user role and intent, topic domain, and dissemination platform. In this paper we introduce and demonstrate a novel methodology to conduct structural content analysis and measure user consistency of document structures and emoji use. Document structure is represented as the order of content types and number of features per document and emoji use is characterized by the attributes, position, order, and repetition of emojis within a document. With these structures we identified user signatures of behavior, clustered users based on consistency of structures utilized, and identified users with similar document structures and emoji use such as those associated with bots, news organizations, and other user types. This research compliments existing text mining and behavior modeling approaches by offering a language agnostic methodology with lower dimensionality than topic modeling, and focuses on three features often overlooked: document structure, emoji use, and consistency of behavior.
Keywords: Data Mining, Social Media, Emojis, User Behavior Modeling.
Emoji attributes

Most common content structures with emojis for non-retweets.

Clusters of users with similar behavior across two factors in non-retweets (left) and retweets (right) Colors indicate cluster assignments.

Full Reference: 
Swartz, M., Crooks, A.T. and Croitoru, A. (2020), Beyond Words: Comparing Structure, Emoji Use, and Consistency Across Social Media Posts, in Thomson, R., Bisgin, H., Dancy, C., Hyder, A. and Hussain, M. (eds), 2020 International Conference on Social Computing, Behavioral-Cultural Modeling & Prediction and Behavior Representation in Modeling and Simulation, Washington DC., pp 1-11. (pdf)

 

Wednesday, July 22, 2020

Diversity from Emojis and Keywords in Social Media

Building on our initial work on emojis  use and and how one can carry out a systematic comparison of emojis across individual user profiles and communication patterns within social media, we have a new paper entitled: "Diversity from Emojis and Keywords in Social Media" which was presented at the 11th International Conference on Social Media and Society

In the paper we present a novel method using a diversity language model to associate diversity related attributes to social media user accounts and content by analyzing the emojis and keywords used (in this case from Twitter). We used this diversity language model to shed light on the groups of social media users and content with similar diversity attributes related to American politics (specifically the 2018 U.S. midterm elections). Our results revealed topics of interest and patterns of social media engagement across political lines among the diverse populations that otherwise would not have been apparent if we only analyzed the key political campaign phrases and slogans (i.e. “Blue Wave” and “Make America Great Again”) without taking diversity into account.

For interested readers, below we provide the abstract to the paper along with some figures from the paper. These include our workflow for diversity analysis of social media content, a high level overview of our diversity language model. These are followed by some of our results. Specifically the presence of diversity keywords and emojis in user profiles, and the composition of users in our collection based on gender for two political campaigns. If this peaks you interest as the conferce was virtual we have also prepared a short movie of the paper. While at the bottom of the post you find the full reference to the paper along with a link to the paper itself.

Abstract:
Social media is a popular source for political communication and user engagement around social and political issues. While the diversity of the population participating in social and political events in person are often considered for social science research, measuring the diversity representation within online communities is not a common part of social media analysis. This paper attempts to fill that gap and presents a methodology for labeling and analyzing diversity in a social media sample based on emojis and keywords associated with gender, skin tone, sexual orientation, religion, and political ideology. We analyze the trends of diversity related themes and the diversity of users engaging in the online political community during the lead up to the 2018 U.S. midterm elections. Our results reveal patterns along diversity themes that otherwise would have been lost in the volume of content. Further, the diversity composition of our sample of online users rallying around political campaigns was similar to those measured in exit polls on election day. The diversity language model and methodology for diversity analysis presented in this paper can be adapted to other languages and applied to other research domains to provide social media researchers a valuable lens to identify the diversity of voices and topics of interest for the less-represented populations participating in an online social community.

Keywords: Social media, emoji, diversity, elections, political campaigns
Workflow for diversity analysis of social media content
Diversity Language Model
Presence of diversity keywords and emojis in user profiles
Composition of users in our collection based on gender for two political campaigns



Full Reference:

Swartz, M., Crooks, A.T. and Kennedy, W.G. (2020), Diversity from Emojis and Keywords in Social Media, in Gruzd, A., Mai, P., Recuero, R., Hernández-García, A., Lee, C.S., Cook, J., Hodson, J., McEwan, B and Hopke, J. (eds.), Proceedings of the 11th International Conference on Social Media & Society, Toronto, Canada, pp 92-100. (pdf)

Wednesday, June 10, 2020

New Paper: A Thematic Similarity Network Approach for Analysis of Places Using VGI

Building upon our work on volunteered geographical information (VGI) and ambient geographic information (AGI) and how such data (e.g. social media) can be used to understand place, Xiaoyi Yuan, Andreas Züfle and myself have a new paper entitled: "A Thematic Similarity Network Approach for Analysis of Places Using Volunteered Geographic Information" in the ISPRS International Journal of Geo-InformationIn this paper we use textual data from crowdsourced reviews originating with TripAdvisor and geo-located Twitter data and leverage this unstructured geographical information to comprehend the complexity of places at scale. Specifically we explore the connectedness and relationships of places through thematic (i.e., topical) similarity networks using Manhattan, New York as a case study. If such work sounds of interest to you, below we provide the abstract to the paper in order for you to gain a greater understanding of work, along with some figures that show our workflow and how communities where connected, before presenting some of our results. Finally at the bottom of the post, the full reference and a link to the paper is provided.  For those interested in extending or utilizing this work. The python code for presented in our analysis is available at: https://bitbucket.org/xiaoyiyuan/network_vgi/

Abstract:
The research presented in this paper proposes a thematic network approach to explore rich relationships between places. We connect places in networks through their thematic similarities by applying topic modeling to the textual volunteered geographic information (VGI) pertaining to the places. The network approach enhances previous research involving place clustering using geo-textual information, which often simplifies relationships between places to be either in-cluster or out-of-cluster. To demonstrate our approach, we use as a case study in Manhattan (New York) that compares networks constructed from three different geo-textural data sources --TripAdvisor attraction reviews, TripAdvisor restaurant reviews, and Twitter data. The results showcase how the thematic similarity network approach enables us to conduct clustering analysis as well as node-to-node and node-to-cluster analysis, which is fruitful for understanding how places are connected through individuals’ experiences. Furthermore, by enriching the networks with geodemographic information as node attributes, we discovered that some low-income communities in Manhattan have distinctive restaurant cultures. Even though geolocated tweets are not always related to place they are posted from, our case study demonstrates that topic modeling is an efficient method to filter out the place-irrelevant tweets and therefore refining how of places can be studied.

Keywords: Geo-Textual Data, Volunteered Geographic Information, Crowdsourcing, Similarity Network Analysis, Topic Modeling

Work flow from data input to the construction of the thematic similarity network and analysis (i.e., community detection and unique nodes discovery).

A stylized network demonstrating the process of community detection from a fully-connected similarity network.


Network visualization of all communities from the thematic similarity networks with major communities highlighted. Only the major communities are shown on the map for the sake of clarity. Major communities in Network visualization and mapping for each network are colored the same and thus the legend applies for both.


Two examples of communities with boundary nodes and their respective topics.

Full Reference:
Yuan X., Crooks, A.T. and Züfle, A. (2020), A Thematic Similarity Network Approach for Analysis of Places Using Volunteered Geographic Information, ISPRS International Journal of Geo-Information,  9(6), 385, https://doi.org/10.3390/ijgi9060385. (pdf)

Friday, January 31, 2020

The Interplay Between the Media and the Public in Mass Shootings

Continuing our work on shootings we recently had a paper published in Criminology and Public Policy entitled: "Responses to Mass Shooting Events: The Interplay Between the Media and the Public." However, here we do not look at bots but instead explore the how the public responds to mass shooting events (e.g. Las Vegas, Sutherland Springs, Marshall County, Parkland, Santa Fe), by seeking additional information or exchanging opinions about them in media coverage (e.g. newspaper articles via LexisNexis) and through online sources of information (e.g. Google Trends, Wikipedia and Online Social Networks (i.e. Twitter)). 

Overall, our results show discernible patterns in both time and space in the public’s online information seeking activities after a mass shooting. In addition we find discernible online information seeking patterns in geographic space, with a focal area of interest in the state in which the shooting event occurs, surrounded by a region of reduced interest. This finding further suggests that online information seeking activities are driven, at least in part, by geographic proximity to mass shooting events.

If you wish to find out more about this research, below we provide the summary and policy implication to the paper along with some figures from our methodology (e.g., how we go about analyzing temporal and geographical trends) and some of the results. Finally at the bottom of the post we provide the full reference and a link to the paper.

Abstract:
Research Summary: Public mass shootings tend to capture the public’s attention and receive substantial coverage in both traditional media and online social networks (OSNs) and have become a salient topic in them. Motivated by this, the overarching objective of this paper is to advance our understanding of how the public responds to mass shooting events in such media outlets. Specifically, it aims to examine whether distinct information seeking patterns emerge over time and space, and whether associations between public mass shooting events emerge in online activities and discourse. Towards this objective, we study a sequence of five public mass shooting events that have occurred in the United States between October 2017 and May 2018 across three major dimensions: the public’s online information seeking activities, the media coverage, and the discourse that emerges in a prominent OSN. To capture these dimensions, respectively, data was collected and analyzed from Google Trends, LexisNexis, Wikipedia Page views, and Twitter. The results of our analysis suggest that distinct temporal patterns emerge in the public’s information seeking activities across different platforms, and that associations between an event and its preceding events emerge both in the media coverage and in OSNs.
Policy Implication: Studying the evolution of discourse in OSNs provides a valuable lens to observe how society’s views on public mass shooting events are formed and evolved over time and space. The ability to analyze such data allows tapping into the dynamics of reshaping and reframing public mass shooting events in the public sphere and enable it to be closely studied and modeled. A deeper understanding of this process, along with the emerging associations drawn between such events, can then provide policy and decision-makers with opportunities to better design policies and communicate the significance of their goals and objectives to the public.
A framework for the analysis of temporal and geographical trends .

The analysis processes of Twitter and LexisNexis data.

Geographic patterns in online search activity in Google Trends for the five events in our study.

Chronologically ordered Google Trends search activity (a, left) and Wikipedia page views (b, right). Each vertical solid black line marks the occurrence of one of four shooting events examined in the analysis (as indicated by the line label).

Mentions of prior events during the first approximately 1-month period following each event in each of the events studied. (a) Sutherland Springs, (b) Marshall County, (c) Parkland, (d) Santa Fe.

Full Reference:
Croitoru, A., Kien, S., Mahabir, R., Radzikowski, J., Crooks, A.T., Schuchard, R., Begay, T., Lee, A., Bettios, A. and Stefanidis, A. (2020), Responses to Mass Shooting Events: The Interplay Between the Media and the Public, Criminology and Public Policy, 19(2): 335–360. (pdf)

Thursday, January 30, 2020

Comparison of Emoji Use in Names, Profiles, and Tweets

In most of our work to date with respect to exploring social media, we have only looked at the text or images from online social media platforms (e.g. Twitter and Flickr) and excluded  emojis from the analysis. However, this has now changed with a new paper co-authored with  Melanie Swartz entitled "Comparison of Emoji Use in Names, Profiles, and Tweets" which will be presented at he Eighth IEEE International Workshop on Semantic Computing for Social Networks and Organization Sciences in conjunction with 14th IEEE International Conference on Semantic Computing

In the paper we discuss how emoji use is becoming more and more popular by users of online social networking sites as they can be an effective way to express sentiment, sarcasm or feelings which are not easily conveyed as text. However, limited research has focused on analysis of the behavior of emoji use or how to compare emoji use across users or documents. To overcome this limitation, in this paper: (1) we present a methodology to extract, aggregate, and compare emoji use across a collection of documents based on Unicode emoji category and subcategories, (2) we present a baseline of statistics of emoji use in user names, profile descriptions, and tweets, and (3) we compare emoji use as categories and subcategories between users and content a user shares in the user name, profile description, retweets and non-retweets.

By considering this semantic grouping of emojis, we move the research on emojis beyond just comparing individual emojis and broad aggregations. In applying our methodology to a set of 44 million tweets and over 3 million user profiles, we find that differences in emoji use emerged based on document type (i.e., user names, profile descriptions, retweets, and non-retweets). As such, our work offers a new lens to study and compare forms of self expression across a variety of digital media content types. If you wish to find out more about this work, below we present the abstract to the paper, our workflow that allows for emoji comparison and some results. Finally at the bottom of the page we provide the full reference and a link to the paper.

Abstract 
Online social networking applications are popular venues for self-expression, communication, and building connections between users. One method of expression is that of emojis, which is becoming more prevalent in online social networking platforms. As emoji use has grown over the last decade, differences in emoji usage by individuals and the way they are used in communication is still relatively unknown. This paper fills this gap by comparing emoji use across users and collectively in user names, profiles, and in original and re-shared content. We present a methodology that enables comparison of semantically similar emojis based on Unicode emoji categories and subcategories. We apply this methodology to a corpus of over 44 million tweets and associated user names and profiles to establish a baseline which reveals differences in emoji use in user names, profile descriptions, non-retweets, and retweets. In addition, our analysis reveals emoji super users who have a significantly higher proportion and diversity of emoji use. Our methodology offers a novel approach for summarizing emoji use and enables systematic comparison of emojis across individual user profiles and communication patterns, thus expanding methods for semantic analysis of social media data beyond just text.
Keywords: emoji; social media analytics; content analysis; online social networks.

Workflow for emoji comparison.

Proportion of emoji use in profiles, names, retweets, and non-retweets, ordered by category.

Proportion of emoji use by subcategory.

Top emoji for each communication type.

Reference:
Swartz, M. and Crooks, A.T. (2020), Comparison of Emoji Use in Names, Profiles, and Tweets, The Eighth IEEE International Workshop on Semantic Computing for Social Networks and Organization Sciences: From User Information to Social Knowledge, San Diego, CA. (pdf)