Showing posts with label Machine Learning. Show all posts
Showing posts with label Machine Learning. Show all posts

Monday, August 17, 2026

New Paper: Online Interactions, Mutual Assistance and the Power of Weak Ties

In the past we have written about disasters and how one can model peoples reactions to them or how one can mine social media or mobility data to explore people's responses to them. In a new paper published in the Annals of the American Association of Geographers, Fuzhen Yin, Lucie Laurian and Emmanuel Boamah and myself continue this line of research. Specifically we explore how people exchanged resources and coordinated mutual aid via Facebook during the 2022 Buffalo Blizzard.

The paper itself is entitled "Surviving the Buffalo Blizzard: Online Interactions, Mutual Assistance and the Power of Weak Ties" In the paper we describe how we manually collected blizzard-related conversations from two Facebook groups "Buffalo S.T.O.R.M" and "Buffalo Blizzard". After which we utilize machine-learning (e.g., support vector machines) to identify mutual-aid messages which were then categorized them into four groups: aid requests, aid offers, emotional support and other. From which we then constructed a social network of users interactions during the blizzard to identify aid requests, aid offers, and emotional support messages during a time of crisis. As such our study contributes to the growing literature on human dynamics by examining spontaneous mutual aid during the blizzard and highlights how online mutual assistance operated through ‘phygital’ (physical–digital) integration.

If this sounds of interest, and you wish to find out more with respect to our findings, below you can read the abstract to the paper, see some of the figures which describe our research methodology and results while at the bottom of the post you can find a link to the paper itself.

Abstract: 
In December 2022, Buffalo, New York, experienced a once-in-a-generation blizzard. The four-day lake-effect snow accompanied by storm-force winds knocked down power lines, halted emergency services in several towns and resulted in forty-seven fatalities of residents who lost heat and power or were trapped in the snow. In response to the storm, Buffalonians demonstrated strong solidarity through quickly self-organized Facebook groups to exchange resources and coordinate mutual aid. Our study examines the emergence of grassroots mutual assistance through online–offline interactions and its impact on resilience in the physical world. We manually collected blizzard-related conversations, used machine learning to identify mutual-aid messages, and applied social network analysis to examine users’ interactions. Our findings reveal that Facebook users delivered life-saving assistance through online conversations involving requesting and offering practical, informational, and emotional support. The Facebook blizzard communities developed networks of weak ties that expanded access to vital resources and facilitated the flow of information and materials among disconnected residents. This research highlights virtual spaces as digital urban commons where strangers can benefit from emerging social capital during crises. It also offers insights for emergency management agencies seeking collaborations with grassroots online communities to develop formal–informal mutual aid strategies for future crises.

Keywords: Mutual aid, winter storm/blizzard, crisis informatics, weak ties, social network analysis, machine learning, social media.

Diagram of analysis workflow.

Geographical distribution of places mentioned in the Blizzard Facebook groups. (a) Kernel density map based on precise point locations. (b) General areas at three spatial levels: neighborhoods in Buffalo (purple), cities and towns (blue), and county subdivisions (green). The orange dashed line delineates the boundary of the kernel density map. Line widths and label sizes are proportional to each area’s prevalence in the online discussions.
Classifying mutual aid messages into four categories: request for support, aid offers, emotional support and other (n=9,599).
Tripartite message network capturing the information flow from posts to comments, from comments to replies, and within replies.

Social network of mutual aid interactions. (1) Users’ degree centrality with a log-transformed x-axis. (2) Users’ betweenness centrality with a log-transformed x-axis. (3) Social networks of users’ online interactions, highlighting four clusters: A, B, C, D. (4) Size distribution of detected communities. (5) Users’ composition in detected communities.

Full Reference: 

Yin, F., Laurian, L., Crooks, A.T. and Boamah, E. (2026), Surviving the Buffalo Blizzard: Online Interactions, Mutual Assistance, and the Power of Weak Ties, Annals of the American Association of Geographers. https://doi.org/10.1080/24694452.2026.2707171 (pdf)


Monday, March 31, 2025

AAG 2025 Talks

As the AAG has just wrapped up I thought I would write brief (well actually quite long) post on the talks that I was involved with at the conference. These talks would not have been possible without the many great students and colleagues who I have been collaborating with over time. Below you will find a brief summary of the talks and if any sound interesting, please reach out and we can give you more details. 

First up (in order in which they were presented) was "Utilizing Streetview Images for Mapping Building Attributes with ChatGPT" with Qingqing Chen and Linda See. In this talk we discussed how multimodal Large Language Models are giving us a new way to study cities, in the sense, lowering the boundary for information extraction. Using ChatGPT and street view images from Mapillary as an example, we showed how one can extract building age, usage (e.g., commercial, mixed use, residential) and estimate building height  which could all be used to inform urban climate models which require detailed information on buildings. 

Abstract: 

With increasing rates of urbanization, many challenges are emerging regarding sustainability such as the energy usage of buildings. Coinciding with this is the growing attention of urban climate models for energy demand estimation and climate adaptation strategies. However, the applicability of these models is constrained by the lack of detailed urban surface information. Therefore, creating comprehensive datasets that capture urban surface information at a granular scale is crucial for responding to our rapidly urbanizing world. Recent advancements in Large Language Models (LLMs) have opened new opportunities in urban studies, offering accessible methods for information extraction. In this talk we explore the feasibility of ChatGPT to extract building attributes from images. Taking New York City as a case study, we collect building images from Mapillary and process them through ChatGPT by posing specific questions to extract building attributes (e.g., height, functions, age). These attributes are then compared with authoritative data. The proposed method helps address the current dearth of fine-grained surface data on urban issues, therefore enhancing the accuracy and utility of urban climate models. Overall, this study demonstrates the practical applications of ChatGPT in geographic knowledge extraction, advancing the understanding of LLMs in geographic contexts, and more broadly to the discourse on Artificial Intelligence (AI) in urban modeling and climate science. 

Keywords: Buildings, ChatGPT, Large Language Models (LLMs), Mapillary, Street View Images (SVI), GeoAI.

Example of Workflow.

Reference

Chen, Q., See, L. and Crooks, A.T. (2025), Utilizing Streetview Images for Mapping Building Attributes with ChatGPT, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)

This was followed by a talk by lead by Qingqing Chen entitled "Multi-sensory Experiences: The Connection Between the Smell and Vision in Understanding Urban Environments" where we explored to what extent can visual data from street view imagery be used as a proxy for capturing large-scale urban smell perceptions when compared to geosocial media. Such as what visual cues evoke specific smell perceptions.

Abstract:

Smell is a crucial transversal sense, which bridges the tangible aspects of urban environments, such as exhaust and garbage, with their intangible impacts on emotions, social interactions and well-being. Despite its crucial role in our everyday life, many urban studies primarily focus on the visual dimension, potentially introducing biases in our understanding of urban spaces. This research transcends this visual-centric bias by integrating the olfactory perceptions to investigate the nuanced relationship between smell and vision in urban environments. Specifically, we utilize advanced semantic segmentation to extract visual elements from street view imagery (i.e., Mapillay) and apply casual forest analysis to examine their causal effects on smell expectations recorded from human participants. These expectations, often tied to personal experiences and/or cultural associations, are compared with real-environment smell experiences derived from geosocial media (i.e., Twitter/X). The results show that visual cues can predict smells in straightforward urban settings, such as small parks or less densely populated areas. However, in complex urban environments, the predictive power of visual cues diminishes as diverse and overlapping scents obscure specific smells, even in visually distinct areas. These findings underscore the importance of a multisensory approach in urban studies, enhancing our understanding of the complex interplay between sensory experiences and informing urban design strategies that integrate multiple senses to create more engaging and inclusive environments. This is especially important for individuals with sensory impairments, such as anosmia or visual impairments, who rely on other senses to compensate for their perception of urban environments. 

Keywords: Multi-sensory Experiences, Smell and Vision; Semantic Segmentation, Causal Effects, Geosocial Media, Street View Imagery (SVI).

Workflow

Reference: 

Chen, Q. and Crooks, A.T. (2025), Multi-sensory Experiences: The Connection Between the Smell and Vision in Understanding Urban Environments, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)

In the Geosimulation session that we organized, we had a talk entitled "Large Language Models for Conceptualizing, Designing, and Generating Agent-based Models" where Na Jiang, Boyu Wang and myself presented our work on exploring using multimodal Large Language Models (LLMs) to create age-based models. In the sense as modelers, we spend a lot of time developing and writing code and we were curious what could be done though the use of LLMs. 

To give a sense of what is possible, below is an example of using ChatGPT for creating a model from a published paper.



Abstract:
Large language models (LLMs) play an important role in AI-powered code assistants such as code completion, debugging, and documentation. Such models can be further fine-tuned on smaller amount of data for specific tasks, often with the improvement of performance compared to generic LLMs. However, such fine-tuning techniques are seldomly used in generating sophisticated agent-based models (ABMs), because they are often implemented as software that demands extra standards such as the “Overview, Design concepts, and Details” (ODD) protocol. This research examines how we can bridge this gap by utilizing LLMs in designing or conceptualizing, building, and running agent-based models in the form of user prompts. . In this work, two models are created to demonstrate the proposed method. Specifically, Sakoda’s checkerboard model of social interaction is created by LLM from explicit design and description through prompts. The other model stimulates consumer preferences and restaurant visits as designed and implemented by a LLM. These models are evaluated by human experts on their code correctness and quality for both verification and validation purposes. This work serves as a first step towards fine-tuned LLMs on existing models and documentations to create high-quality and functional ABMs based on either user prompts or standard protocols, contributing to further exploration on the future of AI-assisted geospatial simulation development. 

Keywords: Agent-Based Modeling, Large Language Models, Geospatial Simulation

Reference: 

Jiang, N., Wang, B. and Crooks, A.T. (2025), Large Language Models for Conceptualizing, Designing, and Generating Agent-based Models, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)
Next up was Ying Zhou who presented our work entitled "Identifying Environmental Characteristics That Influence Perceived Safety in Urban Spaces." In this work we explored how using social media data can be used to study the fear and how this relates to actual crimes within New York city. Broadly speaking we find through our analysis, that fear sentiment may spread out between the neighborhoods and their surrounding areas and that neighborhoods surrounded by crime-clusters may have high sentiments of fear. 

Abstract:
One goal of creating livable cities is to enhance health and safety. While previous research in spatial analysis and urban planning has focused on correlations between physical environments and crime, typically relying on police-reported crime data from sources like the Crime Open Database (CODE), safety perception is inherently subjective and cannot be fully represented by objective crime statistics alone. Also, urban planning today has gradually shifted its focus from a top-down mechanism to a bottom-up mechanism, so understanding and fostering spaces where residents feel safe is essential. This research examines factors that contribute to residents’ perceived insecurity in New York City. In addition to spatial analysis of the open crime data, the research used social media data to acquire people’s perceptions. The result indicates that the aggregations of perceived unsafe locations overlapped with aggregations of crime data's locations, such as in Manhattan’s neighborhoods, but they do not overlap with each other entirely. By adopting Latent Dirichlet Allocation (LDA), a method of topic modeling, the research filtered and summarized the posted texts and contents related to the negative descriptions of places or spaces in the city, and then it identified the related characteristics of the environments. The characteristics are investigated by the method of local Moran’s I, which indicates their spatial autocorrelation in some neighborhoods in the city of New York. This research offers “bottom’s views” about urban safety for both urban planning and decision-makers, which contributes to people-centered consideration for future development and urban resource distribution. 

Keywords: Safety, Crime, Urban Space, Livable Cities, Social Media, Spatial Analysis.
Methodology
Reference: 
Zhou, Y. and Crooks, A.T. (2025), Identifying Environmental Characteristics That Influence Perceived Safety in Urban Spaces, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)
The last day of the conference was another busy day with two talks. First was entitled "PySGN: A Python Package for Constructing Synthetic Geo-social Networks" where Boyu Wang presented our work (with Taylor Anderson and Andreas Züfle) on a Python package that can be used to generate synthetic geo-social networks. As readers of this blog might know we have a an interest in social networks and using them in modeling and this package provides a toolkit for others to easily create their own geosocial networks (e.g., Geospatial Erdős-Rényi, Barabási–Albert and Watts-Strogatz models). For interested readers, the source code available at: https://github.com/wang-boyu/pysgn.

Abstract:
Synthetic population has been widely used in social simulations such as traffic modeling, pedestrian movements, and the spread of infectious diseases. In recent years, much attention was focused on generating synthetic population with social networks, that captures social connections between individuals. While synthetic populations are often geographically explicit, various algorithms have been proposed to create realistic geographic social (geo-social) networks, aiming to integrate spatial information into people’s social links. We build an open-source Python package, namely PySGN, for constructing synthetic geo-social networks that incorporates position information, exhibits small-world network properties, and can be scaled to hundreds of thousands and potentially millions of nodes. We discuss different ways of parametrizing the method, by either a global average node degree, or an expected degree for each individual node. It is demonstrated through a case study with synthetic population in Buffalo, NY. By doing so, we aim to illustrate how such synthetic geo-social networks can be created, utilized, and analyzed in downstream agent-based modeling and network analysis tasks. This work is available as an open-source Python package and integrated with the PyData ecosystem (e.g., GeoPandas, NetworkX), and can be further extended with more synthetic geo-social network algorithms in the future. 
Keywords: Agent-Based Modeling, Synthetic Geo-Social Network, Python, Open-Source Software
Examples of Geosocial Networks Created in PySGN

Reference: 

Wang, B., Crooks, A.T., Anderson, T. and Züfle, A. (2025), PySGN: A Python Package for Constructing Synthetic Geo-social Networks. The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)
The final talk (well for me) was presented by Fuzin Yin who presented our work with Lucie Laurian  and Emmanuel Frimpong Boamah entitled "Analysis of Online Mutual Aid Network during Buffalo Blizzard 2022: Actors and Weak Ties." In this work we explored what kind of support was offered and requested over Facebook groups along with their network structures durring and shortly after the event utilizing machine learning. 

Abstract: 
In December 2022, Buffalo, NY encountered a once-in-a-generation blizzard that dropped over 4 feet of snow. This four-day snow event halted emergency services and left 47 dead. In the face of the devastating blizzard, Buffalonian demonstrated resilience and solidarity by establishing Facebook (FB) groups to share information and coordinate behaviors including donations, wellness checks, and snow removals. These spontaneous behaviors created an essential layer of protection when the major infrastructure was down. This research has collected data from Buffalo blizzard FB groups to analyze community-led self-help behaviors. We have used machine learning to classify FB messages into four categories (e.g., requesting help, offering help, emotional support, and other), and social network analysis to explore users’ communication patterns. Results show that out of all messages (n=9,988), 37% of them express emotional support, which is followed by messages offering help (25%). While requests for help constitute a small proportion (8%), they stimulate more replies than other categories. Network statistics suggest that the mutual aid network is low-density but with a high clustering coefficient. This implies that most group members are strangers with weak ties, but their connections are in the right place to allow efficient communication. However, users do not equally benefit where people requesting or offering help are central in online conversation while pure emotional supporters are at the periphery. We conclude that during the Buffalo blizzard 2022, online interactions translate into offline mutual assistance by establishing weak ties among disconnected users to facilitate the flow of information and resources. 
Keywords: crisis informatics, mutual aid, social network analysis, machine learning, social media

 

Results of Mutual Aid Network during Buffalo Blizzard 2022

Reference: 

Yin, F., Laurian, L., Crooks, A.T. and Boamah, E.F. (2025), Analysis of Online Mutual Aid Network during Buffalo Blizzard 2022: Actors and Weak Ties, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)

While this is a rather longer post than normal, we hope you found it interesting and also as noted at the top of the post, if any of these talks/topics are of interest to you please feel free to reach out. 

Friday, June 07, 2024

A comparison of social surveys and social media for vaccine hesitancy

In the past we have explored various ways to explore vaccine hesitancy and keeping with this theme we have a new paper published in PLOS ONE entitled "Understanding the determinants of vaccine hesitancy in the United States: A comparison of social surveys and social media" with Kuleen Sasse, Ron Mahabir, Olga Gkountouna and Arie Croitoru

In the paper we use social, demographic and economic (e.g., US Censusvariables to predict COVID-19 vaccine hesitancy levels in the ten most populous US metropolitan statistical areas (MSAs). By using  machine learning algorithms (e.g., linear regression, random forest regression, and XGBoost regression) we compare a set of baseline models that contain only these variables with models that incorporate survey data and social media (i.e., Twitter) data separately. 

We find that different algorithms perform differently along with variations in influential variables such as age, ethnicity, occupation, and political inclination across the five hesitancy classes (e.g., “definitely get a vaccine”, “probably get a vaccine”, “unsure”, “probably not get a vaccine”, and “definitely not get a vaccine”).   Further, we find that the application of the models to different MSAs yields mixed results, emphasizing the uniqueness of communities and the need for complementary data approaches. But in summary, this paper shows social media data’s potential for understanding vaccine hesitancy, and tailoring interventions to specific communities. If this sounds of interest, below we provide the abstract to the paper along with our mixed methods matrix, data sources used and the results from the various MSAs. At the bottom of the post, you cans see the full reference and the link to the paper so you can read more if you so desire. 

Abstract:
The COVID-19 pandemic prompted governments worldwide to implement a range of containment measures, including mass gathering restrictions, social distancing, and school closures. Despite these efforts, vaccines continue to be the safest and most effective means of combating such viruses. Yet, vaccine hesitancy persists, posing a significant public health concern, particularly with the emergence of new COVID-19 variants. To effectively address this issue, timely data is crucial for understanding the various factors contributing to vaccine hesitancy. While previous research has largely relied on traditional surveys for this information, recent sources of data, such as social media, have gained attention. However, the potential of social media data as a reliable proxy for information on population hesitancy, especially when compared with survey data, remains underexplored. This paper aims to bridge this gap. Our approach uses social, demographic, and economic data to predict vaccine hesitancy levels in the ten most populous US metropolitan areas. We employ machine learning algorithms to compare a set of baseline models that contain only these variables with models that incorporate survey data and social media data separately. Our results show that XGBoost algorithm consistently outperforms Random Forest and Linear Regression, with marginal differences between Random Forest and XGBoost. This was especially the case with models that incorporate survey or social media data, thus highlighting the promise of the latter data as a complementary information source. Results also reveal variations in influential variables across the five hesitancy classes, such as age, ethnicity, occupation, and political inclination. Further, the application of models to different MSAs yields mixed results, emphasizing the uniqueness of communities and the need for complementary data approaches. In summary, this study underscores social media data’s potential for understanding vaccine hesitancy, emphasizes the importance of tailoring interventions to specific communities, and suggests the value of combining different data sources.
Mixed methods matrix showing the data, processing, and model development steps used in our study.

Data sources used in our study.

MSA model performance (Bolded adjusted R2 values represent the best performing model for each modeling technique and MSA).

Wednesday, October 04, 2023

Leveraging newspapers to understand urban issues

In the past, this blog has explored several aspects of Detroit, such as how well its covered with Volunteered Street View Imagery or how through the use of agent-based models one can explore issues with urban shrinkage. Keeping up with the theme of shrinkage and Detroit but at the same time utilizing our growing interest in natural language processing (especially topic modeling) we (Na (Richard) Jiang, Hamdi Kavak, Wenjing Wang and myself) have a new paper entitled "Leveraging newspapers to understand urban issues: A longitudinal analysis of urban shrinkage in Detroit" published in Environment and Planning B

In the paper, we take 6794 English news articles published by national and local press organizations (e.g., Forbes, The New York Times, Newsweek, The Detroit News) between 1975 to 2021 using the keywords “Detroit”, “shrink” and “decline.” These keywords were selected based on the characteristics of the study area (i.e., Detroit) and the phenomenon of urban shrinkage. With these data we then use BERTopic to detect and classify all collected news articles into certain topics. We chose BERTopic because it captures the semantic relationship among words converting sentences and words to embedding and automatically generates the topic unlike other NLP topic modeling techniques (e.g., LDA). Our topic modeling results identify several insights with respect to Detroit's shrinkage. For example, we can detect the side effects of the 2007-2009 economic recession on Detroit's automobile industry, local employment status, and the housing market. If sounds of interest and you want to find out more, below we provide the abstract, some figures from the paper including the methodology workflow and an example of the resulting topics over time. Finally, at the bottom of the page you can see the full reference and s link to the paper itself.

Abstract 

Today we are awash with data, especially when it comes to studying cities from a diverse data ecosystem ranging from demographic to remotely sensed imagery and social media. This has led to the growth of urban analytics providing new ways to conduct quantitative research within cities. One area that has seen significant growth is using natural language processing techniques on text data from social media to explore various issues relating to urban morphology. However, we would argue that social media only provides limited insights when dealing with longer-term urban phenomena, such as the growth and shrinkage of cities. This relates to the fact that social media is a relatively recent phenomenon compared to longer-term urban problems that take decades to emerge. Concerning longer-term coverage, newspapers, which are increasingly becoming digitized, provide the possibility to overcome the limitations of social media and provide insights over a timeframe that social media does not. To demonstrate the utility of newspapers for urban analytics and to study longer-term urban issues, we utilize an advanced topic modeling technique (i.e., BERTopic) on a large number of newspaper articles from 1975 to 2021 to explore urban shrinkage in Detroit. Our topic modeling results reveal insights related to how Detroit shrinks. For example, side effects of 2007 to 2009 economic recessions on Detroit’s automobile industry, local employment status, and the housing market. 

Key Words: Natural Language Processing, Topic Modeling, Newspapers, Urban Shrinkage, Urban Analytics.

 

 Vacancy status change from 1970 to 2010 for city of Detroit and surrounding area.
Topic modeling work flow.
Topics over time (a) urban, (b) population, (c) shrinkage, (d) economy, (e) job, (f) house.

Full Reference:

Jiang, N., Crooks, A.T., Kavak, H. and Wang, W. (2023), Leveraging Newspapers to Understand Urban Issues: A Longitudinal Analysis of Urban Shrinkage in Detroit, Environment and Planning B. Available at https://doi.org/10.1177/23998083231204695. (pdf)

Thursday, September 07, 2023

Agent-Based Modeling of Consumer Choice

At the upcoming International Conference on Geographic Information Science (GIScience 2023) Boyu Wang and myself have a new paper entitled "Agent-Based Modeling of Consumer Choice by Utilizing Crowdsourced Data and Deep Learning." In the paper we explore how through mining Yelp reviews can inform an agents choices of restaurants. The model itself was created in Mesa and uses Mesa-Geo and  more details about the model can be found at https://github.com/wang-boyu/yelp-abm.  If this sounds of interest, below you can see the abstract to the paper, some fugues including the graphical user interface of the model and a link to the paper.

Abstract: People’s opinions are one of the defining factors that turn spaces into meaningful places. Online platforms such as Yelp allow users to publish their reviews on businesses. To understand reviewers' opinion formation processes and the emergent patterns of published opinions, we utilize natural language processing (NLP) techniques especially that of aspect-based sentiment analysis methods (a deep learning approach) on a geographically explicit Yelp dataset to extract and categorize reviewers' opinion aspects on places within urban areas. Such data is then used as a basis to inform an agent-based model, where consumers' (i.e., agents') choices are based on their characteristics and preferences. The results show the emergent patterns of reviewers' opinions and the influence of these opinions on others. As such this work demonstrates how using deep learning techniques on geospatial data can help advance our understanding of place and cities more generally.


Keywords: aspect-category sentiment analysis, consumer choice, agent-based modeling, online restaurant reviews.

An overview of proposed agent-based model logic.

Average star rating vs. average sentiment by aspect category for 200 randomly selected restaurants in the City of St. Louis, MO.

The prototype agent-based model (a) with simulated (b) and actual visiting patterns (c).

Full reference:

Wang, B. and Crooks, A.T. (2023), Agent-Based Modeling of Consumer Choice by Utilizing Crowdsourced Data and Deep Learning, in Beecham, R., Long, J.A., Smith, D., Zhao, Q., and Wise, S (eds), Proceedings of the 12th International Conference on Geographic Information Science (GIScience 2023), Dagstuhl Publishing, Dagstuhl, Germany., pp. 81:1-81:6. (pdf)


Wednesday, March 22, 2023

AAG 2023 Presentations

At this years Association of American Geographers (AAG) Annual Meeting we have a number of presentations ranging from how one can leverage newspaper articles to study cities over time, to that of how people may chose to become vaccinated. These presentations build on the great work of students and postdocs here at the University at Buffalo and link to our interests in urban analytics, machine learning and agent-based modeling. Below we just give a glimpse at these topics (along with their abstracts) and if you are interested in finding out more please reach out to us.

First up is a presentation with Qingqing Chen and Boyu Wang entitled "Community resilience to wildfires: A network analysis approach utilizing human mobility data."  In this presentation we explore how we can quantify a communities resilience to wildfires utilizing human mobility through network analysis methods. 

Abstract 

Natural disasters, such as earthquakes, floods, and wildfires, have been a long-standing concern to societies at large. With growing attention being paid to sustainable and resilient communities, such concern has been brought to the forefront of resilience studies. However, the definition of disaster resilience is intricate and can vary across the diverse disciplines that study them (e.g., geography, sociology and political science), making its definition and quantification elusive. Moreover, the vast majority of studies often focus on the immediate response to an event, not the long-term recovery of the area impacted by disasters. Thus to date investigating the resilience of an area or a society over a prolonged period of time has remained largely unexplored. To overcome these issues, we propose a novel approach from a social perspective utilizing network analysis and concepts from disaster science (e.g., the resilience triangle) to quantify the long-term impacts of wildfires, especially on collective human behavior. Taking the Camp and Mendocino Complex wildfires - the most deadly and the largest complex wildfires in California to date, respectively - as case studies, we capture the features of resilience, such as robustness and vulnerability, of communities based on human mobility data from 2018 to 2020. The results show that demographic and socioeconomic characteristics alone only partially capture community resilience, however, by leveraging human mobility data and network analysis techniques, we can enhance our understanding of resilience over space and time, which can provide a new lens to study natural disasters and their long-term impacts on society.

Keywords: Community Resilience, Natural Disasters, Wildfires, Social Network Analysis, Human Mobility, Space and Time.

Full Reference

Chen, Q., Wang, B. and Crooks, A.T. (2023), Community Resilience to Wildfires: A Network Analysis Approach Utilizing Human Mobility Data, The Association of American Geographers (AAG) Annual Meeting, 23rd –27th March, Denver, CO. (pdf)

Next up, moving from mobility to textural data, specifically that of newspapers Na (Richard) Jiang and myself have a presentation entitled "Leveraging Newspapers to Understand Urban Issues: A Longitudinal Analysis of Urban Shrinkage in Detroit". In this work we explore how can leverage Bertopic (a topic modeling technique) on newspaper articles spanning the years 1975 to 2021 to explore urban shrinkage in Detroit. 

 

Abstract 

Today we are awash with data especially when it comes to studying cities from a diverse data ecosystem ranging from demographic to that of remotely sensed imagery and social media. This has led to the growth of geographical data science and urban analytics providing new ways to conduct quantitative research within cities. One area that has seen significant growth is that of using natural language processing techniques on text data from social media to explore various issues relating to urban morphology. However, social media only provides limited insights when dealing with longer-term urban phenomena, such as the growth and shrinkage of cities. This relates to the fact that social media is a relatively recent phenomenon compared to more longer-term urban problems that take decades to emerge. With respect to the longer-term coverage, newspapers which are increasingly becoming digitized provide the possibility to overcome the limitations of social media and provide insights over a timeframe that social media does not. To demonstrate the utilization of newspapers within urban analytics and to study longer-term urban issues, we present an advanced topic modeling technique (i.e., Bertopic) on a large number of newspaper articles spanning the years 1975 to 2021 to explore urban shrinkage in Detroit. Our topic modeling results reveal the insights related to Detroit's shrinkage can be linked to the side effects of economic recessions on Detroit's automobile industry, local employment status, and the housing market. As such, this work demonstrates the potential of utilizing newspaper articles to study long-term issues

Keywords: Natural Language Processing, Topic Modeling, Newspapers, Text Data, Urban Shrinkage, Urban Analytics. 

 Full Reference

Jiang, N., and Crooks A.T. (2023), Leveraging Newspapers to Understand Urban Issues: A Longitudinal Analysis of Urban Shrinkage in Detroit, The Association of American Geographers (AAG) Annual Meeting, 23rd –27th March, Denver, CO. (pdf)

Switching gears slightly, we have another presentation that leverages text data, in this case Yelp reviews to help inform decision making within an agent-based model. This presentation with Boyu Wang is entitled "Do people care about others' opinions of places? Utilizing crowdsourced data and deep learning to model peoples’ review patterns."  We use a geospatial artificial intelligence (GeoAI) technique called aspect-based sentiment analysis to extract and categorize reviewers' opinion aspects on places within urban areas and then use this information to inform an agent-based model of peoples choices to which restaurants to go to.


Abstract  

People's opinions are one of the defining factors that turn spaces into meaningful places. While these opinions are subject to individual differences, they can also be influenced by the opinions from others. Online platforms such as Yelp allow users to publish their reviews on businesses. To understand reviewers' opinion formation processes and the emergent patterns of published opinions, we utilize geospatial artificial intelligence (GeoAI) techniques especially that of aspect-based sentiment analysis methods (a deep learning approach) on a geographically explicit Yelp dataset to extract and categorize reviewers' opinion aspects on places within urban areas. Such data is then used as a basis to inform an agent-based model, where reviewers' (i.e., agents') opinions are characterized by opinion dynamics. The parameters of these models are calibrated using extracted opinion aspects from the Yelp dataset. Such a method moves opinion dynamics models away from theoretical concepts to a more data-driven approach, with a specific emphasis being made on place. Focusing on 10 US metropolitan areas which are spread out across the country, we examine the calibrated influence coefficients for each opinion aspect category (e.g., location, experience, service), to compare reviewers' opinion formation processes across different categories. The results show the emergent patterns of reviewers' opinions and the influence of these opinions on others. As such this work demonstrates how using deep learning techniques on geospatial data can help advance our understanding of place and cities more generally.

Keywords: Agent-Based Modeling, Crowdsourcing, Deep Learning, GeoAI, Opinion Dynamics, Urban Analytics

Full Reference

Wang, B. and Crooks, A.T. (2023), Do People Care About Others' Opinions of Places? Utilizing Crowdsourced Data and Deep Learning to Model Peoples’ Review Patterns, The Association of American Geographers (AAG) Annual Meeting, 23rd –27th March, Denver, CO. (pdf)
Following with the agent-based modeling theme, our final presentation with Fuzhen Yin and Li Yin is entitled "How Information Propagation in Physical, Relational and Cyber Spaces Affects Covid-19 Vaccine Uptake: Evidence from Rural Areas." In this work we explore how people may or not be influenced by others (in physical, relational and cyber spaces) with respect to vaccination uptake.


 
Abstract 
With the advent of information and communication technologies, human dynamics studied in a purely physical space increasingly shift to a cyber and relational context. While researchers increasingly recognize the shift and call for attention to the multi-dimensionality of human dynamics (e.g., Splatial framework). Rarely have studies investigated how the information propagated in hybrid spaces affects people’s decision-making process, such as Covid-19 vaccine uptake. Meanwhile, compared to the urban population, the rural population faces greater digital barriers and has been further left out in human dynamics research. To fill this gap, our study investigates Covid-19 vaccine uptake in a rural county (i.e., Chautauqua) in New York State through agent-based modeling. We first generated a synthetic population to match the demographic characteristics of the census data. Then we created home, work, school, and social media networks to represent hybrid spaces. We defined the opinion dynamics of agents based on the social influence network theory. Next, we calibrated and validated our agent-based model based on real-world vaccine update records. Our research helps to elucidate the information propagation mechanism in hybrid spaces and clarify the decision-making process in the digital age. Furthermore, our method can also shed light on how to overcome data limitations for under-represented populations such as those who live in rural areas.

Keywords: Agent-based modeling, Covid-19, Vaccination, Opinion dynamics, Urban informatics, Rural geography

Full Reference 
Yin, F., Crooks, A.T. and Yin, L. (2023), How Information Propagation in Physical, Relational and Cyber Spaces Affects Covid-19 Vaccine Uptake: Evidence from Rural County, The Association of American Geographers (AAG) Annual Meeting, 23rd –27th March, Denver, CO. (pdf)

Thursday, February 09, 2023

Comparison between Online Social Media Discussions and Vaccination Rates

Continuing our work on social media and vaccinationsQingqing Chen, Arie Croitoru, and myself have a new paper entitled "A comparison between online social media discussions and vaccination rates: A tale of four vaccines" published in DIGITAL HEALTH. In the paper we explore online debates among four prominent vaccines (i.e., COVID-19, Influenza, MMR, and HPV) as captured on Twitter in the United States (US) from 2015 to 2021.

By using machine learning models (e.g., Naive Bayes, support vector machine (SVM), logistic regression, and extreme gradient boosting (XGBoost)) on over  11.7 million Twitter messages sent by approximately 2.6 million distinct users we found that while the COVID-19, it has come to dominate the vaccination discussion, there was an apparent discrepancy between the online debates and the actual vaccination rates in the US. 

If this sounds of interest and you wish to find out more, below we provide the abstract to to the paper, some figures which captures our workflow and a sample of the results such as a comparison between different vaccine discussions on Twitter and the actual vaccination rate. Finally at the bottom of the page you can find the full reference and a link to the paper.

Abstract

The recent COVID-19 pandemic has brought the debate around vaccinations to the forefront of public discussion. In this discussion, various social media platforms have a key role. While this has long been recognized, the way by which the public assigns attention to such topics remains largely unknown. Furthermore, the question of whether there is a discrepancy between people's opinions as expressed online and their actual decision to vaccinate remains open. To shed light on this issue, in this paper we examine the dynamics of online debates among four prominent vaccines (i.e., COVID-19, Influenza, MMR, and HPV) through the lens of public attention as captured on Twitter in the United States from 2015 to 2021. We then compare this to actual vaccination rates from governmental reports, which we argue serve as a proxy for real-world vaccination behaviors. Our results demonstrate that since the outbreak of COVID-19, it has come to dominate the vaccination discussion, which has led to a redistribution of attention from the other three vaccination themes. The results also show an apparent discrepancy between the online debates and the actual vaccination rates. These findings are in line with existing theories, that of agenda-setting and zero-sum theory. Furthermore, our approach could be extended to assess the public's attention toward other health-related issues, and provide a basis for quantifying the effectiveness of health promotion policies.

Keywords:  COVID-19, Influenza, MMR, HPV, Social media, Vaccination.

 

The workflow for comparing between online social media discussion and vaccination rates.

The quarterly distribution of percentage of users by different vaccine discussion from 2015 to 2021.

 The comparison between different vaccine discussions on Twitter and growth rate of the actual vaccination rate collected from the CDC (a) COVID-19; (b) Influenza; (c) HPV; (d) MMR.

The changes of emotion over time for different vaccines.

Full reference: 

Chen Q, Croitoru A. and Crooks A.T (2023), A Comparison between Online Social Media Discussions and Vaccination Rates: A tale of four vaccines. DIGITAL HEALTH: 9. doi:10.1177/20552076231155682. (pdf)

Thursday, December 08, 2022

Simulating Geographical Systems using CA and ABMs

 
In the chapter we discuss how thinking and studying of geographical systems like cities has changed over time from top down aggregate analysis to more bottom up approaches which captures the complex nature of such systems. We then discuss how we can model such systems from a cellular automata and agent-based perspectives. and how these styles of models have evolved and how they can be used to model future systems. If this sounds of interest below we provide the abstract to the chapter, some of the figures that accompany it and at the  bottom of the page we provide the full reference to the paper along with a link to the chapter itself.
"Abstract: How we view and understand the processes driving and shaping geographical systems is constantly evolving. This is due to the appearance of new rich data sources, increased computing power and storage, and the development of individual-level approaches. This allows us to explore geographical systems (from the bottom up) at scales not possible in the past. In this chapter, we examine the utility of two of the most commonly used individual-level modelling approaches, cellular automata and agent-based modelling. We outline their key differences and how these models are being used to further our understanding of geographical systems through simulation. We conclude with a discussion about the challenges that both approaches need to meet to continue developing into the future.
Keywords: Cellular automata; Agent-based models; Geographical systems; Machine learning
 

A SLEUTH like model stylized on Santa Fe, New Mexico denoting how land use charges over time from undeveloped (grey) to urban (red).

Example applications of agent-based models at different spatial and temporal scales

Full reference:

Heppenstall, A., Crooks, A.T., Manley, E. and Malleson, N. (2022) Simulating Geographical Systems using Cellular Automata and Agent-based Models, in Rey S. and Franklin, R. (eds.), Handbook of Spatial Analysis in the Social Sciences, Edward Elgar Publishing, Cheltenham, UK, pp. 142-157. (pdf)

Monday, September 19, 2022

Information propagation on cyber, relational and physical spaces about covid-19 vaccine

It seems that its been a quite some time that we posted about geosocial analysis but in a recent paper with  Fuzhen Yin  and Li Yin entitled "Information Propagation on Cyber, Relational and Physical Spaces about Covid-19 Vaccine: Using Social Media and the Splatial Framework" published in Computers, Environment and Urban Systems we revisit this line of work while at the same time linking it to Covid and vaccination debates. 

Specifically we examine the interaction between cyber, relational (i.e, networks between objects), and physical spaces using the Splatial framework. Through our analysis focused on New York State, we find that non-polarized vaccination debates were observed in cyber, relational, and physical spaces. Furthermore,  we found that while physical space users had less anti-vaccine stance than relational and cyber space users there were strong interactions are observed between physical–relational, and relational-cyber spaces.If this sort of thing interests you. Below we provide the abstract to the paper along with some figures which show the study area, our methodology and some of the results. While at the bottom of the post we provide the full reference and the link to the paper.

Abstract:

With the advent of social media, human dynamics studied in purely physical space have been extended to that of a cyber and relational context. However, connections and interactions between these hybrid spaces have not been sufficiently investigated. The “space-place (Splatial)” framework proposed in recent years allows capturing human activities in the hybrid of spaces. This study applies the Splatial framework to examine the information propagation between cyber, relational, and physical spaces through a case study of Covid-19 vaccine debates in New York State (NYS). Whereby the physical space represents the regional boundaries and locations of social media (i.e., Twitter) users in NYS, the relational space indicates the social networks of these NYS users, and the cyber space captures the larger conversational context of the vaccination debate. Our results suggest that the Covid-19 vaccine debate is not polarized across all three spaces as compared to that of other vaccines. However, the rate of users with a pro-vaccine stance decreases from physical to relational and cyber spaces. We also found that while users from different spaces interact with each other, they also engage in local communications with users from the same region or same space, and distance-based and boundary-confined clusters exist in cyber and relational space communities. These results based on the Splatial framework not only shed light on the vaccination debates but also help to define and elucidate the relationships between the three spaces. The intense interactions between spaces suggest incorporating people’s relational network and cyber presence in physical place-making.

Keywords: Covid-19, Vaccination, Social media, Social network analysis, Community detection, Urban informatics
Schematic representation of the three spaces: cyber, relational and physical spaces.

Map of study area (NYS) with the primary road system. Red dots denote collected vaccine-related tweets in NYS.

Research workflow to investigate the propagation of different opinions between three spaces: cyber, relational and physical spaces.

Network visualization of the eight top large communities in relational space. (A) Visualization of communities using ForceAtlas layout. (B) Project communities into physical space. Nodes without location information are placed outside of NYS.

The hybrid space network shows the information propagation between physical and relational spaces. (A) shows the network of all tweets, (B) shows the pro-vaccine tweets, and (C) shows the anti-vaccine tweets.
 
Full Reference:

Yin, F., Crooks, A.T. and Yin, L. (2022), Information Propagation on Cyber, Relational and Physical Spaces about Covid-19 Vaccine: Using Social Media and the Splatial Framework, Computers, Environment and Urban Systems. Available at: https://doi.org/10.1016/j.compenvurbsys.2022.101887.  (pdf)