Showing posts with label Crowdsourcing. Show all posts
Showing posts with label Crowdsourcing. Show all posts

Tuesday, May 13, 2025

Crowdsourcing dust storms utilizing social media data

In the past we have explored how social media can be used to delineate earthquakes, study human-wildlife interactions, understand urban morphology, urban smells or  locating wildfires among many other things. 

Keeping with the last topic (i.e., locating things), in a new paper published in GeoJournal entitled "Crowdsourcing dust storms in the United States utilizing social media data," Stuart EvansFestus Adegbola and myself explore how we can use X (formerly Twitter) and Flickr  to source observations of windblown dust. 

As such the paper demonstrates how social media data can act as supplementary source for dust events monitoring and captures the seasonal trends of such events. Furthermore, the paper highlights the potential of using crowdsourced data for the often overlooked field of dust monitoring that has substantial health and economic impacts. 

If this sounds of interest, below we provide the abstract to the paper along with some figures which showcase our methodology and comparison with National Weather Service dust advisories and VIIRS satellite data. At the bottom of the post, you can find the full reference to the paper along with a link to it. 

Abstract: 

Dust storms and other dust events are natural phenomena characterized by strong winds carrying large amounts of fine particles which have significant environmental and human impacts. However, capturing the occurrence of such phenomena is a challenge. Previous studies have limitations due to available data, especially regarding short-lived, intense dust storms and events that are not captured by observing stations and satellite instruments. In recent years, the advent of social media platforms has provided a unique opportunity to access vast amounts of crowdsourced data. This paper explores the utilization of Flickr and X (Twitter) data to study dust event occurrences within the United States and their correlation with National Weather Service (NWS) advisories. The work ascertains the reliability of using crowdsourced data as a supplementary source for dust events monitoring. Our analysis of Flickr and X indicates that the Southwest region is most susceptible to dust events, with Arizona leading in the highest number of occurrences. On the other hand, the Great Plains show a scarcity of crowdsourced data related to dust events, which can be attributed to the sparsely populated nature of the region. Furthermore, seasonal analysis reveals that dust events are prevalent during the Summer months followed by Spring. These results are consistent with previous traditional studies that did not use social media of dust occurrences in the U.S., and Flickr-identified images of dust events show substantial co-occurrence with regions of NWS dust warnings. This paper highlights the potential of using crowdsourced data for the often overlooked field of dust monitoring that has substantial health and economic impacts.
Keywords: Dust storms, Crowdsourcing, Social media, Weather. 

 

Flowchart of our workflow
Selected posts retrieved from X showing active dust events.

Selected images retrieved from Flickr showing active dust events.

Map showing the distribution of flickr-identified dust event occurrences, X-identified dust event occurrences, National Weather Service dust advisories, including dust storm (DS) warnings and blowing dust (DU) advisories.

Seasonal cycle of dust events using social media metadata, the National Weather Service advisories, and the VIIRS satellite data.

Examples of social media identified dust events and satellite observations for the same day. Brown shaded pixels indicate locations Suomi-VIIRS observed dust particles. Any VTEC warnings issued by NWS for the location are shown after the date of each dust event, with HWW and DSW indicating High Wind Warning and Dust Storm Warning, respectively.

Full Referece: 
Adegbola, F., Crooks, A.T. and Evans, S.M. (2025). Crowdsourcing dust storms in the United States utilizing social media data. GeoJournal, 90(3), pp.1-18. Available at https://doi.org/10.1007/s10708-025-11359-9 (pdf)

Friday, January 31, 2025

New Directions in Mapping the Earth’s Surface with Citizen Science and Generative

In previous posts, we have written how large language models (LLMs) like ChatGPT can be used in various urban analytical applications. We have kept exploring this potential especially with respect to citizen science applications. To this end we have just published a new paper in iScience, entitled "New Directions in Mapping the Earth’s Surface with Citizen Science and Generative AI". In the paper, lead by Linda See, we discuss how multi-modal LLMs (MLLMs) which are like LMMs but can take different forms of inputs (e.g., text, images, video) and output multi-modal information (e.g., take an image and output a description) could be leveraged to enhance citizen science land cover/land use mapping campaigns. If this sounds of interest, below you can read the abstract to the paper, see some of the figures we use to build our argument, while at the bottom of the post you can see the full reference and a link to the actual paper.
Abstract: 
As more satellite imagery has become openly available, efforts in mapping the Earth’s surface have accelerated. Yet the accuracy of these maps is still limited by the lack of in-situ data needed to train machine learning algorithms. Citizen science has proven to be a valuable approach for collecting in-situ data through applications like Geo-Wiki and Picture Pile, but better approaches for optimizing volunteer time are still required. Although machine learning is being used in some citizen science projects, advances in generative Artificial Intelligence (AI) are yet to be fully exploited. This paper discusses how generative AI could be harnessed for land cover/land use mapping by enhancing citizen science approaches with multi-modal large language models (MLLMs), including improvements to the spatial awareness of AI.
Visual interpretation tasks undertaken by ChatGPT for (a) a wetland/mangrove landscape in South America (b) an agricultural area in central Europe.
Visual interpretation tasks undertaken by ChatGPT for identification of natural and non-natural ecosystems where ChatGPT misclassified the images as non-natural for locations in (a) Chad and (b) Austria. In (c), the image from Colombia was classified as unsure by validators but natural by ChatGPT.
Integrating multi-modal Large Language Models (MLLMs) in a citizen science visual interpretation workflow.
Full reference : 
See, L., Chen, Q., Crooks, A., Bayas, J.C.L., Fraisl, D., Fritz, S., Georgieva, I., Hager, G., Hofer, M., and Lesiv, M., Malek, Ž., Milenković, M., Moorthy, I., Orduña-Cabrera, F., Pérez-Guzmán, K., Schepaschenko, D., Shchepashchenko, M., Steinhauser, J.and McCallum, I. (2025), New Directions in Mapping the Earth’s Surface with Citizen Science and Generative AI, iScience, doi: https://doi.org/10.1016/j.isci.2025.111919(pdf)

Tuesday, December 19, 2023

Crowdsourcing Dust Storms in the United States Utilizing Flickr

In the past on this site we have written about how one can use social media to study the world around us. Often the focus has been on Twitter but that is not the only social media platform available.  Another is Flickr, and while in past posts have show how we can use this platform to explore bird sightings, wildfires and human migration we are now turning our attention to other phenomena. One of which is dust storms. Working with Festus Adegbola and Stuart  Evans we have just presented a poster at the 2023 American Geophysical Union Fall Meeting entitled "Crowdsourcing Dust Storms in the United States Utilizing Flickr"

In this research we compare Flickr images with National Weather Service  advisories and the VIIRS Deep Blue aerosol product data from the Suomi-NPP satellite. Our preliminary findings show that Flickr images of dust storms have a substantial co-occurrence with regions of NWS blowing dust advisories. If this sounds of interest, below you can read our abstract, see our workflow and the poster itself. 

Abstract

Dust storms are natural phenomena characterized by strong winds carrying large amounts of fine particles which have significant environmental and human impacts. Previous studies have limitations due to available data, especially regarding short-lived, intense dust storms that are not captured by observing stations and satellite instruments. In recent years, the advent of social media platforms has provided a unique opportunity to access a vast amount of user-generated data. This research explores the utilization of Flickr data to study dust storm occurrences within the United States and their correlation with National Weather Service (NWS) advisories. The work ascertains the reliability of using crowdsourced data as a supplementary tool for dust storm monitoring. Our analysis of Flickr metadata indicates that the Southwest is most susceptible to dust storm events, with Arizona leading in the highest number of occurrences. On the other hand, the Great Plains show a scarcity of Flickr data related to dust storms, which can be attributed to the sparsely populated nature of the region. Furthermore, seasonal analysis reveals that dust storm events are prevalent during the Summer months, specifically from June to August, followed by Spring. These results are consistent with previous studies of dust occurrence in the US, and Flickr-identified images of dust storms show substantial co-occurrence with regions of NWS blowing dust advisories. This research highlights the potential of unconventional user-generated data sources to crowdsource environmental monitoring and research.

Data collection and workflow.
Distribution of Flickr identified dust storm occurrences and NWS dust storm advisories.

Full Reference: 

Adegbola, F., Crooks, A.T. and Evans, S. (2023), Crowdsourcing Dust Storms in the United States Utilizing Flickr, American Geophysical Union (AGU) Fall Meeting, 11th – 15th December, San Francisco, CA. (abstract, poster)

Wednesday, March 16, 2022

Leveraging Street Level Imagery for Urban Planning

Just a short post that say that  Linda See and myself have a new editorial in Environment and Planning B: Urban Analytics and City Science entitled " Leveraging Street Level Imagery for Urban Planning." While in the in the past we have written about street view imagery and how there are initiatives like KartaView (previously named OpenStreetView and OpenStreetCam) and Mapillary which allow for the collection of volunteered street view imagery (VSVI) using just smartphones. But we have not really delved much into how such initiatives could be used to assist assist urban planning (e.,g. change detection, augmented reality (AR) and urban navigation).  If this sounds of interest,  please feel free to check out our editorial here

Exploring urban change in Buffalo, New York with Google Street View in October 2020 and the same location in the 2007 inset.

 Full Reference:

Crooks, A.T. and See, L. (2022), Leveraging Street Level Imagery for Urban Planning, Environment and Planning B, https://doi.org/10.1177/23998083221083364

Wednesday, June 10, 2020

New Paper: A Thematic Similarity Network Approach for Analysis of Places Using VGI

Building upon our work on volunteered geographical information (VGI) and ambient geographic information (AGI) and how such data (e.g. social media) can be used to understand place, Xiaoyi Yuan, Andreas Züfle and myself have a new paper entitled: "A Thematic Similarity Network Approach for Analysis of Places Using Volunteered Geographic Information" in the ISPRS International Journal of Geo-InformationIn this paper we use textual data from crowdsourced reviews originating with TripAdvisor and geo-located Twitter data and leverage this unstructured geographical information to comprehend the complexity of places at scale. Specifically we explore the connectedness and relationships of places through thematic (i.e., topical) similarity networks using Manhattan, New York as a case study. If such work sounds of interest to you, below we provide the abstract to the paper in order for you to gain a greater understanding of work, along with some figures that show our workflow and how communities where connected, before presenting some of our results. Finally at the bottom of the post, the full reference and a link to the paper is provided.  For those interested in extending or utilizing this work. The python code for presented in our analysis is available at: https://bitbucket.org/xiaoyiyuan/network_vgi/

Abstract:
The research presented in this paper proposes a thematic network approach to explore rich relationships between places. We connect places in networks through their thematic similarities by applying topic modeling to the textual volunteered geographic information (VGI) pertaining to the places. The network approach enhances previous research involving place clustering using geo-textual information, which often simplifies relationships between places to be either in-cluster or out-of-cluster. To demonstrate our approach, we use as a case study in Manhattan (New York) that compares networks constructed from three different geo-textural data sources --TripAdvisor attraction reviews, TripAdvisor restaurant reviews, and Twitter data. The results showcase how the thematic similarity network approach enables us to conduct clustering analysis as well as node-to-node and node-to-cluster analysis, which is fruitful for understanding how places are connected through individuals’ experiences. Furthermore, by enriching the networks with geodemographic information as node attributes, we discovered that some low-income communities in Manhattan have distinctive restaurant cultures. Even though geolocated tweets are not always related to place they are posted from, our case study demonstrates that topic modeling is an efficient method to filter out the place-irrelevant tweets and therefore refining how of places can be studied.

Keywords: Geo-Textual Data, Volunteered Geographic Information, Crowdsourcing, Similarity Network Analysis, Topic Modeling

Work flow from data input to the construction of the thematic similarity network and analysis (i.e., community detection and unique nodes discovery).

A stylized network demonstrating the process of community detection from a fully-connected similarity network.


Network visualization of all communities from the thematic similarity networks with major communities highlighted. Only the major communities are shown on the map for the sake of clarity. Major communities in Network visualization and mapping for each network are colored the same and thus the legend applies for both.


Two examples of communities with boundary nodes and their respective topics.

Full Reference:
Yuan X., Crooks, A.T. and Züfle, A. (2020), A Thematic Similarity Network Approach for Analysis of Places Using Volunteered Geographic Information, ISPRS International Journal of Geo-Information,  9(6), 385, https://doi.org/10.3390/ijgi9060385. (pdf)

Tuesday, May 26, 2020

Crowdsourcing Street View Imagery: A Comparison of Mapillary and OpenStreetCam


In the past we have written extensively on Volunteered Geographic Information (VGI) such as OpenStreetMap or Twitter. However, we have not really explored Street View Imagery  (SVI), well not until now. Within the realm of VGI, SVI has emerged in recent years as a novel and rich source of data on cities from which geographic information can be derived.

Perhaps the most well-known example of SVI utilization is that of Google Street View (GSV). While SVI has been traditionally collected by governmental agencies and companies alike, we are now also witnessing the emergence of Volunteered Street View Imagery (VSVI), which relies on a crowdsourced effort to provide geotagged street-level imagery coverage of traversable pathways (e.g., a street or trail). Such imagery, similar to GSV, provides detailed information about the location of objects such as cars, road markings, traffic lights and signs, and allows for the automatic extraction of features at scale. Such imagery can also be mined using machine learning algorithms to automatically derive points of interest (POI) databases (e.g., locations of coffee shops and fire hydrants) without the intervention of the citizen.

To explore VSVI we have just published a new paper entitled: "Crowdsourcing Street View Imagery: A Comparison of Mapillary and OpenStreetCam" in the ISPRS International Journal of Geo-Information. In this paper we examine VSVI data collected from two different platforms: Mapillary and OpenStreetCam (OSC) for four metropoiltan areas in the United States (i.e., Washington (District of Columbia), San Francisco (California), Phoenix (Arizona), and Detroit (Michigan)). Both of these online platforms accept sequences of images captured from mobile devices and uploaded via an app on the device (like those shown in the image to the right). Images are geolocated using the device’s global positioning system (GPS). More specifically the paper examines:
  • the level of spatial coverage of each platform in order to assess the overall potential of such platforms to provide adequate coverage of geographic information.
  • user contribution patterns in Mapillary and OSC in order to understand how users are contributing to these platforms.
Results from our systematic and quantitative analysis of these two emerging VGI sources indicate that most Mapillary and OSC contributions occurred along control-access highways and local roads, and that the overall coverage in these sources is variable in comparison to an authoritative source (i.e., TIGER). Furthermore, our results showed that while the number of contributors varied across sites, only a few contributors were responsible for producing most of the raw data. User contribution patterns were also different in Mapillary and OSC. Specifically, we found that while patterns in coverage were variable for the different OSC sites, coverage patterns in Mapillary tended to be similar among sites. This finding may be linked to several factors, including differences in mapping practice, or issues with participation inequality, a topic that has been highly researched for other VGI platforms such as OSM, but which is still lacking within VSVI. Lastly, user contributions in Mapillary tended to be higher around 8:00 am, 1:00 pm and 5:00 pm (local time). This finding suggests that VSVI contributions tend to coincide with the morning and afternoon commute, and the lunch hour of the contributors.

If you wish to find out more about this work below we provide the abstract to the paper, a visual flowchart of our workflow and some of our our results. The full reference and link to the paper is provided at the bottom of the post.

Abstract:
Over the last decade, Volunteered Geographic Information (VGI) has emerged as a viable source of information on cities. During this time, the nature of VGI has been evolving, with new types and sources of data continually being added. In light of this trend, this paper explores one such type of VGI data: Volunteered Street View Imagery (VSVI). Two VSVI sources, Mapillary and OpenStreetCam, were extracted and analyzed to study road coverage and contribution patterns for four US metropolitan areas. Results show that coverage patterns vary across sites, with most contributions occurring along local roads and in populated areas. We also found that a few users contributed most of the data. Moreover, the results suggest that most data are being collected during three distinct times of day (i.e., morning, lunch and late afternoon). The paper concludes with a discussion that while VSVI data is still relatively new, it has the potential to be a rich source of spatial and temporal information for monitoring cities.

Keywords: Crowdsourcing; Volunteered Geographic Information; Street View Imagery; Mapillary, OpenStreetCam
Overview of methodology

Spatial distribution of road networks.
Spatial comparison of roads in kilometers.


Full Reference: 
Mahabir, R., Schuchard, R., Crooks, A.T., Croitoru, A. and Stefanidis, A. (2020), Crowdsourcing Street View Imagery: A Comparison of Mapillary and OpenStreetCam, ISPRS International Journal of Geo-Information. 9(6), 341; https://doi.org/10.3390/ijgi9060341 (pdf)

Friday, January 31, 2020

The Interplay Between the Media and the Public in Mass Shootings

Continuing our work on shootings we recently had a paper published in Criminology and Public Policy entitled: "Responses to Mass Shooting Events: The Interplay Between the Media and the Public." However, here we do not look at bots but instead explore the how the public responds to mass shooting events (e.g. Las Vegas, Sutherland Springs, Marshall County, Parkland, Santa Fe), by seeking additional information or exchanging opinions about them in media coverage (e.g. newspaper articles via LexisNexis) and through online sources of information (e.g. Google Trends, Wikipedia and Online Social Networks (i.e. Twitter)). 

Overall, our results show discernible patterns in both time and space in the public’s online information seeking activities after a mass shooting. In addition we find discernible online information seeking patterns in geographic space, with a focal area of interest in the state in which the shooting event occurs, surrounded by a region of reduced interest. This finding further suggests that online information seeking activities are driven, at least in part, by geographic proximity to mass shooting events.

If you wish to find out more about this research, below we provide the summary and policy implication to the paper along with some figures from our methodology (e.g., how we go about analyzing temporal and geographical trends) and some of the results. Finally at the bottom of the post we provide the full reference and a link to the paper.

Abstract:
Research Summary: Public mass shootings tend to capture the public’s attention and receive substantial coverage in both traditional media and online social networks (OSNs) and have become a salient topic in them. Motivated by this, the overarching objective of this paper is to advance our understanding of how the public responds to mass shooting events in such media outlets. Specifically, it aims to examine whether distinct information seeking patterns emerge over time and space, and whether associations between public mass shooting events emerge in online activities and discourse. Towards this objective, we study a sequence of five public mass shooting events that have occurred in the United States between October 2017 and May 2018 across three major dimensions: the public’s online information seeking activities, the media coverage, and the discourse that emerges in a prominent OSN. To capture these dimensions, respectively, data was collected and analyzed from Google Trends, LexisNexis, Wikipedia Page views, and Twitter. The results of our analysis suggest that distinct temporal patterns emerge in the public’s information seeking activities across different platforms, and that associations between an event and its preceding events emerge both in the media coverage and in OSNs.
Policy Implication: Studying the evolution of discourse in OSNs provides a valuable lens to observe how society’s views on public mass shooting events are formed and evolved over time and space. The ability to analyze such data allows tapping into the dynamics of reshaping and reframing public mass shooting events in the public sphere and enable it to be closely studied and modeled. A deeper understanding of this process, along with the emerging associations drawn between such events, can then provide policy and decision-makers with opportunities to better design policies and communicate the significance of their goals and objectives to the public.
A framework for the analysis of temporal and geographical trends .

The analysis processes of Twitter and LexisNexis data.

Geographic patterns in online search activity in Google Trends for the five events in our study.

Chronologically ordered Google Trends search activity (a, left) and Wikipedia page views (b, right). Each vertical solid black line marks the occurrence of one of four shooting events examined in the analysis (as indicated by the line label).

Mentions of prior events during the first approximately 1-month period following each event in each of the events studied. (a) Sutherland Springs, (b) Marshall County, (c) Parkland, (d) Santa Fe.

Full Reference:
Croitoru, A., Kien, S., Mahabir, R., Radzikowski, J., Crooks, A.T., Schuchard, R., Begay, T., Lee, A., Bettios, A. and Stefanidis, A. (2020), Responses to Mass Shooting Events: The Interplay Between the Media and the Public, Criminology and Public Policy, 19(2): 335–360. (pdf)

Tuesday, January 14, 2020

New Paper: Insights into Human-wildlife Interactions in Cities from Bird Sightings Recorded Online

In the past we have explored how social media can be used to delineate earthquakes, locate wildfires or be used to understand urban morphology. However, recently we have also started to explore how social media and crowdsourced data can be utilized to to study socio-environmental systems. Keeping with this them, Bianca Lopez, Emily Minor, and myself have recently had a paper published in Landscape and Urban Planning entitled "Insights into Human-wildlife Interactions in Cities from Bird Sightings Recorded Online."  

In the paper we explore where do people observe birds, using the city of Chicago as our case study. By utilizing urban bird observations collected from eBird, iNaturalist, and Flickr we find that most bird observations occurred in open space zoned for recreational use. Further analysis revealed that the number of bird observations varied with income, population size, and proximity to Lake Michigan. If you want to find out more, below is the abstract to the paper, along with some figures of the results and at the bottom of the post, the full reference and a link to the paper. 

Abstract:
Interactions with nature can improve the wellbeing of urban residents and increase their interest in biodiversity. Many places within cities offer opportunities for people to interact with wildlife, including open space and residential yards and gardens, but little is known about which places within a city people use to observe wildlife. In this study, we used publicly available spatial data on people’s observations of birds from three online platforms—eBird, iNaturalist, and Flickr—to determine where people observe birds within the city of Chicago, Illinois (USA). Specifically, we investigated whether land use or neighborhood demographics explained where people observe birds. We expected that more observations would occur in open spaces, and especially conservation areas, than land uses where people tend to spend more time, but biodiversity is often lower (e.g., residential land). We also expected that more populated neighborhoods and those with higher median age and income of residents would have more bird observations recorded online. We found that bird observations occurred more often in open spaces than in residential areas, with high proportions of observations in recreation areas. In addition, a linear regression model showed that neighborhoods with higher median incomes, those with larger populations, and those located closer to Lake Michigan had more bird observations recorded online. These results have implications for conservation and environmental education efforts in Chicago and demonstrate the potential for social media and citizen science data to provide insight into urban human-wildlife interactions.
Keywords: Urban biodiversity, human-nature interaction, open space, residential, spatial analysis, birdwatching.

Map of bird observations from the three web platforms (Flickr, eBird, and iNaturalist) across the city of Chicago, in relation to mean median income of community areas (left panel) and open space, residential land use, highways, and waterways (right panel).

Proportions of observations recorded in different land uses on the three different online platforms (n = 7944 eBird; n = 474 iNaturalist; n = 561 Flickr). There was a significant difference between the three distributions (simulated p-value less than 0.001), including in the proportions of observations in conservation, recreation, and residential land uses.

Full Reference:
Lopez, B.E., Minor, E.S. and Crooks, A.T. (2020), Insights into Human-wildlife Interactions in Cities from Bird Sightings Recorded Online, Landscape and Urban Planning. 196: 103742. (pdf)

Tuesday, November 05, 2019

New Paper: Assessing the Placeness of Locations through User-contributed Content

In the past we have written about how one can use crowdsourced data to gain a collective sense of place from Twitter contributions and also from corresponding Wikipedia entries (e.g. here). In a new paper with Xiaoyi Yuan, we extend this work to explore how user-contributed data can be used to explore if urban places are becoming inauthentic due to urban commodification and standardization by chain stores such as restaurants. To this end, at the at 3rd ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery (GeoAI) we have a paper entitled: "Assessing the Placeness of Locations through User-contributed Content"

In the paper we attempt to understand the relationship between restaurants and urban identities via user-contributed content. We extracted and analyzed information from over 3 million Yelp reviews from 37,000 restaurants using a Convolutional Neural Network (CNN) model in order to study places from the bottom up. Specifically we were interested to what extent cities share similarities or differences in their Yelp restaurant reviews. Furthermore, we wanted to explore how opinion aspects (i.e. what reviewers care about the most) are mentioned differently in urban chain and independent restaurants. Through the analysis of the Yelp reviews we find that online geo-tagged text data is fruitful for understanding places and aspect-based sentiment analysis helps us understand the large volumes of text. Not only did we discover that cities show homogeneity in terms of restaurant reviews, but for chain restaurants, “location” often emphasizes the differences between different stores of the same chain whereas for independent restaurant reviews, the aspect “location” reflects the characteristics of the places the restaurants are situated. If this is of interest to you, below we provide the abstract to the paper, along with some of the key findings and a link to the paper.

Abstract
Previous research has argued that urban places are becoming “placeless” and inauthentic. Many local policies have also proposed to encourage more independent stores in order to restore urban identity. Others argue, however, that chain stores provide affordable merchandise and different locations of the same chain may have different meanings to an individual. The research presented in this paper uses a Convolutional Neural Networks model to extract opinion aspects from more than 3 million user-contributed Yelp restaurant reviews. The results show high homogeneity among cities in terms of the average proportions of aspects in restaurant reviews. In addition, for fast food chains, “location” is the only aspect category reviewed proportionally higher than independent fast food restaurants. An analysis of the co-occurrences of “location” indicates that the identity of chain restaurants stems from the comparison between the same chain of different locations whereas the identity of the independent restaurants is more diverse, implying the intricacies of placeness of urban stores. This research demonstrates that fine-grained sentiment analysis (i.e., opinion aspect extraction and analysis) with geo-tagged text data is fruitful for studying nuanced place perceptions on a large scale.
KEYWORDS: Urban Places, Convolutional Neural Networks, Aspect-based Sentiment Analysis
Figure 1: Illustration of an example of a CNN layer.
Figure 3: Mapping restaurants in NV, AZ, PA, NC, WI, IL. Not all cities are shown in each state. Only cities have data that accounts for the majority of the restaurants in that state are mapped, for the sake of visual clarity.
Figure 6: Average proportions of aspect categories for chain and independent fast food restaurants for two kinds of cuisine (American, Mexican) in Las Vegas, Phoenix, and Charlotte, normalized by dividing the mean for comparison.
Reference:
Yuan X. and Crooks A.T. (2019), Assessing the Placeness of Locations through User-contributed Content, in Gao, S., Newsam, S., Zhao, L., Lunga, D., Hu, Y., Martins, B., Zhou, X. and Chen, F. (eds.), Proceedings of the 3rd ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery (GeoAI), Chicago, IL. pp. 15-23. (pdf)

Friday, April 26, 2019

Computational Social Science of Disasters: Opportunities and Challenges

Figure 1: Relation of computational social science of
disasters (CSSD) with other fields.
Past posts have discussed or demonstrated how  computational social science (CSS) (i.e. the study of social science through computational methods) can be utilized explore disasters or diseases but this has not really been  formalized.  To this end, Annetta Burger, Talha Oz, William Kennedy and myself have just had a paper published in Future Internet entitled "Computational Social Science of Disasters: Opportunities and Challenges". In the paper we introduce computational social science of disasters (CSSD). CSSD is defined as an approach to explain the social dynamics of disasters via computational means by adopting the relevant parts of CSS, social sciences in disaster, and crisis informatics as depicted in Figure 1. Specifically, we briefly review the domains and the approaches of each of the traditional social science disciplines to disasters (e.g. sociology, psychology, anthropology, political science, and economics). Next we describe the fields of CSS and crisis informatics before discussing the components of CSSD. We highlight some exemplar studies which capture certain elements of CSSD along with the challenges and opportunities it brings to the study of disasters. If you would like to find out more, below is the abstract to the paper along with the full reference and link to the paper.

Abstract
Disaster events and their economic impacts are trending, and climate projection studies suggest that the risks of disaster will continue to increase in the near future. Despite the broad and increasing social effects of these events, the empirical basis of disaster research is often weak, partially due to the natural paucity of observed data. At the same time, some of the early research regarding social responses to disasters have become outdated as social, cultural, and political norms have changed. The digital revolution, the open data trend, and the advancements in data science provide new opportunities for social science disaster research. We introduce the term computational social science of disasters (CSSD), which can be formally defined as the systematic study of the social behavioral dynamics of disasters utilizing computational methods. In this paper, we discuss and showcase the opportunities and the challenges in this new approach to disaster research. Following a brief review of the fields that relate to CSSD, namely traditional social sciences of disasters, computational social science, and crisis informatics, we examine how advances in Internet technologies offer a new lens through which to study disasters. By identifying gaps in the literature, we show how this new field could address ways to advance our understanding of the social and behavioral aspects of disasters in a digitally connected world. In doing so, our goal is to bridge the gap between data science and the social sciences of disasters in rapidly changing environments.

Keywords: Disasters; Computational Social Science; Crisis Informatics; Disaster Modeling, Web 2.0; Social Media; Big Data; Volunteered Geographical Information; Crowdsourcing.
Figure 2: Interactions of data analysis, computational models, and social theory
in computational social science of disasters.

Full Reference:
Burger, A., Oz, T., Kennedy, W.G. and Crooks, A.T. (2019), Computational Social Science of Disasters: Opportunities and Challenges, Future Internet, 11(5): 103; https://doi.org/10.3390/fi11050103. (pdf)

Thursday, September 22, 2016

The study of slums as social and physical constructs: challenges and emerging research opportunities

Conceptual model for integrating social
and physical constructs to monitor,
analyze and model slums.


Continuing our research on slums, we have just had a paper published in the journal Regional Studies, Regional Science entitled "The Study of Slums as Social and Physical Constructs: Challenges and Emerging Research Opportunities". In this open access publication we review past lines of research with respect to studying slums which often focus on one of three constructs: (1) exploring the socio-economic and policy issues; (2) exploring the physical characteristics; and, lastly, (3) those modelling slums. We argue that while such lines of inquiry have proved invaluable with respect to studying slums, there is a need for  a  more  holistic  approach  for  studying  slums  to truly understand  them at the local, national and regional scales. Below you can read the abstract of our paper:
"Over 1 billion people currently live in slums, with the number of slum dwellers only expected to grow in the coming decades. The vast majority of slums are located in and around urban centres in the less economically developed countries, which are also experiencing greater rates of urbanization compared with more developed countries. This rapid rate of urbanization is cause for significant concern given that many of these countries often lack the ability to provide the infrastructure (e.g., roads and affordable housing) and basic services (e.g., water and sanitation) to provide adequately for the increasing influx of people into cities. While research on slums has been ongoing, such work has mainly focused on one of three constructs: exploring the socio-economic and policy issues; exploring the physical characteristics; and, lastly, those modelling slums. This paper reviews these lines of research and argues that while each is valuable, there is a need for a more holistic approach for studying slums to truly understand them. By synthesizing the social and physical constructs, this paper provides a more holistic synthesis of the problem, which can potentially lead to a deeper understanding and, consequently, better approaches for tackling the challenge of slums at the local, national and regional scales."

Keywords: Slums; informal settlements; socio-economic; remote sensing; crowdsourced information; modelling.
Framework for studying and understanding slums.


We hope you enjoy this paper and we wound be interested in receiving any feedback.

Full Reference:
Mahabir, R., Crooks, A.T., Croitoru, A. and Agouris, P. (2016), “The Study of Slums as Social and Physical Constructs: Challenges and Emerging Research Opportunities”, Regional Studies, Regional Science, 3(1): 737-757. (pdf)

Friday, May 13, 2016

A Semester with Urban Analytics

This past semester I gave a new class at GMU entitled "Urban Analytics". In a nutshell the class was about introducing students to a broad interdisciplinary field that focuses on the use of data to study cities. More specifcally the emphasis of the class was to provide students with a understanding of what methods, tools and theory can be used to monitor, analyze and model cities. 

From my past research and also when preparing the class material,  I have come to the realization that to study cities (like many others, you know who you are) that there is no one general model, tool or dataset. Therefore, one needs to maintain a toolbox of specialized tools than can be applied to different aspects of urban problems and questions. 

The toolbox that we used in class included a variety of software such as ArcGIS, QGIS, GeoDa, SANET along with programing and scripting in Python and R to modeling  cities via UrbanSim, NetLogo and MASON. Data we used ranged from crowdsourced (e.g. volunteered geographical information) data such as from OpenStreetMap or Wikipedia, to crowd harvested (ambient geographical information) data such as Twitter and Flickr, as-well as more traditional sources of data such as the US Census.

The Urban Analytics Toolbox

As an introduction to urban analytics, the course had the following objectives:
  1. to understand the motivation for the use of data to study cities, including some historical aspects; 
  2. to learn about the variety of Urban Analytics research programs across the several disciplines (urban planning, regional science, public policy, geography, computational social science etc.), through a survey of the literature and case studies. 
  3. to understand the distinct contribution that Urban Analytics can make by providing specific insights about cities at multiple scales. 
  4. to provide the foundations for more advanced work in the area of Urban Analytics. 
As with many of my courses, students were expected to complete a end of semester project. Below is a selection of these projects which explored some aspect of urban life.



I would like to thank the students for participating in this new class. It was a fun trip.

Wednesday, April 06, 2016

Crowdsourcing A Collective Sense of Place


Following on with our GeoSocial Analysis work, we recently had a paper published in  PLOS ONE entitled "Crowdsourcing A Collective Sense of Place." In the paper we discuss and showcase how one can take a quantitative approach to derive a collective sense of place from Twitter contributions and also from corresponding Wikipedia entries.

To illustrate this we present a brief study three cities, that of New York, Los Angeles and Singapore. Below you read the abstract of the paper, see some images from the paper especially the flowchart describing overall process used to discover platial alignment, along with the full reference to the paper.
Place can be generally defined as a location that has been assigned meaning through human experience, and as such it is of multidisciplinary scientific interest. Up to this point place has been studied primarily within the context of social sciences as a theoretical construct. The availability of large amounts of user-generated content, e.g. in the form of social media feeds or Wikipedia contributions, allows us for the first time to computationally analyze and quantify the shared meaning of place. By aggregating references to human activities within urban spaces we can observe the emergence of unique themes that characterize different locations, thus identifying places through their discernible sociocultural signatures. In this paper we present results from a novel quantitative approach to derive such sociocultural signatures from Twitter contributions and also from corresponding Wikipedia entries. By contrasting the two we show how particular thematic characteristics of places (referred to herein as platial themes) are emerging from such crowd-contributed content, allowing us to observe the meaning that the general public, either individually or collectively, is assigning to specific locations. Our approach leverages probabilistic topic modelling, semantic association, and spatial clustering to find locations are conveying a collective sense of place. Deriving and quantifying such meaning allows us to observe how people transform a location to a place and shape its characteristics.  

Keywords: Online Encyclopedias, Twitter, Data Visualization, Social Media, Semantics.

Flowchart describing overall process used to discover platial alignment.

Statistically significant clusters of recreation and entertainment categories concentrated over Manhattan, NYC

Maps depict significant hotspots for each of the high-level categories for (A) Singapore, (B) London, (C) Los Angeles, and (D) New York City.

Full Reference: 
Jenkins A., Croitoru, A. Crooks, A.T. and Stefanidis, A. (2016), Crowdsourcing A Collective Sense of Place, PLoS ONE 11(4): e0152932. doi:10.1371/journal.pone.0152932