Showing posts with label AGI. Show all posts
Showing posts with label AGI. Show all posts

Thursday, July 04, 2019

Challenges and Opportunities of Social Media Data for Socio-environmental Systems Research

SES diagram with examples of topics that
have been researched using social media data
While I have written about how one can use social media data to study cities, health issues etc... more recently we have been looking into how such data can be used to aid  Socio-environmental Systems (SES) research. SES are defined as tightly linked social and biophysical subsystems that mutually influence one another through positive and negative feedbacks.  To this end, Bianca Lopez, Nick Magliocca and myself just ahd a paper published in Land entitled "Challenges and Opportunities of Social Media Data for Socio-environmental Systems Research." 

In this paper we discuss SES and how research into them poses many challenges, not least of which are collecting or compiling data at the appropriate scales and aligning social and environmental data to address SES questions.  We discuss how SES have been studied using more traditional sources of data (e.g. census data, remote sensing etc.) and explore how social media can be used in the context of SES research. Specifically we ask three specific questions. 1) How can feedback between social and environmental systems be meaningfully studied using social media data? 2) How can using social media data re-frame or compliment current SES research questions and methods? and 3) Are there best practices for collecting and validating social media data for use in SES research? If these questions sound interesting to you, we encourage you to read the abstract below or the full paper.

Abstract:
Social media data provide an unprecedented wealth of information on people’s perceptions, attitudes, and behaviors at fine spatial and temporal scales and over broad extents. Social media data produce insight into relationships between people and the environment at scales that are generally prohibited by the spatial and temporal mismatch between traditional social and environmental data. These data thus have great potential for use in socio-environmental systems (SES) research. However, biases in who uses social media platforms and what they use them for create uncertainty in the potential insights from these data. Here, we describe ways that social media data have been used in SES research, including tracking land-use and environmental changes, natural resource use, and ecosystem service provisioning. We also highlight promising areas for future research and present best practices for SES research using social media data.

Keywords: social media; socio-ecological systems; human-environment interactions; geospatial analysis; crowdsourced data.
Example of information provided by social media posts and how it is used in analyses. A single post from a social media user.

Example of information on people’s use of natural resources from social media data based on key word searches fish and oyster from Twitter, Instagram and Foursquare.

Full Reference:
Lopez, B., Magliocca, N. and Crooks, A.T. (2019), Challenges and Opportunities of Social Media Data for Socio-environmental Systems Research, Land, 8(7), 107; https://www.mdpi.com/2073-445X/8/7/107/htm  (pdf)

 

Update: This review paper was awarded second prize for Best Review Paper Award in Land 2019.


Friday, September 21, 2018

Exodus 2.0: Crowdsourcing Geographical and Social Trails of Mass Migration

Readers of the blog might know we have an interest in volunteered geographic information, social media and Web 2.0 technologies and how they can be used to explore urban systems. Recently however, we turned our focus on how such information and technologies can be used to explore and understand mass migrations.

To this end we recently had a paper published in the Journal of Geographical Systems entitled "Exodus 2.0: Crowdsourcing Geographical and Social Trails of Mass Migration". We adopt the term Exodus 2.0 to refer to this new migration paradigm in the digital age, whereby information is a commodity in the migration process.

Given the nature of migration processes, it is possible to explore them across two key dimensions: geographical and situational. The geographical dimension is associated with the physical migration pathways migrants take from a country of origin to a destination site (often through a number of intermediate “stop” sites). The situational dimension is associated with the social connectivity of moving migrant populations, the conditions on the ground, and the activities that take place as part of migration efforts (including the root conditions, proximate conditions and triggering events).
Factors that potentially cause refugee production and
 mass movement based on identified factors detailed by
Clark (1989) and Zottarelli (1998).
In the paper, we use the ongoing Syrian humanitarian crisis as a case study to to explore how the factors that potentially causes refugee production and mass movement  can be gleamed from new sources of data. Specifically, the potential of crowd-generated data—especially open data, volunteered geographic information and social media content (e.g. OpenStreetMap, Flickr, Twitter and Instagram)  to provide information about migration processes.  Through a series of case studies  we show how such data (when combined with more traditional data sources) offers a new lens to study such the geographical and situational dimensions of mass migration. Finally we discuss  how such data could be used to inform migration modeling. If we have not bored you yet and you are interested in finding out more about this line of inquiry, below we provide the abstract to the paper, some of the figures which go along with our analysis for studying the refugee production and movement. Finally, we also provide the full reference and a link to the paper. 

Abstract:
The exodus of displaced populations is a recurring historical phenomenon, and the ongoing Syrian humanitarian crisis is its latest incarnation. During such mass migration events, information is an essential commodity. Of particular importance is geographical (e.g., pathways and refugee camps) and social (e.g., refugee activities and networking) information. Traditionally, such information had been produced and disseminated by authorities, but a new paradigm is emerging: Web 2.0 and mobile computing technologies enable the involved stakeholder communities to produce, access, and consume migration-related information. The purpose of this article is to put forward a new typology for understanding the factors around migration and to examine the potential of crowd-generated data—especially open data and volunteered geographic information—to study such events. Using the recent wave of migration to Europe from the Middle East and northern Africa as a case study, we examine how migration-related information can be dynamically mined and analyzed to study the migrants’ pathways from their home countries to their destination sites, as well as the conditions and activities that evolve during the migration process. These new data sources can provide a deeper and more fine-grained understanding of the migration process, often in real-time, and often through the eyes of the communities affected by it. Nevertheless, this also raises significant methodological and technical challenges for their future use associated with potential biases, data quality issues, and data processing.

Keywords: Refugees, Forced migration, Humanitarian crisis, Volunteered geographic information, Crowdsourcing, Social media, GIS, Web 2.0.
Cumulative flow (2011–2015) illustrating Syrian forced migration to neighboring countries and other destination countries. Line thickness indicates increasing number of persons migrating.

Retweet network of geolocated Twitter microblogs that are discussing opinions, news and retweeting information related to “refugee” in multiple languages from May to August 2017.

A concept graph illustrating the associations between a keyword related to root factors of mass migration such as poverty (“welfare”) to other keywords, as they appear in our Twitter data corpus. The color of the node refers to specific themes: locations (green), actors (dark red), topics (red), entities and individuals (blue), concepts (white), and events (yellow). Red edges represent active associations between terms; gray edges represent inactive associations between terms.

An agent-based model of migration: top: the spatial environment, where the lines represent migration pathways, and the nodes represent number of migrants. Purple nodes represent final destination sites, red nodes show migrant deaths, and green nodes show migrants en route (source: Hu 2016).

Full Reference: 
Curry, T., Croitoru, A., Crooks, A.T. and Stefanidis, A. (in press), Exodus 2.0: Crowdsourcing Geographical and Social Trails of Mass Migration, Journal of Geographical Systems. DOI: https://doi.org/10.1007/s10109-018-0278-1 (pdf)

Wednesday, January 24, 2018

A Review of High and Very High Resolution Remote Sensing Approaches for Detecting and Mapping Slums

Regular readers of this site might of noticed that we have an interest in slums. In the past this has focused on modeling them from an agent-based perspective, comparing volunteered geographical information to more authoritative data on slums, to that of attempting to come up with a Slum Severity Index. However, more recently we have taken to looking at how remote sensing approaches have been and can be used to detect and map slums.

To this end we recently had a review paper accepted in Urban Systems entitled "A Critical Review of High and Very High Resolution Remote Sensing Approaches for Detecting and Mapping Slums: Trends, Challenges and Emerging Opportunities". In this paper we carry out a comprehensive review of studies that have used high and very high resolution (H/VH-R) remote sensing techniques to detect and map slums (along with their global footprint). We discuss approaches used (e.g. multi-scale, image texture analysis, landscape analysis, object-based image analysis, building feature extraction, data mining, socio-economic measures) using H/VH-R imagery for identifying and mapping slums, listing what are the limitations and advantages of each. After this, we  discuss emerging sources of geospatial data that should we thing should be considered (e.g., volunteer geographic information, VGI, social media) in conjunction with growing trends and advancements in technology (e.g., geosensor networks, unmanned aerial vehicles (UAVs) or “drones) when trying to map and monitor slums. We argue that it is only through such data integration and analysis that we can then create a benchmark for determining the most suitable methods for mapping slums in a given locality. Below you can read the abstract of the paper and see some of the figures we use to support our discussion, along with the full reference.

Abstract: Slums are a global urban challenge, with less developed countries being particularly impacted. To adequately detect and map them, data is needed on their location, spatial extent and evolution. High- and very high-resolution remote sensing imagery has emerged as an important source of data in this regard. The purpose of this paper is to critically review studies that have used such data to detect and map slums. Our analysis shows that while such studies have been increasing over time, they tend to be concentrated to a few geographical areas and often focus on the use of a single approach (e.g., image texture and object-based image analysis), thus limiting generalizability to understand slums, their population, and evolution within the global context. We argue that to develop a more comprehensive framework that can be used to detect and map slums, other emerging sourcing of geospatial data should be considered (e.g., volunteer geographic information) in conjunction with growing trends and advancements in technology (e.g., geosensor networks). Through such data integration and analysis we can then create a benchmark for determining the most suitable methods for mapping slums in a given locality, thus fostering the creation of new approaches to address this challenge.
Keywords: high and very high resolution imagery; remote sensing, slums; geosensor networks; image analysis.

Global distribution of urban and slum populations.

Country level distribution of H/VH-R studies (studies published between 1997-2016).

OSM and Google Maps views of Kibera slum (a) Top:Left OSM and right Google Maps (b) Bottom:Left OSM and right Google Maps.

Full Reference:
Mahabir, R., Croitoru, A., Crooks, A.T., Agouris, P. and Stefanidis, A. (2018), A Critical Review of High and Very High Resolution Remote Sensing  Approaches for Detecting and Mapping Slums: Trends, Challenges and Emerging Opportunities, Urban Science. 2(1), 8; doi:10.3390/urbansci2010008 (pdf)
As always, any thoughts or comments are most welcome.

Friday, March 10, 2017

Geovisualization of Social Media

Figure 1: Map Mashup of Twitter data, where eachdot
represents a tweet, the text corresponds to the selected
 tweet marked with a star
In the recently released "The International Encyclopedia of Geography: People, the Earth, Environment, and Technology" we were asked to write a brief entry entitled "geovisualization of social media". Below is a summary of  our chapter:

The proliferation of social media over the last decade is presenting substantial computational challenges associated with the management, processing, analysis and visualization of the corresponding massive volumes of data. Furthermore, this new form of information also imposes new-found challenges upon the geographical community due to the unique nature of its content, as analyzing such data calls for a hybrid mix of spatial and social analysis. The spatial content of social media comprises primarily coordinates from which the contributions originate, or references to specific locations. At the same time, these data have a strong social component, as they can reveal the underlying social structure of the user community through manifestations of their interactions. Analyzing both the spatial and social content of social media feeds is referred to as geosocial analysis. Within this entry we explore the geovisualization opportunities and challenges that are emerging as social media are becoming the subject of study of the geographical community.
In more detail, we start off discussing how the geographic content of social media feeds represents a new type of geographic information. It transcends the early definitions of crowdsourcing or volunteered geographic information as it is not the product of a process through which citizens explicitly and purposefully contribute geographic information to update or expand geographic databases. Instead, the type of geographic information that can be harvested from social media feeds can be referred to as Ambient Geographic Information; it is embedded in the content of these feeds, often across the content of numerous entries rather than within a single one, and has to be somehow extracted. Nevertheless, it is of great importance as it communicates instantaneously information about emerging issues. At the same time, it provides an unparalleled view of the complex social networking and cultural dynamics within a society, and captures the temporal evolution of the human landscape.

In many cases, the geovisualization of social media feeds predominately take the appearance of web map mashups, in essence portraying the location of social media usage on a map. Such an early attempt to visualize social media is shown Figure 1. We argue that while this approach is informative, it often falls short of capturing the depth, richness, and complexity of the information that can be gleaned from social data. As a result, a need for more advanced geovisualization approaches that are capable of better capturing and communicating the complexity and multidimensionality of social media arises. And this is the focus of our chapter. We discuss briefly the geovisualization of network structures (such as shown in Figure 2), the geovisualization of network structure dynamics, the geovisualization of social media content (such as shown in Figure 3) along with the visualization of social media analysis (Figure 4) and conclude the chapter with a list of emerging research challenges.

Figure 2: Visualizing communities: a social network of an interest group (a), and the geovisualization of the  largest community shown over the contiguous U.S (B).

Figure 3: Visualizing social media content dynamics by coupling a Twitter stream viewer (A), a Twitter activity density map (B), and a ranked list top hash-tags (C) and top authors (E), a time slider (D), and author/hash-tags time series graphs.
Figure 4: Visualizing spatiotemporal clusters of tweets following the 2013 Boston bombing. Red circles indicate the approximate radius of each cluster, and color is used to indicate time.


We hope you enjoy. As always any feedback or comments most welcome. Please note this chapter was written a couple of years ago and more recent work by us has been done, click here to see some.

Full Reference:
Croitoru, A., Crooks, A.T., Radzikowski, J. and Stefanidis, A. (2017), Geovisualization of Social Media, in Richardson, D., Castree, N., Goodchild, M. F., Kobayashi, A. L., Liu, W. and Marston, R. (eds.), The International Encyclopedia of Geography: People, the Earth, Environment, and Technology, Wiley Blackwell. DOI: 10.1002/9781118786352.wbieg0605 (PDF)

Wednesday, April 06, 2016

Crowdsourcing A Collective Sense of Place


Following on with our GeoSocial Analysis work, we recently had a paper published in  PLOS ONE entitled "Crowdsourcing A Collective Sense of Place." In the paper we discuss and showcase how one can take a quantitative approach to derive a collective sense of place from Twitter contributions and also from corresponding Wikipedia entries.

To illustrate this we present a brief study three cities, that of New York, Los Angeles and Singapore. Below you read the abstract of the paper, see some images from the paper especially the flowchart describing overall process used to discover platial alignment, along with the full reference to the paper.
Place can be generally defined as a location that has been assigned meaning through human experience, and as such it is of multidisciplinary scientific interest. Up to this point place has been studied primarily within the context of social sciences as a theoretical construct. The availability of large amounts of user-generated content, e.g. in the form of social media feeds or Wikipedia contributions, allows us for the first time to computationally analyze and quantify the shared meaning of place. By aggregating references to human activities within urban spaces we can observe the emergence of unique themes that characterize different locations, thus identifying places through their discernible sociocultural signatures. In this paper we present results from a novel quantitative approach to derive such sociocultural signatures from Twitter contributions and also from corresponding Wikipedia entries. By contrasting the two we show how particular thematic characteristics of places (referred to herein as platial themes) are emerging from such crowd-contributed content, allowing us to observe the meaning that the general public, either individually or collectively, is assigning to specific locations. Our approach leverages probabilistic topic modelling, semantic association, and spatial clustering to find locations are conveying a collective sense of place. Deriving and quantifying such meaning allows us to observe how people transform a location to a place and shape its characteristics.  

Keywords: Online Encyclopedias, Twitter, Data Visualization, Social Media, Semantics.

Flowchart describing overall process used to discover platial alignment.

Statistically significant clusters of recreation and entertainment categories concentrated over Manhattan, NYC

Maps depict significant hotspots for each of the high-level categories for (A) Singapore, (B) London, (C) Los Angeles, and (D) New York City.

Full Reference: 
Jenkins A., Croitoru, A. Crooks, A.T. and Stefanidis, A. (2016), Crowdsourcing A Collective Sense of Place, PLoS ONE 11(4): e0152932. doi:10.1371/journal.pone.0152932

Thursday, January 08, 2015

Crowdsourcing Urban Form and Function

We have just had published a new paper entitled: "Crowdsourcing Urban Form and Function" in International Journal of Geographical Information Science which showcases some of our recent work with respect to cities and how new sources of information can be used to study urban morphology at a variety of spatial and temporal scales. Below is the abstract for the paper: 

"Urban form and function have been studied extensively in urban planning and geographic information science. However, gaining a greater understanding of how they merge to define the urban morphology remains a substantial scientific challenge. Towards this goal, this paper addresses the opportunities presented by the emergence of crowdsourced data to gain novel insights into form and function in urban spaces. We are focusing in particular on information harvested from social media and other open-source and volunteered datasets (e.g. trajectory and OpenStreetMap data). These data provide a first-hand account of form and function from the people who define urban space through their activities. This novel bottom-up approach to study these concepts complements traditional urban studies work to provide a new lens for studying urban activity. By synthesizing recent advancements in the analysis of open-source data we provide a new typology for characterizing the role of crowdsourcing in the study of urban morphology. We illustrate this new perspective by showing how social media, trajectory, and traffic data can be analyzed to capture the evolving nature of a city’s form and function. While these crowd contributions may be explicit or implicit in nature, they are giving rise to an emerging research agenda for monitoring, analyzing and modeling form and function for urban design and analysis."
This paper builds and extends considerably our prior work, with respect to crowdsourcing, volunteered and ambient geographic information. In the scope of this paper we use the term ‘urban form’ to refer to the aggregate of the physical shape of the city, its buildings, streets, and all other elements that make up the urban space. In essence, the geometry of the city. In contrast, we use the term ‘urban function’ to refer to the activities that are taking place within this space. To this end we contrast how crowdsourced data can related to more traditional sources of such information both explicitly and implicitly as shown in the table below. 

A typology of implicit and explicit form and function content

In addition, we also discuss in the paper how these new sources of data, which are often at finer resolutions than more authoritative data are allowing us to to customize the we we aggregate the data  at various geographical levels as shown below. Such aggregations can range from building footprints and addresses to street blocks (e.g. for density analysis), or street networks (e.g. for accessibility analysis). For large-scale urban analysis we can revert to the use of zonal geographies or grid systems.  
Aggregation methods for varied scales of built environment analysis

In the application section of the paper we highlight how we can extract implicit form and function from crowdsourced data. The image below for example, shows how we can take information from Twitter, and differentiate different neighborhoods over space and time.

Neighborhood map and topic modeling results showing the mixture of social functions in each area.

Finally in the paper, we outline an emerging research agenda related to the "persistent urban morphology concept" as shown below. Specifically how crowdsourcing is changing how we collect, analyze and model urban morphology. Moreover, how this new paradigm provides a new lens for studying the conceptualization of how cities operate, at much finer temporal, spatial, and social scales than we had been able to study so far.

The persistent urban morphology concept.

We hope you enjoy the paper.

Full Reference:  
Crooks, A.T., Pfoser, D., Jenkins, A., Croitoru, A., Stefanidis, A., Smith, D. A., Karagiorgou, S., Efentakis, A. and Lamprianidis, G. (2015), Crowdsourcing Urban Form and Function, International Journal of Geographical Information Science. DOI: 10.1080/13658816.2014.977905 (pdf)
 

Friday, April 13, 2012

#Earthquake: Twitter as a Distributed Sensor System

Our work on using social media continues to develop and we have recently had a paper accepted in Transactions in GIS, entitled "#Earthquake: Twitter as a Distributed Sensor System". Below we present our abstract and some of the results.
Social media feeds are rapidly emerging as a novel avenue for the contribution and dissemination of information that is often geographic. Their content often includes references to events occurring at, or affecting specific locations. Within this paper we analyze the spatial and temporal characteristics of the twitter feed activity responding to a 5.8 magnitude earthquake which occurred on the East Coast of the United States (US) on August 23, 2011. We argue that these feeds represent a hybrid form of a sensor system that allows for the identification and localization of the impact area of the event. By contrasting this to comparable content collected through the dedicated crowdsourcing ‘Did You Feel It?’ (DYFI) website of the US Geological Survey we assess the potential of the use of harvested social media content for event monitoring. The experiments support the notion that people act as sensors to give us comparable results in a timely manner, and can complement other sources of data to enhance our situational awareness and improve our understanding and response to such events.
The movie below show geolocated tweets with references to the earthquake through keyword (earthquake or earth and quake) and hashtag search (#earthquake or #quake) for the first hour after the earthquake.





The following images give a glimpse at some of our analysis.
Response pattern as function of distance from epicenter for the first 400 seconds after the earthquake. At the top we see a plot of (reaction time, distance) of all tweets during that period. At the bottom we show the histogram of the number of tweets as a function of distance.
Locations of the 40 tweets in the shaded area of the figure above overlaid over the USGS CDI scale map. Tweet locations are marked as green circles. Color-coding in the graph is ranging from red (high perceived intensity) to yellow (lower perceived intensity). The dashed line shows a distance of approximately 950 km (8.5 degrees of angular distance) from the epicenter.
The movie below gives you an idea of some of the tweet content:





Full reference to this paper is:
Crooks, A.T., Croitoru, A., Stefanidis, A. and Radzikowski, J. (2013), #Earthquake: Twitter as a Distributed Sensor System, Transactions in GIS, 17(1): 124-147. (pdf)

Friday, December 16, 2011

Occupy Wall Street movement via Twitter

Following on from our work on harvesting ambient geospatial information (AGI) from social media feeds we have started to explore the Occupy Wall Street movement. The movie below shows just one part of this work, specifically the movement of the protesters in New York during the Action Day (November 17) from Wall Street to Brooklyn Bridge. The red dots denote locations of the tweets. Selected tweets are displayed at the bottom of the screen. Active tweets are marked with a white star.






More ananylsis to follow...

Tuesday, December 06, 2011

Harvesting ambient geospatial information from social media feeds

A paper I  recently co-authored with Anthony Stefanidis and Jacek Radzikowski from George Mason University entitled "Harvesting ambient geospatial information from social media feeds" is now available in  GeoJournal. 
 
The abstract for the paper reads as follows: "Social media generated from many individuals is playing a greater role in our daily lives and provides a unique opportunity to gain valuable insight on information flow and social networking within a society. Through data collection and analysis of its content, it supports a greater mapping and understanding of the evolving human landscape. The information disseminated through such media represents a deviation from volunteered geography, in the sense that it is not geographic information per se. Nevertheless, the message often has geographic footprints, for example, in the form of locations from where the tweets originate, or references in their content to geographic entities. We argue that such data conveys ambient geospatial information, capturing for example, people’s references to locations that represent momentary social hotspots. In this paper we address a framework to harvest such ambient geospatial information, and resulting hybrid capabilities to analyze it to support situational awareness as it relates to human activities. We argue that this emergence of ambient geospatial analysis represents a second step in the evolution of geospatial data availability, following on the heels of volunteered geographical information."

Geolocating pairs of tweeters and retweeters