Showing posts with label Volunteered Geographic Information. Show all posts
Showing posts with label Volunteered Geographic Information. Show all posts

Wednesday, June 10, 2020

New Paper: A Thematic Similarity Network Approach for Analysis of Places Using VGI

Building upon our work on volunteered geographical information (VGI) and ambient geographic information (AGI) and how such data (e.g. social media) can be used to understand place, Xiaoyi Yuan, Andreas Züfle and myself have a new paper entitled: "A Thematic Similarity Network Approach for Analysis of Places Using Volunteered Geographic Information" in the ISPRS International Journal of Geo-InformationIn this paper we use textual data from crowdsourced reviews originating with TripAdvisor and geo-located Twitter data and leverage this unstructured geographical information to comprehend the complexity of places at scale. Specifically we explore the connectedness and relationships of places through thematic (i.e., topical) similarity networks using Manhattan, New York as a case study. If such work sounds of interest to you, below we provide the abstract to the paper in order for you to gain a greater understanding of work, along with some figures that show our workflow and how communities where connected, before presenting some of our results. Finally at the bottom of the post, the full reference and a link to the paper is provided.  For those interested in extending or utilizing this work. The python code for presented in our analysis is available at: https://bitbucket.org/xiaoyiyuan/network_vgi/

Abstract:
The research presented in this paper proposes a thematic network approach to explore rich relationships between places. We connect places in networks through their thematic similarities by applying topic modeling to the textual volunteered geographic information (VGI) pertaining to the places. The network approach enhances previous research involving place clustering using geo-textual information, which often simplifies relationships between places to be either in-cluster or out-of-cluster. To demonstrate our approach, we use as a case study in Manhattan (New York) that compares networks constructed from three different geo-textural data sources --TripAdvisor attraction reviews, TripAdvisor restaurant reviews, and Twitter data. The results showcase how the thematic similarity network approach enables us to conduct clustering analysis as well as node-to-node and node-to-cluster analysis, which is fruitful for understanding how places are connected through individuals’ experiences. Furthermore, by enriching the networks with geodemographic information as node attributes, we discovered that some low-income communities in Manhattan have distinctive restaurant cultures. Even though geolocated tweets are not always related to place they are posted from, our case study demonstrates that topic modeling is an efficient method to filter out the place-irrelevant tweets and therefore refining how of places can be studied.

Keywords: Geo-Textual Data, Volunteered Geographic Information, Crowdsourcing, Similarity Network Analysis, Topic Modeling

Work flow from data input to the construction of the thematic similarity network and analysis (i.e., community detection and unique nodes discovery).

A stylized network demonstrating the process of community detection from a fully-connected similarity network.


Network visualization of all communities from the thematic similarity networks with major communities highlighted. Only the major communities are shown on the map for the sake of clarity. Major communities in Network visualization and mapping for each network are colored the same and thus the legend applies for both.


Two examples of communities with boundary nodes and their respective topics.

Full Reference:
Yuan X., Crooks, A.T. and Züfle, A. (2020), A Thematic Similarity Network Approach for Analysis of Places Using Volunteered Geographic Information, ISPRS International Journal of Geo-Information,  9(6), 385, https://doi.org/10.3390/ijgi9060385. (pdf)

Tuesday, May 26, 2020

Crowdsourcing Street View Imagery: A Comparison of Mapillary and OpenStreetCam


In the past we have written extensively on Volunteered Geographic Information (VGI) such as OpenStreetMap or Twitter. However, we have not really explored Street View Imagery  (SVI), well not until now. Within the realm of VGI, SVI has emerged in recent years as a novel and rich source of data on cities from which geographic information can be derived.

Perhaps the most well-known example of SVI utilization is that of Google Street View (GSV). While SVI has been traditionally collected by governmental agencies and companies alike, we are now also witnessing the emergence of Volunteered Street View Imagery (VSVI), which relies on a crowdsourced effort to provide geotagged street-level imagery coverage of traversable pathways (e.g., a street or trail). Such imagery, similar to GSV, provides detailed information about the location of objects such as cars, road markings, traffic lights and signs, and allows for the automatic extraction of features at scale. Such imagery can also be mined using machine learning algorithms to automatically derive points of interest (POI) databases (e.g., locations of coffee shops and fire hydrants) without the intervention of the citizen.

To explore VSVI we have just published a new paper entitled: "Crowdsourcing Street View Imagery: A Comparison of Mapillary and OpenStreetCam" in the ISPRS International Journal of Geo-Information. In this paper we examine VSVI data collected from two different platforms: Mapillary and OpenStreetCam (OSC) for four metropoiltan areas in the United States (i.e., Washington (District of Columbia), San Francisco (California), Phoenix (Arizona), and Detroit (Michigan)). Both of these online platforms accept sequences of images captured from mobile devices and uploaded via an app on the device (like those shown in the image to the right). Images are geolocated using the device’s global positioning system (GPS). More specifically the paper examines:
  • the level of spatial coverage of each platform in order to assess the overall potential of such platforms to provide adequate coverage of geographic information.
  • user contribution patterns in Mapillary and OSC in order to understand how users are contributing to these platforms.
Results from our systematic and quantitative analysis of these two emerging VGI sources indicate that most Mapillary and OSC contributions occurred along control-access highways and local roads, and that the overall coverage in these sources is variable in comparison to an authoritative source (i.e., TIGER). Furthermore, our results showed that while the number of contributors varied across sites, only a few contributors were responsible for producing most of the raw data. User contribution patterns were also different in Mapillary and OSC. Specifically, we found that while patterns in coverage were variable for the different OSC sites, coverage patterns in Mapillary tended to be similar among sites. This finding may be linked to several factors, including differences in mapping practice, or issues with participation inequality, a topic that has been highly researched for other VGI platforms such as OSM, but which is still lacking within VSVI. Lastly, user contributions in Mapillary tended to be higher around 8:00 am, 1:00 pm and 5:00 pm (local time). This finding suggests that VSVI contributions tend to coincide with the morning and afternoon commute, and the lunch hour of the contributors.

If you wish to find out more about this work below we provide the abstract to the paper, a visual flowchart of our workflow and some of our our results. The full reference and link to the paper is provided at the bottom of the post.

Abstract:
Over the last decade, Volunteered Geographic Information (VGI) has emerged as a viable source of information on cities. During this time, the nature of VGI has been evolving, with new types and sources of data continually being added. In light of this trend, this paper explores one such type of VGI data: Volunteered Street View Imagery (VSVI). Two VSVI sources, Mapillary and OpenStreetCam, were extracted and analyzed to study road coverage and contribution patterns for four US metropolitan areas. Results show that coverage patterns vary across sites, with most contributions occurring along local roads and in populated areas. We also found that a few users contributed most of the data. Moreover, the results suggest that most data are being collected during three distinct times of day (i.e., morning, lunch and late afternoon). The paper concludes with a discussion that while VSVI data is still relatively new, it has the potential to be a rich source of spatial and temporal information for monitoring cities.

Keywords: Crowdsourcing; Volunteered Geographic Information; Street View Imagery; Mapillary, OpenStreetCam
Overview of methodology

Spatial distribution of road networks.
Spatial comparison of roads in kilometers.


Full Reference: 
Mahabir, R., Schuchard, R., Crooks, A.T., Croitoru, A. and Stefanidis, A. (2020), Crowdsourcing Street View Imagery: A Comparison of Mapillary and OpenStreetCam, ISPRS International Journal of Geo-Information. 9(6), 341; https://doi.org/10.3390/ijgi9060341 (pdf)

Thursday, July 04, 2019

Challenges and Opportunities of Social Media Data for Socio-environmental Systems Research

SES diagram with examples of topics that
have been researched using social media data
While I have written about how one can use social media data to study cities, health issues etc... more recently we have been looking into how such data can be used to aid  Socio-environmental Systems (SES) research. SES are defined as tightly linked social and biophysical subsystems that mutually influence one another through positive and negative feedbacks.  To this end, Bianca Lopez, Nick Magliocca and myself just ahd a paper published in Land entitled "Challenges and Opportunities of Social Media Data for Socio-environmental Systems Research." 

In this paper we discuss SES and how research into them poses many challenges, not least of which are collecting or compiling data at the appropriate scales and aligning social and environmental data to address SES questions.  We discuss how SES have been studied using more traditional sources of data (e.g. census data, remote sensing etc.) and explore how social media can be used in the context of SES research. Specifically we ask three specific questions. 1) How can feedback between social and environmental systems be meaningfully studied using social media data? 2) How can using social media data re-frame or compliment current SES research questions and methods? and 3) Are there best practices for collecting and validating social media data for use in SES research? If these questions sound interesting to you, we encourage you to read the abstract below or the full paper.

Abstract:
Social media data provide an unprecedented wealth of information on people’s perceptions, attitudes, and behaviors at fine spatial and temporal scales and over broad extents. Social media data produce insight into relationships between people and the environment at scales that are generally prohibited by the spatial and temporal mismatch between traditional social and environmental data. These data thus have great potential for use in socio-environmental systems (SES) research. However, biases in who uses social media platforms and what they use them for create uncertainty in the potential insights from these data. Here, we describe ways that social media data have been used in SES research, including tracking land-use and environmental changes, natural resource use, and ecosystem service provisioning. We also highlight promising areas for future research and present best practices for SES research using social media data.

Keywords: social media; socio-ecological systems; human-environment interactions; geospatial analysis; crowdsourced data.
Example of information provided by social media posts and how it is used in analyses. A single post from a social media user.

Example of information on people’s use of natural resources from social media data based on key word searches fish and oyster from Twitter, Instagram and Foursquare.

Full Reference:
Lopez, B., Magliocca, N. and Crooks, A.T. (2019), Challenges and Opportunities of Social Media Data for Socio-environmental Systems Research, Land, 8(7), 107; https://www.mdpi.com/2073-445X/8/7/107/htm  (pdf)

 

Update: This review paper was awarded second prize for Best Review Paper Award in Land 2019.


Wednesday, December 05, 2018

Detecting and Mapping Slums using Open Data

Urban and slum areas in Nairobi (False composite image created
by stacking image bands 7, 6 and 4 from the Landsat 8 satellite.
Turning back to slums, we just published paper entitled "Detecting and Mapping Slums using Open Data: A Case Study in Kenya" in the International Journal of Digital Earth. This work builds and extends our previous research on using new sources of data to explore the slum settlements in 3 cities in Kenya (i.e. Nairobi, Mombasa and Kisumu). Specifically, we examine how the fusion of Volunteered Geographical Information, Social Media, and other open data sources can complement remote sensing imagery in supporting slum detection, mapping and monitoring. 

We do this by using data mining tools (e.g. logistic regression, discriminant analysis and the See5 decision tree), to develop context-sensitive definitions for slums based on location, as well as for testing the generalizability of indicators and derived slum models. The end result is an indicator database for slums using open sources of physical and socio-economic data that can be used to characterize slum settlements. If you wish to know more, below we provide the abstract to the paper along with some of the figures and the full citation with a link to the paper itself.

Abstract:
The worldwide slum population currently stands at over one billion, with substantial growth expected in the coming decades. Traditionally, slums have been mapped using information derived mainly from either physical indicators using remote sensing data, or socio-economic indicators using census data. Each data source on its own provides only a partial view of slums, an issue further compounded by data poverty in less developed countries. To overcome such issues, this paper explores the fusion of traditional with emerging open data sources and data mining tools to identify additional indicators that can be used to detect and map the presence of slums, map their footprint, and map their evolution. Towards this goal, we develop an indicator database for slums using open sources of physical and socio-economic data that can be used to characterize slum settlements. Using this database, we then leverage data mining techniques to identify the most suitable combination of these indicators for mapping slums. Using three cities in Kenya as test cases, results show that the fusion of these data can improve the mapping accuracy of slums. These results suggest that the proposed approach can provide a viable solution to the emerging challenge of monitoring the growth of slums.
Keywords: Slums; Remote Sensing; Socio-economic; Urban sustainability; Data mining; Kenya

Study areas in Kenya

Methodology workflow

Distribution of positive classified cases for slums for (a) logistic regression, (b) discriminant analysis and (c) the See5 decision tree.
Full Reference:
Mahabir, R., Agouris, P., Stefanidis, A., Croitoru, A. and Crooks, A.T. (2018), Detecting and Mapping Slums using Open Data: A Case Study in Kenya, International Journal of Digital Earth. DOI: https://doi.org/10.1080/17538947.2018.1554010. (pdf)

Thursday, November 08, 2018

Refugee Camps and Volunteered Geographical Information

Fig. 7. Stimulus-Awareness-Activism (SA2) framework
Previously we have posted on how one can use new sources of data  (e.g. Volunteered Geographical Information) to explore and understand the world around us, such as mass migration, urban form and function, or be used for the basis of a model. Continuing on with this research theme we recently had a paper published in PLoS ONE entitled: "News Coverage, Digital Activism, and Geographical Saliency: A Case Study of Refugee Camps and Volunteered Geographical Information."

In this paper we explore the relationship between news coverage (via Google news), search trends (via Google trends) and user edit contribution patterns in OpenStreetMap and  Wikipedia for refugee camps from around the world. Specifically we are interested in how news media coverage (and in particular digital media) impacts digital activism (i.e.  volunteers who contribute content to online communities). Based on our analysis we find that digital activism bursts tend to take place during periods of sustained build-up of public awareness deficit or surplus.

These findings are in line with two prominent mass communication theories: agenda setting and corrective action, and suggest the emergence of a novel Stimulus-Awareness-Activism (SA2) framework in today’s participatory digital age. We argue that this paper brings us one step closer to understanding the underlying mechanisms that drive digital activism in particular in the geospatial domain. Below you can read the abstract of the paper, see the refugee camps we studied and some of the results. At the bottom of the post we also provide the full reference and a link to the paper.

Abstract:
The last several decades have witnessed a shift in the way in which news is delivered and consumed by users. With the growth and advancements in mobile technologies, the Internet, and Web 2.0 technologies users are not only consumers of news, but also producers of online content. This has resulted in a novel and highly participatory cyber-physical news awareness ecosystem that fosters digital activism, in which volunteers contribute content to online communities. While studies have examined the various components of this news awareness ecosystem, little is still known about how news media coverage (and in particular digital media) impacts digital activism. In order to address this challenge and develop a greater understanding of it, this paper focuses on a specific form of digital activism, that of the production of digital geographical content through crowdsourcing efforts. Using refugee camps from around the world as a case study, we examine the relationship between news coverage (via Google news), search trends (via Google trends) and user edit contribution patterns in OpenStreetMap, a prominent geospatial data crowdsourcing platform. In addition, we compare and contrast these patterns with user edit patterns in Wikipedia, a well-known non-geospatial crowdsourcing platform. Using Google news and Google trends to derive a measure of thematic public awareness, our findings indicate that digital activism bursts tend to take place during periods of sustained build-up of public awareness deficit or surplus. These findings are in line with two prominent mass communication theories: agenda setting and corrective action, and suggest the emergence of a novel stimulus-awareness-activism framework in today’s participatory digital age. Moreover, these findings further complement existing research examining the motivational factors that drive users to contribute to online collaborative communities. This paper brings us one step closer to understanding the underlying mechanisms that drive digital activism in particular in the geospatial domain.

Figure 1. Study areas  (centroid location of camp).

Figure 5. OSM, Wikipedia, Google News, and Google Trends time series during a -/+4 months period around the strongest extremum point of each camp. The figures show that whereas OSM and Wikipedia entries tend to come in bursts, Google News and Trends display a more sustained type of activity.

Figure 6. The public awareness curve versus the cumulative OSM and Wikipedia edit activity during a -/+4 months period around the strongest extremum point of each camp. For camps such as Nyarugusu, OSM and Wikipedia bursts overlap with public awareness surplus. In other camps, such as Bidibidi, OSM edit activity bursts coincide with public awareness deficit.  

Full Reference: 
Mahabir, R., Croitoru, A., Crooks, A.T., Agouris, P. and Stefanidis, A. (2018), News Coverage, Digital Activism, and Geographical Saliency: A Case Study of Refugee Camps and Volunteered Geographical Information, PLoS ONE, 13(11): e0206825.   https://doi.org/10.1371/journal.pone.0206825 (pdf)

Friday, September 21, 2018

Exodus 2.0: Crowdsourcing Geographical and Social Trails of Mass Migration

Readers of the blog might know we have an interest in volunteered geographic information, social media and Web 2.0 technologies and how they can be used to explore urban systems. Recently however, we turned our focus on how such information and technologies can be used to explore and understand mass migrations.

To this end we recently had a paper published in the Journal of Geographical Systems entitled "Exodus 2.0: Crowdsourcing Geographical and Social Trails of Mass Migration". We adopt the term Exodus 2.0 to refer to this new migration paradigm in the digital age, whereby information is a commodity in the migration process.

Given the nature of migration processes, it is possible to explore them across two key dimensions: geographical and situational. The geographical dimension is associated with the physical migration pathways migrants take from a country of origin to a destination site (often through a number of intermediate “stop” sites). The situational dimension is associated with the social connectivity of moving migrant populations, the conditions on the ground, and the activities that take place as part of migration efforts (including the root conditions, proximate conditions and triggering events).
Factors that potentially cause refugee production and
 mass movement based on identified factors detailed by
Clark (1989) and Zottarelli (1998).
In the paper, we use the ongoing Syrian humanitarian crisis as a case study to to explore how the factors that potentially causes refugee production and mass movement  can be gleamed from new sources of data. Specifically, the potential of crowd-generated data—especially open data, volunteered geographic information and social media content (e.g. OpenStreetMap, Flickr, Twitter and Instagram)  to provide information about migration processes.  Through a series of case studies  we show how such data (when combined with more traditional data sources) offers a new lens to study such the geographical and situational dimensions of mass migration. Finally we discuss  how such data could be used to inform migration modeling. If we have not bored you yet and you are interested in finding out more about this line of inquiry, below we provide the abstract to the paper, some of the figures which go along with our analysis for studying the refugee production and movement. Finally, we also provide the full reference and a link to the paper. 

Abstract:
The exodus of displaced populations is a recurring historical phenomenon, and the ongoing Syrian humanitarian crisis is its latest incarnation. During such mass migration events, information is an essential commodity. Of particular importance is geographical (e.g., pathways and refugee camps) and social (e.g., refugee activities and networking) information. Traditionally, such information had been produced and disseminated by authorities, but a new paradigm is emerging: Web 2.0 and mobile computing technologies enable the involved stakeholder communities to produce, access, and consume migration-related information. The purpose of this article is to put forward a new typology for understanding the factors around migration and to examine the potential of crowd-generated data—especially open data and volunteered geographic information—to study such events. Using the recent wave of migration to Europe from the Middle East and northern Africa as a case study, we examine how migration-related information can be dynamically mined and analyzed to study the migrants’ pathways from their home countries to their destination sites, as well as the conditions and activities that evolve during the migration process. These new data sources can provide a deeper and more fine-grained understanding of the migration process, often in real-time, and often through the eyes of the communities affected by it. Nevertheless, this also raises significant methodological and technical challenges for their future use associated with potential biases, data quality issues, and data processing.

Keywords: Refugees, Forced migration, Humanitarian crisis, Volunteered geographic information, Crowdsourcing, Social media, GIS, Web 2.0.
Cumulative flow (2011–2015) illustrating Syrian forced migration to neighboring countries and other destination countries. Line thickness indicates increasing number of persons migrating.

Retweet network of geolocated Twitter microblogs that are discussing opinions, news and retweeting information related to “refugee” in multiple languages from May to August 2017.

A concept graph illustrating the associations between a keyword related to root factors of mass migration such as poverty (“welfare”) to other keywords, as they appear in our Twitter data corpus. The color of the node refers to specific themes: locations (green), actors (dark red), topics (red), entities and individuals (blue), concepts (white), and events (yellow). Red edges represent active associations between terms; gray edges represent inactive associations between terms.

An agent-based model of migration: top: the spatial environment, where the lines represent migration pathways, and the nodes represent number of migrants. Purple nodes represent final destination sites, red nodes show migrant deaths, and green nodes show migrants en route (source: Hu 2016).

Full Reference: 
Curry, T., Croitoru, A., Crooks, A.T. and Stefanidis, A. (in press), Exodus 2.0: Crowdsourcing Geographical and Social Trails of Mass Migration, Journal of Geographical Systems. DOI: https://doi.org/10.1007/s10109-018-0278-1 (pdf)

Wednesday, January 24, 2018

A Review of High and Very High Resolution Remote Sensing Approaches for Detecting and Mapping Slums

Regular readers of this site might of noticed that we have an interest in slums. In the past this has focused on modeling them from an agent-based perspective, comparing volunteered geographical information to more authoritative data on slums, to that of attempting to come up with a Slum Severity Index. However, more recently we have taken to looking at how remote sensing approaches have been and can be used to detect and map slums.

To this end we recently had a review paper accepted in Urban Systems entitled "A Critical Review of High and Very High Resolution Remote Sensing Approaches for Detecting and Mapping Slums: Trends, Challenges and Emerging Opportunities". In this paper we carry out a comprehensive review of studies that have used high and very high resolution (H/VH-R) remote sensing techniques to detect and map slums (along with their global footprint). We discuss approaches used (e.g. multi-scale, image texture analysis, landscape analysis, object-based image analysis, building feature extraction, data mining, socio-economic measures) using H/VH-R imagery for identifying and mapping slums, listing what are the limitations and advantages of each. After this, we  discuss emerging sources of geospatial data that should we thing should be considered (e.g., volunteer geographic information, VGI, social media) in conjunction with growing trends and advancements in technology (e.g., geosensor networks, unmanned aerial vehicles (UAVs) or “drones) when trying to map and monitor slums. We argue that it is only through such data integration and analysis that we can then create a benchmark for determining the most suitable methods for mapping slums in a given locality. Below you can read the abstract of the paper and see some of the figures we use to support our discussion, along with the full reference.

Abstract: Slums are a global urban challenge, with less developed countries being particularly impacted. To adequately detect and map them, data is needed on their location, spatial extent and evolution. High- and very high-resolution remote sensing imagery has emerged as an important source of data in this regard. The purpose of this paper is to critically review studies that have used such data to detect and map slums. Our analysis shows that while such studies have been increasing over time, they tend to be concentrated to a few geographical areas and often focus on the use of a single approach (e.g., image texture and object-based image analysis), thus limiting generalizability to understand slums, their population, and evolution within the global context. We argue that to develop a more comprehensive framework that can be used to detect and map slums, other emerging sourcing of geospatial data should be considered (e.g., volunteer geographic information) in conjunction with growing trends and advancements in technology (e.g., geosensor networks). Through such data integration and analysis we can then create a benchmark for determining the most suitable methods for mapping slums in a given locality, thus fostering the creation of new approaches to address this challenge.
Keywords: high and very high resolution imagery; remote sensing, slums; geosensor networks; image analysis.

Global distribution of urban and slum populations.

Country level distribution of H/VH-R studies (studies published between 1997-2016).

OSM and Google Maps views of Kibera slum (a) Top:Left OSM and right Google Maps (b) Bottom:Left OSM and right Google Maps.

Full Reference:
Mahabir, R., Croitoru, A., Crooks, A.T., Agouris, P. and Stefanidis, A. (2018), A Critical Review of High and Very High Resolution Remote Sensing  Approaches for Detecting and Mapping Slums: Trends, Challenges and Emerging Opportunities, Urban Science. 2(1), 8; doi:10.3390/urbansci2010008 (pdf)
As always, any thoughts or comments are most welcome.

Friday, January 20, 2017

Authoritative and VGI in a Developing Country: A Comparative Case Study of Road Datasets in Nairobi


The motivation behind the paper was that while there are numerous studies comparing VGI to authoritative data in the developed world, there are very few that do so in developing world. In order to address this issue in the paper we compare the quality of authoritative road data (i.e. from the Regional Center for Mapping of Resources for Development - RCMRD) and non-authoritative crowdsourced road data (i.e. from OpenStreetMap (OSM) and Google’s Map Maker) in conjunction with population data in and around Nairobi, Kenya.

Results from our analysis show variability in coverage between all these datasets. RCMRD provided the most complete, albeit less current, coverage when taking into account the entire study area, while OSM and Map Maker showed a degradation of coverage as one moves from central Nairobi towards more rural areas. Further information including the abstract to our paper, some figures and full reference is given below.

Abstract:
With volunteered geographic information (VGI) platforms such as OpenStreetMap (OSM) becoming increasingly popular, we are faced with the challenge of assessing the quality of their content, in order to better understand its place relative to the authoritative content of more traditional sources. Until now, studies have focused primarily on developed countries, showing that VGI content can match or even surpass the quality of authoritative sources, with very few studies in developing countries. In this paper we compare the quality of authoritative (data from the Regional Center for Mapping of Resources for Development - RCMRD) and non-authoritative (data from OSM and Google’s Map Maker) road data in conjunction with population data in and around Nairobi, Kenya. Results show variability in coverage between all these datasets. RCMRD provided the most complete, albeit less current, coverage when taking into account the entire study area, while OSM and Map Maker showed a degradation of coverage as one moves from central Nairobi towards rural areas. Furthermore, OSM had higher content density in large slums, surpassing the authoritative datasets at these locations, while Map Maker showed better coverage in rural housing areas. These results suggest a greater need for a more inclusive approach using VGI to supplement gaps in authoritative data in developing nations.

Keywords: Volunteered Geographic Information; Crowdsourcing; Road Networks; Population Data; Kenya  
Road Coverage per km2
Pairwise difference in road coverage. Clockwise from top left: i) RCMRD 2011 versus Map Maker 2014; ii) RCMRD 2011 versus OSM 2011; iii) RCMRD 2011 versus OSM 2014; iv) OSM 2014 versus Map Maker 2014 (Red cells: first layer has higher coverage; Green cells: second layer has higher coverage).

Full Reference:
Mahabir, R., Stefanidis, A., Croitoru, A., Crooks, A.T. and Agouris, P. (2017), “Authoritative and Volunteered Geographical Information in a Developing Country: A Comparative Case Study of Road Datasets in Nairobi, Kenya”, ISPRS International Journal of Geo-Information, 6(1): 24, doi:10.3390/ijgi6010024.
As always any thoughts or comments about this work are welcome.

Thursday, January 08, 2015

Crowdsourcing Urban Form and Function

We have just had published a new paper entitled: "Crowdsourcing Urban Form and Function" in International Journal of Geographical Information Science which showcases some of our recent work with respect to cities and how new sources of information can be used to study urban morphology at a variety of spatial and temporal scales. Below is the abstract for the paper: 

"Urban form and function have been studied extensively in urban planning and geographic information science. However, gaining a greater understanding of how they merge to define the urban morphology remains a substantial scientific challenge. Towards this goal, this paper addresses the opportunities presented by the emergence of crowdsourced data to gain novel insights into form and function in urban spaces. We are focusing in particular on information harvested from social media and other open-source and volunteered datasets (e.g. trajectory and OpenStreetMap data). These data provide a first-hand account of form and function from the people who define urban space through their activities. This novel bottom-up approach to study these concepts complements traditional urban studies work to provide a new lens for studying urban activity. By synthesizing recent advancements in the analysis of open-source data we provide a new typology for characterizing the role of crowdsourcing in the study of urban morphology. We illustrate this new perspective by showing how social media, trajectory, and traffic data can be analyzed to capture the evolving nature of a city’s form and function. While these crowd contributions may be explicit or implicit in nature, they are giving rise to an emerging research agenda for monitoring, analyzing and modeling form and function for urban design and analysis."
This paper builds and extends considerably our prior work, with respect to crowdsourcing, volunteered and ambient geographic information. In the scope of this paper we use the term ‘urban form’ to refer to the aggregate of the physical shape of the city, its buildings, streets, and all other elements that make up the urban space. In essence, the geometry of the city. In contrast, we use the term ‘urban function’ to refer to the activities that are taking place within this space. To this end we contrast how crowdsourced data can related to more traditional sources of such information both explicitly and implicitly as shown in the table below. 

A typology of implicit and explicit form and function content

In addition, we also discuss in the paper how these new sources of data, which are often at finer resolutions than more authoritative data are allowing us to to customize the we we aggregate the data  at various geographical levels as shown below. Such aggregations can range from building footprints and addresses to street blocks (e.g. for density analysis), or street networks (e.g. for accessibility analysis). For large-scale urban analysis we can revert to the use of zonal geographies or grid systems.  
Aggregation methods for varied scales of built environment analysis

In the application section of the paper we highlight how we can extract implicit form and function from crowdsourced data. The image below for example, shows how we can take information from Twitter, and differentiate different neighborhoods over space and time.

Neighborhood map and topic modeling results showing the mixture of social functions in each area.

Finally in the paper, we outline an emerging research agenda related to the "persistent urban morphology concept" as shown below. Specifically how crowdsourcing is changing how we collect, analyze and model urban morphology. Moreover, how this new paradigm provides a new lens for studying the conceptualization of how cities operate, at much finer temporal, spatial, and social scales than we had been able to study so far.

The persistent urban morphology concept.

We hope you enjoy the paper.

Full Reference:  
Crooks, A.T., Pfoser, D., Jenkins, A., Croitoru, A., Stefanidis, A., Smith, D. A., Karagiorgou, S., Efentakis, A. and Lamprianidis, G. (2015), Crowdsourcing Urban Form and Function, International Journal of Geographical Information Science. DOI: 10.1080/13658816.2014.977905 (pdf)
 

Wednesday, July 09, 2014

New Paper: Assessing the impact of demographic characteristics on spatial error in VGI features

LISA analysis of positional accuracy for the OSM  data set
Building upon our interest in volunteered geographic information (VGI) and extending our previous paper  "Assessing Completeness and Spatial Error of Features in Volunteered Geographic Information" we have just published the paper with the rather long title "Assessing the impact of demographic characteristics on spatial error in volunteered geographic information features" where we explore how demographics impact on the quality of VGI data

Below is the abstract of the paper: 
The proliferation of volunteered geographic information (VGI), such as OpenStreetMap (OSM) enabled by technological advancements, has led to large volumes of user-generated geographical content. While this data is becoming widely used, the understanding of the quality characteristics of such data is still largely unexplored. An open research question is the relationship between demographic indicators and VGI quality. While earlier studies have suggested a potential relationship between VGI quality and population density or socio-economic characteristics of an area, such relationships have not been rigorously explored, and mainly remained qualitative in nature. This paper addresses this gap by quantifying the relationship between demographic properties of a given area and the quality of VGI contributions. We study specifically the demographic characteristics of the mapped area and its relation to two dimensions of spatial data quality, namely positional accuracy and completeness of the corresponding VGI contributions with respect to OSM using the Denver (Colorado, US) area as a case study. We use non-spatial and spatial analysis techniques to identify potential associations among demographics data and the distribution of positional and completeness errors found within VGI data. Generally, the results of our study show a lack of statistically significant support for the assumption that demographic properties affect the positional accuracy or completeness of VGI. While this research is focused on a specific area, our results showcase the complex nature of the relationship between VGI quality and demographics, and highlights the need for a better understanding of it. By doing so, we add to the debate of how demographics impact on the quality of VGI data and lays the foundation to further work.

The analysis workflow
Full Reference:
Mullen W., Jackson, S. P., Croitoru, A., Crooks, A. T., Stefanidis, A. and Agouris, P., (2014), Assessing the Impact of Demographic Characteristics on Spatial Error in Volunteered Geographic Information Features, GeoJournal. DOI: 10.1007/s10708-014-9564-8

Thursday, June 05, 2014

The Evolving GeoWeb

We recently contributed a chapter to Geocomputation (2nd edition) entitled "The Evolving GeoWeb". What is interesting is the marked difference between the first edition (which was published in 2000) and the second. For example, in the latest edition, there is a chapter on agent-based modeling (ABM), while in the first, only cellular automata (CA) models were covered and ABMs only briefly discussed. We also see in the second edition new chapters including ours on the GeoWeb which shows how the field of geocomputation has changed with advances in Web 2.0 technology, greater computational power, new devices (such as GPS enabled smart phones) and the rise in new sources of data (volunteered and ambient geographical information, VGI and AGI). The abstract of our chapter is copied below, while examples of early and current web mapping is provided in the figures below.

"The Internet and its World Wide Web (WWW) have revolutionised many aspects of our daily lives from how we access and retrieve information to how we communicate with friends and peers. Over the past two decades, the Web has evolved from a system aimed primarily towards data access to a medium that fosters information contribution and interaction within large, globally distributed communities. Just as the Web evolved, so too did Web-based GeoComputation (GC), which we refer to here as the Geographic World Wide Web or the GeoWeb for short. Whereas the generation and viewing of geographical information was initially limited to the purview of specialists and dedicated workstations, it has now become of interest to the general public and is accessed using a variety of devices such as GPS-enabled smartphones and tablets. Accordingly, in order to meet the needs of this expanded constituency, the GeoWeb has evolved from displaying static maps to a dynamic environment where diverse datasets can be accessed, exchanged and mashed together. Within this chapter, we trace this evolution and corresponding paradigm shifts within the GeoWeb with a particular focus on Web 2.0 technologies. Furthermore, we explore the role of the crowd in consuming and producing geographical information and how this is influencing GeoWeb developments. Specifically, we are interested in how location provides a means to index and access information over the Internet. Next, we discuss the role of Digital Earth and virtual world paradigms for storing, manipulating and displaying geographical information in an immersive environment. We then discuss how GIS software is changing towards GIS services and the rise in location-based services (LBS) and lightweight software applications (so-called apps). Finally, we conclude with a summary of this chapter and discuss how the GeoWeb might evolve with the rise in massive amounts of locational data being generated."

PARC Map Viewer (Source: Putz, 1994)

Google Earth as a base layer for possible trajectories of the radioactive plume from the Fukushima Daiichi nuclear disaster. The different color lines represent different possible paths of the plume (Source: http://forecast.chapman.edu/images/japan/tranj.kmz).

A proof of our chapter can be downloaded from here. We hope you enjoy it!

Full reference:
Crooks, A.T., Hudson-Smith, A., Croitoru, A. and Stefanidis, A. (2014), The Evolving GeoWeb, in Abrahart R. J. and See, L. M. (eds.), Geocomputation (Second Edition), CRC Press, Boca Raton, FL, pp. 69-96. (pdf)

Tuesday, June 04, 2013

Completeness and Spatial Error of Features in VGI

I have had an interest in volunteered geographic information (VGI) for quite some time (see my publications or blog posts) but only recently have I had an opportunity to look at the spatial error of features within VGI. To this end, our paper entitled "Assessing Completeness and Spatial Error of Features in Volunteered Geographic Information" has just been published in ISPRS International Journal of Geo-Information. Below is the abstract of the paper along with some figures. Further details about the paper can be seen at the bottom of the page.
The assessment of the quality and accuracy of Volunteered Geographic Information (VGI) contributions, and by extension the ultimate utility of VGI data has fostered much debate within the geographic community. The limited research to date has been focused on VGI data of linear features and has shown that the error in the data is heterogeneously distributed. Some have argued that data produced by numerous contributors will produce a more accurate product than an individual and some research on crowd-sourced initiatives has shown that to be true, although research on VGI is more infrequent. This paper proposes a method for quantifying the completeness and accuracy of a select subset of infrastructure-associated point datasets of volunteered geographic data within a major metropolitan area using a national geospatial dataset as the reference benchmark with two datasets from volunteers used as test datasets. The results of this study illustrate the benefits of including quality control in the collection process for volunteered data. 

Keywords: volunteered geographic information (VGI); OpenStreetMap; quality; error; point.
Comparison of OSM, OSMCP, and ORNL data.
Various identified locations of Southwest Early College
Full reference:
Jackson, S. P., Mullen W., Agouris, P., Crooks, A., Croitoru, A. and Stefanidis, A. (2013), Assessing Completeness and Spatial Error of Features in Volunteered Geographic Information, ISPRS International Journal of Geo-Information, 2 (2): 507-530. Download from here.

Tuesday, December 06, 2011

Harvesting ambient geospatial information from social media feeds

A paper I  recently co-authored with Anthony Stefanidis and Jacek Radzikowski from George Mason University entitled "Harvesting ambient geospatial information from social media feeds" is now available in  GeoJournal. 
 
The abstract for the paper reads as follows: "Social media generated from many individuals is playing a greater role in our daily lives and provides a unique opportunity to gain valuable insight on information flow and social networking within a society. Through data collection and analysis of its content, it supports a greater mapping and understanding of the evolving human landscape. The information disseminated through such media represents a deviation from volunteered geography, in the sense that it is not geographic information per se. Nevertheless, the message often has geographic footprints, for example, in the form of locations from where the tweets originate, or references in their content to geographic entities. We argue that such data conveys ambient geospatial information, capturing for example, people’s references to locations that represent momentary social hotspots. In this paper we address a framework to harvest such ambient geospatial information, and resulting hybrid capabilities to analyze it to support situational awareness as it relates to human activities. We argue that this emergence of ambient geospatial analysis represents a second step in the evolution of geospatial data availability, following on the heels of volunteered geographical information."

Geolocating pairs of tweeters and retweeters