Showing posts with label Data Mining. Show all posts
Showing posts with label Data Mining. Show all posts

Thursday, July 14, 2022

Drone strikes and radicalization

In the past we had posted on models of radicalization, but such models were rather abstract.  Building on this previous work Brandon Shapiro and myself have a new paper entitled "Drone Strikes and Radicalization: An Exploration Utilizing Agent-Based Modeling and Data Applied to Pakistan" which has recently been published in Computational and Mathematical Organization Theory journal. In the paper we develop and present an agent-based model informed by theory and calibrated using empirical data to explore the relationship between kinetic actions (i.e., drone strikes) and terrorist attacks in Pakistan from 2004 through 2018. 

The data itself came from the Bureau of Investigative Journalism data as our source for Pakistan drone strikes (i.e., kinetic actions) and the National Consortium for the Study of Terrorism and Responses to Terrorism( START)  Global Terrorism Database (GTD) as our source for terrorist incidents. Rather than try to pinpoint and define the motivating factors which might influence somebody down a path toward radicalization, our model that incorporated a distributed lag model to characterize the inter-dependencies between drone strikes and terrorist attacks observed in Pakistan. Based on parametric and validation tests, the model simulates a terrorist attack curve which approximates the rate and magnitude observed in Pakistan from 2007 through 2018. 

If this sounds of interest, below we provide the abstract to the paper, along with some images of model graphical user interface, the model logic and some of the results. The model itself was created in NetLogo and is available at: https://www.comses.net/codebase-release/30540ae3-486b-44e4-8ff0-785575433af0/  (along with the data and detailed ODD of the model). At the bottom of the page you can find the full citation and a link to the paper.

Abstract:

The employment of drone strikes has been ongoing and the public continues to debate their perceived benefits. A question that persists is whether drone strikes contribute to an increase in radicalization. This paper presents a data-driven approach to explore the relationship between drone strikes conducted in Pakistan and subsequent responses, often in the form of terrorist attacks carried out by those in the communities targeted by these particular counter terrorism measures. Our exploration and analysis of news reports which discussed drone strikes and radicalization suggest that government-sanctioned drone strikes in Pakistan appear to drive terrorist events with a distributed lag that can be determined analytically. We leverage news reports to inform and calibrate an agent-based model grounded in radicalization and opinion dynamics theory. This enabled us to simulate terrorist attacks that approximated the rate and magnitude observed in Pakistan from 2007 through 2018. We argue that this research effort advances the field of radicalization and lays the foundation for further work in the area of data-driven modeling and drone strikes.  
Keywords: Radicalization, Data-driven modeling, Drone strikes, Terrorism, Pakistan , Agent-based modeling.
Pakistan radicalization model’s graphical user interface. From left to right: model input param- eters, the agents’ social network and resulting model outputs

The agent-based model flow diagram.

Terrorist attacks simulated by Pakistan radicalization model qualitatively agree with real-world system.

Full Reference: 

Shapiro, B. and Crooks, A.T. (2022) Drone Strikes and Radicalization: An Exploration Utilizing Agent-Based Modeling and Data Applied to Pakistan, Computational and Mathematical Organization Theory. Available at https://doi.org/10.1007/s10588-022-09364-1. (pdf)


Tuesday, November 09, 2021

GIS and ABM: Past, Present and Future

The other day, Alison Heppenstall and myself were invited to give a keynote at the 2021 International Conference on Geospatial Information Sciences. Its not hard to guess what we chose to be the title of our talk: "GIS and Agent-Based-Modelling: Past, Present and Future."

The talk was a synthesis of two publications (Crooks et al. 2019 and Heppenstall et al. 2021) along with some things we are currently working on. For those who are interested the conference has released not only our talk but also the other keynotes.


Tuesday, July 13, 2021

Kinetic Action and Radicalization

In the past we had posted on models of radicalization, but such models were rather abstract.  However in a recent paper entitled "Kinetic Action and Radicalization: A Case Study of Pakistan" which was presented at the  International Conference on Social Computing, Behavioral-Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation (or SBP-BRiMS for short) we take such work a step further. 

In the paper, Brandon Shapiro and myself develop and present a simple agent-based model informed by theory and calibrated using empirical data to explore the relationship between kinetic actions (i.e., drone strikes) and terrorist attacks in Pakistan from 2004 through 2018. The data itself came from the Bureau of Investigative Journalism data as our source for Pakistan drone strikes (i.e., kinetic actions) and the National Consortium for the Study of Terrorism and Responses to Terrorism( START)  Global Terrorism Database (GTD) as our source for terrorist incidents. 

Rather than try to pinpoint and define the motivating factors which might influence somebody down a path toward radicalization, our model that incorporated a distributed lag model to characterize the inter-dependencies between drone strikes and terrorist attacks observed in Pakistan. Based on parametric and validation tests, the model simulates a terrorist attack curve which approximates the rate and magnitude observed in Pakistan from 2007 through 2018. If this sounds of interest, below we provide the abstract to the paper, along with some images of model graphical user interface, the model logic and some of the results. The model itself was created in NetLogo and is available at: https://bit.ly/3qLJynv  (along with the data and detailed ODD of the model). At the bottom of the page you can find the full citation and a link to the paper.

Abstract. Drone strikes have been ongoing and there is a debate about their benefits. One major question is what is their role with respect to radicalization. This paper presents a data-driven approach to explore the relationship between drone strikes in Pakistan and subsequent responses, often in the form of terrorist attacks carried out by those in the communities targeted by these counter-terrorism measures. Our analysis of news reports which dis-cussed drone strikes and radicalization suggests that government-sanctioned drone strikes in Pakistan appear to drive terrorist events with a distributed lag that can be determined analytically. We then utilize these news reports to inform and calibrate an agent-based model which is ground-ed in radicalization and opinion dynamics theory. In doing so, we were able to simulate terrorist attacks that approximated the rate and magnitude ob-served in Pakistan from 2007 through 2018. We argue that this research effort advances the field of radicalization and lays the foundation for further work in the area of data-driven modeling and kinetic actions.

Keywords: Radicalization, Data-driven modeling, Drone strikes, Terrorism, Pakistan, Agent-based modeling.

Radicalization model’s graphical user interface.

The agent-based model flow diagram.

Terrorist attacks simulated by radicalization model qualitatively agree with real-world system.

Full Reference:
Shapiro, B. and Crooks, A.T. (2021), Kinetic Action and Radicalization: A Case Study of Pakistan, in Thomson, R., Hussain, M.N., Dancy, C.L. and Pyke, A. (eds), Proceedings of 2021 International Conference on Social Computing, Behavioral-Cultural Modeling & Prediction and Behavior Representation in Modeling and Simulation, Washington DC., pp 321-330. (pdf)

Wednesday, January 06, 2021

Elections and Bots

Continuing our work on botsRoss Schuchard and myself have a new paper in PLOS ONE entitled "Insights into elections: An ensemble bot detection coverage framework applied to the 2018 U.S. midterm elections." Our motivation for the work came from the fact that during elections internet-based technological platforms (e.g., online social networks (OSNs), online political blogs etc.) are gaining more power compared to mainstream media sources (e.g., print, television and radio). While such technologies are reducing the barrier for individuals to actively participate in political dialogue, the relatively unsupervised nature of OSNs increases susceptibility to misinformation campaigns, especially with respect to political and election dialogue. This is especially the case for social bots—automated software agents designed to mimic or impersonate humans which are prevalent actors in OSN platforms and have proven to amplify misinformation.  

The issue however is that no single detection algorithm is able to account for the myriad of social bots operating in OSNs. To overcome this issue, this research incorporates multiple social bot detection services to determine the prevalence and relative importance of social bots within an OSN conversation of tweets. Through the lens of the 2018 U.S. midterm elections, 43.5 million tweets were harvested capturing the election conversation which were then analyzed for evidence of bots using three bot detection platform services: Botometer, DeBot and Bot-hunter.

We found that bot and human accounts contributed temporally to our tweet election corpus at relatively similar cumulative rates. The multi-detection platform comparative analysis of intra-group and cross-group interactions showed that bots detected by DeBot and Bot-hunter persistently engaged humans at rates much higher than bots detected by Botometer. Furthermore, while bots accounted for less than 8% of all unique accounts in the election conversation retweet network, bots accounted for more than 20% of the top-100 and top-25 ranking out-degree centrality, thus suggesting persistent activity to engage with human accounts. Finally, the bot coverage overlap analysis shows that minimal overlap existed among the bots detected by the three bot detection platforms, with only eight total bot accounts detected by all (out of a total of 254,492 unique bots in the overall tweet corpus ).

If this research sounds interesting to you, below we provide the abstract to the paper along with some figures outlining our methodology and some of the results. While at the bottom of the post you can see the full reference and there is a link to the paper were you can read more.

Abstract:

The participation of automated software agents known as social bots within online social network (OSN) engagements continues to grow at an immense pace. Choruses of concern speculate as to the impact social bots have within online communications as evidence shows that an increasing number of individuals are turning to OSNs as a primary source for information. This automated interaction proliferation within OSNs has led to the emergence of social bot detection efforts to better understand the extent and behavior of social bots. While rapidly evolving and continually improving, current social bot detection efforts are quite varied in their design and performance characteristics. Therefore, social bot research efforts that rely upon only a single bot detection source will produce very limited results. Our study expands beyond the limitation of current social bot detection research by introducing an ensemble bot detection coverage framework that harnesses the power of multiple detection sources to detect a wider variety of bots within a given OSN corpus of Twitter data. To test this framework, we focused on identifying social bot activity within OSN interactions taking place on Twitter related to the 2018 U.S. Midterm Election by using three available bot detection sources. This approach clearly showed that minimal overlap existed between the bot accounts detected within the same tweet corpus. Our findings suggest that social bot research efforts must incorporate multiple detection sources to account for the variety of social bots operating in OSNs, while incorporating improved or new detection methods to keep pace with the constant evolution of bot complexity.

 

Fig 1. Social bot analysis framework employing multiple bot detection platforms. The framework enables the application of ensemble analysis methods to determine the prevalence and relative importance of social bots within Twitter conversations discussing the 2018 U.S. midterm elections.
 

Fig 3. Cumulative tweet contribution rates for the 2018 U.S. midterm OSN conversation (October 10 – November 6, 2018) from the (a) human (blue) / bot (red) and (b) DeBot (green) / Botometer (pink) / Bot-hunter (orange) account classification perspectives.

Fig 4. Intra-group and cross-group retweet communication patterns of human (blue) and social bot (red) users within the 2018 U.S. midterm election Twitter conversation according to each bot detection classification platform: (a) Combined Bot Sources (b) DeBot (c) Botometer (d) Bot-hunter. The combined bot sources results (shown in gray) classified an account as a bot in aggregate fashion if any of the three detection platforms classified the account as a bot.

Fig 5. Social bot account evidence within the top-N (where, N = 1000 / 500 / 100 / 25) centrality rankings [(a) eigenvector (b) in-degree (c) out-degree (d) PageRank] according to bot classification results from Bot-hunter (orange), Botometer (pink) and DeBot (green).
 
Fig 7. Bot detection coverage analysis for bots detected within the 2018 U.S. midterm election Twitter conversation using the Botometer, Bot-hunter and DeBot bot detection platforms.

 

Full reference:

Schuchard, R.J. and Crooks, A.T. (2021), Insights into Elections: An Ensemble Bot Detection Coverage Framework Applied to the 2018 U.S. Midterm Elections, PLoS ONE, 16(1): e0244309. Available at  https://doi.org/10.1371/journal.pone.0244309. (pdf).

Tuesday, May 26, 2020

Crowdsourcing Street View Imagery: A Comparison of Mapillary and OpenStreetCam


In the past we have written extensively on Volunteered Geographic Information (VGI) such as OpenStreetMap or Twitter. However, we have not really explored Street View Imagery  (SVI), well not until now. Within the realm of VGI, SVI has emerged in recent years as a novel and rich source of data on cities from which geographic information can be derived.

Perhaps the most well-known example of SVI utilization is that of Google Street View (GSV). While SVI has been traditionally collected by governmental agencies and companies alike, we are now also witnessing the emergence of Volunteered Street View Imagery (VSVI), which relies on a crowdsourced effort to provide geotagged street-level imagery coverage of traversable pathways (e.g., a street or trail). Such imagery, similar to GSV, provides detailed information about the location of objects such as cars, road markings, traffic lights and signs, and allows for the automatic extraction of features at scale. Such imagery can also be mined using machine learning algorithms to automatically derive points of interest (POI) databases (e.g., locations of coffee shops and fire hydrants) without the intervention of the citizen.

To explore VSVI we have just published a new paper entitled: "Crowdsourcing Street View Imagery: A Comparison of Mapillary and OpenStreetCam" in the ISPRS International Journal of Geo-Information. In this paper we examine VSVI data collected from two different platforms: Mapillary and OpenStreetCam (OSC) for four metropoiltan areas in the United States (i.e., Washington (District of Columbia), San Francisco (California), Phoenix (Arizona), and Detroit (Michigan)). Both of these online platforms accept sequences of images captured from mobile devices and uploaded via an app on the device (like those shown in the image to the right). Images are geolocated using the device’s global positioning system (GPS). More specifically the paper examines:
  • the level of spatial coverage of each platform in order to assess the overall potential of such platforms to provide adequate coverage of geographic information.
  • user contribution patterns in Mapillary and OSC in order to understand how users are contributing to these platforms.
Results from our systematic and quantitative analysis of these two emerging VGI sources indicate that most Mapillary and OSC contributions occurred along control-access highways and local roads, and that the overall coverage in these sources is variable in comparison to an authoritative source (i.e., TIGER). Furthermore, our results showed that while the number of contributors varied across sites, only a few contributors were responsible for producing most of the raw data. User contribution patterns were also different in Mapillary and OSC. Specifically, we found that while patterns in coverage were variable for the different OSC sites, coverage patterns in Mapillary tended to be similar among sites. This finding may be linked to several factors, including differences in mapping practice, or issues with participation inequality, a topic that has been highly researched for other VGI platforms such as OSM, but which is still lacking within VSVI. Lastly, user contributions in Mapillary tended to be higher around 8:00 am, 1:00 pm and 5:00 pm (local time). This finding suggests that VSVI contributions tend to coincide with the morning and afternoon commute, and the lunch hour of the contributors.

If you wish to find out more about this work below we provide the abstract to the paper, a visual flowchart of our workflow and some of our our results. The full reference and link to the paper is provided at the bottom of the post.

Abstract:
Over the last decade, Volunteered Geographic Information (VGI) has emerged as a viable source of information on cities. During this time, the nature of VGI has been evolving, with new types and sources of data continually being added. In light of this trend, this paper explores one such type of VGI data: Volunteered Street View Imagery (VSVI). Two VSVI sources, Mapillary and OpenStreetCam, were extracted and analyzed to study road coverage and contribution patterns for four US metropolitan areas. Results show that coverage patterns vary across sites, with most contributions occurring along local roads and in populated areas. We also found that a few users contributed most of the data. Moreover, the results suggest that most data are being collected during three distinct times of day (i.e., morning, lunch and late afternoon). The paper concludes with a discussion that while VSVI data is still relatively new, it has the potential to be a rich source of spatial and temporal information for monitoring cities.

Keywords: Crowdsourcing; Volunteered Geographic Information; Street View Imagery; Mapillary, OpenStreetCam
Overview of methodology

Spatial distribution of road networks.
Spatial comparison of roads in kilometers.


Full Reference: 
Mahabir, R., Schuchard, R., Crooks, A.T., Croitoru, A. and Stefanidis, A. (2020), Crowdsourcing Street View Imagery: A Comparison of Mapillary and OpenStreetCam, ISPRS International Journal of Geo-Information. 9(6), 341; https://doi.org/10.3390/ijgi9060341 (pdf)

Wednesday, December 05, 2018

Detecting and Mapping Slums using Open Data

Urban and slum areas in Nairobi (False composite image created
by stacking image bands 7, 6 and 4 from the Landsat 8 satellite.
Turning back to slums, we just published paper entitled "Detecting and Mapping Slums using Open Data: A Case Study in Kenya" in the International Journal of Digital Earth. This work builds and extends our previous research on using new sources of data to explore the slum settlements in 3 cities in Kenya (i.e. Nairobi, Mombasa and Kisumu). Specifically, we examine how the fusion of Volunteered Geographical Information, Social Media, and other open data sources can complement remote sensing imagery in supporting slum detection, mapping and monitoring. 

We do this by using data mining tools (e.g. logistic regression, discriminant analysis and the See5 decision tree), to develop context-sensitive definitions for slums based on location, as well as for testing the generalizability of indicators and derived slum models. The end result is an indicator database for slums using open sources of physical and socio-economic data that can be used to characterize slum settlements. If you wish to know more, below we provide the abstract to the paper along with some of the figures and the full citation with a link to the paper itself.

Abstract:
The worldwide slum population currently stands at over one billion, with substantial growth expected in the coming decades. Traditionally, slums have been mapped using information derived mainly from either physical indicators using remote sensing data, or socio-economic indicators using census data. Each data source on its own provides only a partial view of slums, an issue further compounded by data poverty in less developed countries. To overcome such issues, this paper explores the fusion of traditional with emerging open data sources and data mining tools to identify additional indicators that can be used to detect and map the presence of slums, map their footprint, and map their evolution. Towards this goal, we develop an indicator database for slums using open sources of physical and socio-economic data that can be used to characterize slum settlements. Using this database, we then leverage data mining techniques to identify the most suitable combination of these indicators for mapping slums. Using three cities in Kenya as test cases, results show that the fusion of these data can improve the mapping accuracy of slums. These results suggest that the proposed approach can provide a viable solution to the emerging challenge of monitoring the growth of slums.
Keywords: Slums; Remote Sensing; Socio-economic; Urban sustainability; Data mining; Kenya

Study areas in Kenya

Methodology workflow

Distribution of positive classified cases for slums for (a) logistic regression, (b) discriminant analysis and (c) the See5 decision tree.
Full Reference:
Mahabir, R., Agouris, P., Stefanidis, A., Croitoru, A. and Crooks, A.T. (2018), Detecting and Mapping Slums using Open Data: A Case Study in Kenya, International Journal of Digital Earth. DOI: https://doi.org/10.1080/17538947.2018.1554010. (pdf)