Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Wednesday, January 06, 2021

Elections and Bots

Continuing our work on botsRoss Schuchard and myself have a new paper in PLOS ONE entitled "Insights into elections: An ensemble bot detection coverage framework applied to the 2018 U.S. midterm elections." Our motivation for the work came from the fact that during elections internet-based technological platforms (e.g., online social networks (OSNs), online political blogs etc.) are gaining more power compared to mainstream media sources (e.g., print, television and radio). While such technologies are reducing the barrier for individuals to actively participate in political dialogue, the relatively unsupervised nature of OSNs increases susceptibility to misinformation campaigns, especially with respect to political and election dialogue. This is especially the case for social bots—automated software agents designed to mimic or impersonate humans which are prevalent actors in OSN platforms and have proven to amplify misinformation.  

The issue however is that no single detection algorithm is able to account for the myriad of social bots operating in OSNs. To overcome this issue, this research incorporates multiple social bot detection services to determine the prevalence and relative importance of social bots within an OSN conversation of tweets. Through the lens of the 2018 U.S. midterm elections, 43.5 million tweets were harvested capturing the election conversation which were then analyzed for evidence of bots using three bot detection platform services: Botometer, DeBot and Bot-hunter.

We found that bot and human accounts contributed temporally to our tweet election corpus at relatively similar cumulative rates. The multi-detection platform comparative analysis of intra-group and cross-group interactions showed that bots detected by DeBot and Bot-hunter persistently engaged humans at rates much higher than bots detected by Botometer. Furthermore, while bots accounted for less than 8% of all unique accounts in the election conversation retweet network, bots accounted for more than 20% of the top-100 and top-25 ranking out-degree centrality, thus suggesting persistent activity to engage with human accounts. Finally, the bot coverage overlap analysis shows that minimal overlap existed among the bots detected by the three bot detection platforms, with only eight total bot accounts detected by all (out of a total of 254,492 unique bots in the overall tweet corpus ).

If this research sounds interesting to you, below we provide the abstract to the paper along with some figures outlining our methodology and some of the results. While at the bottom of the post you can see the full reference and there is a link to the paper were you can read more.

Abstract:

The participation of automated software agents known as social bots within online social network (OSN) engagements continues to grow at an immense pace. Choruses of concern speculate as to the impact social bots have within online communications as evidence shows that an increasing number of individuals are turning to OSNs as a primary source for information. This automated interaction proliferation within OSNs has led to the emergence of social bot detection efforts to better understand the extent and behavior of social bots. While rapidly evolving and continually improving, current social bot detection efforts are quite varied in their design and performance characteristics. Therefore, social bot research efforts that rely upon only a single bot detection source will produce very limited results. Our study expands beyond the limitation of current social bot detection research by introducing an ensemble bot detection coverage framework that harnesses the power of multiple detection sources to detect a wider variety of bots within a given OSN corpus of Twitter data. To test this framework, we focused on identifying social bot activity within OSN interactions taking place on Twitter related to the 2018 U.S. Midterm Election by using three available bot detection sources. This approach clearly showed that minimal overlap existed between the bot accounts detected within the same tweet corpus. Our findings suggest that social bot research efforts must incorporate multiple detection sources to account for the variety of social bots operating in OSNs, while incorporating improved or new detection methods to keep pace with the constant evolution of bot complexity.

 

Fig 1. Social bot analysis framework employing multiple bot detection platforms. The framework enables the application of ensemble analysis methods to determine the prevalence and relative importance of social bots within Twitter conversations discussing the 2018 U.S. midterm elections.
 

Fig 3. Cumulative tweet contribution rates for the 2018 U.S. midterm OSN conversation (October 10 – November 6, 2018) from the (a) human (blue) / bot (red) and (b) DeBot (green) / Botometer (pink) / Bot-hunter (orange) account classification perspectives.

Fig 4. Intra-group and cross-group retweet communication patterns of human (blue) and social bot (red) users within the 2018 U.S. midterm election Twitter conversation according to each bot detection classification platform: (a) Combined Bot Sources (b) DeBot (c) Botometer (d) Bot-hunter. The combined bot sources results (shown in gray) classified an account as a bot in aggregate fashion if any of the three detection platforms classified the account as a bot.

Fig 5. Social bot account evidence within the top-N (where, N = 1000 / 500 / 100 / 25) centrality rankings [(a) eigenvector (b) in-degree (c) out-degree (d) PageRank] according to bot classification results from Bot-hunter (orange), Botometer (pink) and DeBot (green).
 
Fig 7. Bot detection coverage analysis for bots detected within the 2018 U.S. midterm election Twitter conversation using the Botometer, Bot-hunter and DeBot bot detection platforms.

 

Full reference:

Schuchard, R.J. and Crooks, A.T. (2021), Insights into Elections: An Ensemble Bot Detection Coverage Framework Applied to the 2018 U.S. Midterm Elections, PLoS ONE, 16(1): e0244309. Available at  https://doi.org/10.1371/journal.pone.0244309. (pdf).

Friday, December 04, 2020

Future Developments in Geographical Agent-Based Models: Challenges and Opportunities

Its been a while since (to say the least), that we wrote a position paper about agent-based modeling. But with agent-based modeling becoming more widely accepted  and the growth of machine learning within the geographical sciences we thought we would revisit some of the existing challenges  (e.g. validation, representing behavior) and discuss how machine learning and data might help here. To this end, Alison HeppenstallNick Malleson, Ed Manley, Jiaqi Ge and Mike Batty, have recently published a paper entitled "Future Developments in Geographical Agent-Based Models: Challenges and Opportunities" in Geographical Analysis.  Below we provide the abstract to the paper, and if this is of interest please follow the links to the paper itself.

Abstract

Despite reaching a point of acceptance as a research tool across the geographical and social sciences, there remain significant methodological challenges for agent-based models. These include recognizing and simulating emergent phenomena, agent representation, construction of behavioral rules, calibration and validation. Whilst advances in individual-level data and computing power have opened up new research avenues, they have also brought with them a new set of challenges. This paper reviews some of the challenges that the field has faced, the opportunities available to advance the state-of-the-art, and the outlook for the field over the next decade. We argue that although agent-based models continue to have enormous promise as a means of developing dynamic spatial simulations, the field needs to fully embrace the potential offered by approaches from machine learning to allow us to fully broaden and deepen our understanding of geographical systems.

Full Reference:

Heppenstall, A., Crooks, A.T., Malleson, N., Manley, E., Ge, J. and Batty, M. (2021), Future Developments in Geographical Agent-Based Models: Challenges and Opportunities, Geographical Analysis. https://doi.org/10.1111/gean.12267 (pdf)

Thursday, October 15, 2020

The Impact of Mandatory Remote Work during the COVID-19 Pandemic

In the past we have written about using agent-based modeling to study human resources management issues and how workplace the layout might impact subordinates interactions with managers but with growing amounts data we can explore how employees communicate with each other. To this end, Talha Oz and myself have a  new paper entitled "Exploring the Impact of Mandatory Remote Workduring the COVID-19 Pandemic" which will be presented in a special session on COVID-19 at the 2020 International Conference on Social Computing, Behavioral-Cultural Modeling, & Prediction and Behavior Representation in Modeling and Simulation (or SBP-Brims 2020 for short). 

In this study we exploit metadata (and not content) emitted from commonplace workplace technologies such as calendar and workplace messaging apps collected from a tech company in order to see how mandatory remote work changed communication patterns and how such data can be used to measure organizational health. If this is of interest to you, below we provide the abstract to the paper along with some of the results with respect to how meetings and communication patterns changed from  business as usual (BAU), pre pandemic to that when people were forced to work from home (WFH). Finally at the bottom of the post we provide the full reference and the link to the paper.

Abstract. During the early months of the COVID-19 pandemic, millions of people had to work from home. We examine the ways in which COVID-19 affect organizational communication by analyzing five months of calendar and messaging metadata from a technology company. We found that: (i) cross-level communication increased more than that of same-level, (ii) while within-team messaging increased considerably, meetings stayed the same, (iii) off-hours messaging became much more frequent, and that this effect was stronger for women; (iv) employees respond to non-managers faster than managers; finally, (v) the number of short meetings increased while long meetings decreased. These findings contribute to theories on organizational communication, remote work, management, and flexibility stigma. Besides, this study exemplifies a strategy to measure organizational health using an objective (not self-report based) method. To the best of our knowledge, this is the first study using workplace communication metadata to examine the heterogeneous effects of mandatory remote work. 

Keywords: Work from Home, Communication, COVID-19, Organization.




Full Reference:

Oz, T. and Crooks, A.T. (2020), Exploring the Impact of Mandatory Remote Work during the COVID-19 Pandemic, 2020 International Conference on Social Computing, Behavioral-Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation, Washington DC. (pdf)

 If you would like a pre-print of  paper, just let us know and we can email you one. 

Friday, April 26, 2019

Computational Social Science of Disasters: Opportunities and Challenges

Figure 1: Relation of computational social science of
disasters (CSSD) with other fields.
Past posts have discussed or demonstrated how  computational social science (CSS) (i.e. the study of social science through computational methods) can be utilized explore disasters or diseases but this has not really been  formalized.  To this end, Annetta Burger, Talha Oz, William Kennedy and myself have just had a paper published in Future Internet entitled "Computational Social Science of Disasters: Opportunities and Challenges". In the paper we introduce computational social science of disasters (CSSD). CSSD is defined as an approach to explain the social dynamics of disasters via computational means by adopting the relevant parts of CSS, social sciences in disaster, and crisis informatics as depicted in Figure 1. Specifically, we briefly review the domains and the approaches of each of the traditional social science disciplines to disasters (e.g. sociology, psychology, anthropology, political science, and economics). Next we describe the fields of CSS and crisis informatics before discussing the components of CSSD. We highlight some exemplar studies which capture certain elements of CSSD along with the challenges and opportunities it brings to the study of disasters. If you would like to find out more, below is the abstract to the paper along with the full reference and link to the paper.

Abstract
Disaster events and their economic impacts are trending, and climate projection studies suggest that the risks of disaster will continue to increase in the near future. Despite the broad and increasing social effects of these events, the empirical basis of disaster research is often weak, partially due to the natural paucity of observed data. At the same time, some of the early research regarding social responses to disasters have become outdated as social, cultural, and political norms have changed. The digital revolution, the open data trend, and the advancements in data science provide new opportunities for social science disaster research. We introduce the term computational social science of disasters (CSSD), which can be formally defined as the systematic study of the social behavioral dynamics of disasters utilizing computational methods. In this paper, we discuss and showcase the opportunities and the challenges in this new approach to disaster research. Following a brief review of the fields that relate to CSSD, namely traditional social sciences of disasters, computational social science, and crisis informatics, we examine how advances in Internet technologies offer a new lens through which to study disasters. By identifying gaps in the literature, we show how this new field could address ways to advance our understanding of the social and behavioral aspects of disasters in a digitally connected world. In doing so, our goal is to bridge the gap between data science and the social sciences of disasters in rapidly changing environments.

Keywords: Disasters; Computational Social Science; Crisis Informatics; Disaster Modeling, Web 2.0; Social Media; Big Data; Volunteered Geographical Information; Crowdsourcing.
Figure 2: Interactions of data analysis, computational models, and social theory
in computational social science of disasters.

Full Reference:
Burger, A., Oz, T., Kennedy, W.G. and Crooks, A.T. (2019), Computational Social Science of Disasters: Opportunities and Challenges, Future Internet, 11(5): 103; https://doi.org/10.3390/fi11050103. (pdf)

Wednesday, December 05, 2018

Detecting and Mapping Slums using Open Data

Urban and slum areas in Nairobi (False composite image created
by stacking image bands 7, 6 and 4 from the Landsat 8 satellite.
Turning back to slums, we just published paper entitled "Detecting and Mapping Slums using Open Data: A Case Study in Kenya" in the International Journal of Digital Earth. This work builds and extends our previous research on using new sources of data to explore the slum settlements in 3 cities in Kenya (i.e. Nairobi, Mombasa and Kisumu). Specifically, we examine how the fusion of Volunteered Geographical Information, Social Media, and other open data sources can complement remote sensing imagery in supporting slum detection, mapping and monitoring. 

We do this by using data mining tools (e.g. logistic regression, discriminant analysis and the See5 decision tree), to develop context-sensitive definitions for slums based on location, as well as for testing the generalizability of indicators and derived slum models. The end result is an indicator database for slums using open sources of physical and socio-economic data that can be used to characterize slum settlements. If you wish to know more, below we provide the abstract to the paper along with some of the figures and the full citation with a link to the paper itself.

Abstract:
The worldwide slum population currently stands at over one billion, with substantial growth expected in the coming decades. Traditionally, slums have been mapped using information derived mainly from either physical indicators using remote sensing data, or socio-economic indicators using census data. Each data source on its own provides only a partial view of slums, an issue further compounded by data poverty in less developed countries. To overcome such issues, this paper explores the fusion of traditional with emerging open data sources and data mining tools to identify additional indicators that can be used to detect and map the presence of slums, map their footprint, and map their evolution. Towards this goal, we develop an indicator database for slums using open sources of physical and socio-economic data that can be used to characterize slum settlements. Using this database, we then leverage data mining techniques to identify the most suitable combination of these indicators for mapping slums. Using three cities in Kenya as test cases, results show that the fusion of these data can improve the mapping accuracy of slums. These results suggest that the proposed approach can provide a viable solution to the emerging challenge of monitoring the growth of slums.
Keywords: Slums; Remote Sensing; Socio-economic; Urban sustainability; Data mining; Kenya

Study areas in Kenya

Methodology workflow

Distribution of positive classified cases for slums for (a) logistic regression, (b) discriminant analysis and (c) the See5 decision tree.
Full Reference:
Mahabir, R., Agouris, P., Stefanidis, A., Croitoru, A. and Crooks, A.T. (2018), Detecting and Mapping Slums using Open Data: A Case Study in Kenya, International Journal of Digital Earth. DOI: https://doi.org/10.1080/17538947.2018.1554010. (pdf)

Thursday, February 01, 2018

The Future of GEOINT: Data Science Will Not Be Enough







In the 2018 State and Future of GEOINT Report published by the The United States Geospatial Intelligence Foundation (USGIF) we had a paper accepted entitled "The Future of GEOINT: Data Science Will Not Be Enough". In the paper we discuss how there has been a deluge of spatial-temporally enabled data in the last several years with no signs of slowing down (e.g. by the year 2020, many experts predict the global universe of accessible data to be on the order of 44 zettabytes—44 trillion gigabytes). With this growth in data there has been steady uptake of data scientists in the GEOINT community because of their  ability to navigate petabytes of raw and unstructured data, then clean, analyze, and visualize the data. However,  we argue that we must go beyond just statistically analyzing data collected on the world around us to truly gain an understanding of the people who inhabit the world.  In order to do this, we suggest that future GEOINT analysts should not only have skills in data science but also be able to apply advanced computational methods, such as agent-based modeling, social network analysis, geographic information science, and deep learning algorithms (i.e. geospatial computational social science) to explore and test hypotheses based on social and geographic theory to truly achieve an understanding of human interactions.  

Full Reference: 
Parrett, C.M., Crooks, A.T. and Pike, T. (2018), The Future of GEOINT: Data Science Will Not Be Enough. The State and Future of GEOINT 2018 Report, The United States Geospatial Intelligence Foundation, Herndon, VA. pp 12-15. (pdf)
As always, any thoughts or comments are welcome. 

Friday, October 13, 2017

AAG2018: Innovations in Urban Analytics

Call for Papers, AAG2018: Innovations in Urban Analytics

We welcome paper submissions for our session at the Association of American Geographers Annual Meeting on 10-14 April, 2018, in New Orleans.

Session Description

New forms of data about people and cities, often termed ‘Big’, are fostering research that is disrupting many traditional fields. This is true in geography, and especially in those more technical branches of the discipline such as computational geography / geocomputation, spatial analytics and statistics, geographical data science, etc. These new forms of micro-level data have lead to new methodological approaches in order to better understand how urban systems behave. Increasingly, these approaches and data are being used to ask questions about how cities can be made more sustainable and efficient in the future.

This session will bring together the latest research in urban analytics. We are particularly interested in papers that engage with the following domains:
  • Agent-based modelling (ABM) and individual-based modelling;
  • Machine learning for urban analytics;
  • Innovations in consumer data analytics for understanding urban systems;
  • Real-time model calibration and data assimilation;
  • Spatio-temporal data analysis;
  • New data, case studies, demonstrators, and tools for the study of urban systems;
  • Complex systems analysis;
  • Geographic data mining and visualization;
  • Frequentist and Bayesian approaches to modelling cities.

Please e-mail the abstract and key words with your expression of intent to Nick Malleson (n.s.malleson@leeds.ac.uk) by 18 October, 2017 (one week before the AAG abstract deadline). Please make sure that your abstract conforms to the AAG guidelines in relation to title, word limit and key words and as specified at: http://annualmeeting.aag.org/submit_an_abstract. An abstract should be no more than 250 words that describe the presentation’s purpose, methods, and conclusions.

For those interested specifically in the interface between research and policy, they might consider submitting their paper to the session “Computation for Public Engagement in Complex Problems” (http://www.gisagents.org/2017/10/call-for-papers-computation-for-public.html).

Key Dates
  • 18 October, 2017: Abstract submission deadline. E-mail Nick Malleson by this date if you are interested in being in this session. Please submit an abstract and key words with your expression of intent.
  • 23 October, 2017: Session finalization and author notification.
  • 25 October, 2017: Final abstract submission to AAG, via the link above. All participants must register individually via this site. Upon registration you will be given a participant number (PIN). Send the PIN and a copy of your final abstract to Nick Malleson (n.s.malleson@leeds.ac.uk). Neither the organizers nor the AAG will edit the abstracts.
  • 8 November, 2017: AAG session organization deadline. Sessions submitted to AAG for approval.
  • 9-14 April, 2018: AAG Annual Meeting.

Session Organizers

Thursday, August 17, 2017

Big Data, Agents and the City


In the recently published book "Big Data for Regional Science" edited by Laurie Schintler and  Zhenhua Chen, Nick Malleson, Sarah Wise, and Alison Heppenstall and myself have a chapter entitled: Big Data, Agents and the City. In the chapter we discuss how big data can be used with respect to building more powerful agent-based models. Specifically how data from say social media could be used to inform agents behaviors and their dynamics; along with helping with the calibration and validation of such models with a emphasis on urban systems. 

Below you can read the abstract of the chapter, see some of the figures we used to support our discussion, along with the full reference and a pdf proof of the chapter. As always any thoughts or comments are welcome.


Abstract:
Big Data (BD) offers researchers the scope to simulate population behavior through vastly more powerful Agent Based Models (ABMs), presenting exciting opportunities in the design and appraisal of policies and plans. Agent-based simulations capture system richness by representing micro-level agent choices and their dynamic interactions. They aid analysis of the processes which drive emergent population level phenomena, their change in the future, and their response to interventions. The potential of ABMs has led to a major increase in applications, yet models are limited in that the individual-level data required for robust, reliable calibration are often only available in aggregate form. New (‘big’) sources of data offer a wealth of information about the behavior (e.g. movements, actions, decisions) of individuals. By building ABMs with BD, it is possible to simulate society across many application areas, providing insight into the behavior, interactions, and wider social processes that drive urban systems. This chapter will discuss, in context of urban simulation, how BD can unlock the potential of ABMs, and how ABMs can leverage real value from BD.  In particular, we will focus on how BD can improve an agent’s abstract behavioral representation and suggest how combining these approaches can both reveal new insights into urban simulation, and also address some of the most pressing issues in agent-based modeling; particularly those of calibration and validation.

Keywords: Agent-based models, Big Data, Emergence, Cities.

The growth in Agent-based modeling -from search results of Web of Science and Google Scholar.

Hotspots of activity of Tweeter Users: Tweet locations and associated densities for a selection of prolific users.

Full Reference:
Crooks, A.T., Malleson, N., Wise, S. and Heppenstall, A. (2018), Big Data, Agents and the City, in Schintler, L.A. and Chen, Z. (eds.), Big Data for Urban and Regional Science, Routledge, New York, NY, pp. 204-213. (pdf)

Tuesday, October 04, 2016

Agent-based Modeling in Geographical Systems

Recently Alison Heppenstall and myslef wrote a short introductory chapter entitled "Agent-based Modeling in Geographical Systems" for AccessScience (a online version of McGraw-Hill Encyclopedia of Science and Technology).

In the chapter we trace the rise in agent-based modeling within geographical systems with a specific emphasis of cities. We briefly outline how thinking and modeling cities has changed and how agent-based models align with this thinking along with giving a selection of example applications. We discuss the current limitations of agent-based models and ways of overcoming them and how such models can and have been used to support real world decision-making.

Conceptualization of an agent-based model where people are connected to each other and take actions when a specific condition is met

 Full Reference:
Heppenstall, A. and Crooks, A.T. (2016). Agent-based Modeling in Geographical Systems, AccessScience, McGraw-Hill Education, Columbus, OH. DOI: http://dx.doi.org/10.1036/1097-8542.YB160741. (pdf)
 

Saturday, October 01, 2016

New Paper: User-Generated Big Data and Urban Morphology

Continuing our work with crowdsourcing and geosocial analysis we recently had a paper published in a special issue of the  Built Environment journal entitled "User-Generated Big Data and Urban Morphology."

The theme of the special issue is: "Big Data and the City" which was guest edited by Mike Batty and includes 12 papers.  To quote from the website

"This cutting edge special issue responds to the latest digital revolution, setting out the state of the art of the new technologies around so-called Big Data, critically examining the hyperbole surrounding smartness and other claims, and relating it to age-old urban challenges. Big data is everywhere, largely generated by automated systems operating in real time that potentially tell us how cities are performing and changing. A product of the smart city, it is providing us with novel data sets that suggest ways in which we might plan better, and design more sustainable environments. The articles in this issue tell us how scientists and planners are using big data to better understand everything from new forms of mobility in transport systems to new uses of social media. Together, they reveal how visualization is fast becoming an integral part of developing a thorough understanding of our cities."
Table of Contents

In our paper we discuss and show how crowdsourced data is leading to the emergence of alternate views of urban morphology that better capture the intricate nature of urban environments and their dynamics. Specifically how such data can provide us information pertaining to linked spaces and geosocial neighborhoods. We argue that a geosocial neighborhood is not defined by its administrative boundaries, planning zones, or physical barriers, but rather by its emergence as an organic self-organized social construct that is embedded in geographical spaces that are linked by human activity. Below is the abstract of the paper and some of the figures we have in it which showcase our work.
"Traditionally urban morphology has been the study of cities as human habitats through the analysis of their tangible, physical artefacts. Such artefacts are outcomes of complex social and economic forces, and their study is primarily driven by traditional modes of data collection (e.g. based on censuses, physical surveys, and mapping). The emergence of Web 2.0 and through its applications, platforms and mechanisms that foster user-generated contributions to be made, disseminated, and debated in cyberspace, is providing a new lens in the study of urban morphology. In this paper, we showcase ways in which user-generated ‘big data’ can be harvested and analyzed to generate snapshots and impressionistic views of the urban landscape in physical terms. We discuss and support through representative examples the potential of such analysis in revealing how urban spaces are perceived by the general public, establishing links between tangible artefacts and cyber-social elements. These links may be in the form of references to, observations about, or events that enrich and move beyond the traditional physical characteristics of various locations. This leads to the emergence of alternate views of urban morphology that better capture the intricate nature of urban environments and their dynamics."

Keywords: Urban Morphology, Social Media, GeoSocial, Cities, Big Data.
City Infoscapes – Fusing Data from Physical (L1, L2), Social, Perceptual (L3) Spaces to Derive Place Abstractions (L4) for Different Locations (N1, N2).


Recreational Hotspots Composed of “Locals” and “Tourists” with Perceived Artifacts Indicating “Use” and “Need”. (A) High Line Park (B) Madison Square Garden.



Moving from Spatial Neighborhoods to Geosocial Neighborhoods via Links.

The Emergence of Geosocial Neighborhoods after the in the
Aftermath of the 2013 Boston Marathon Bombing


Full  Reference: 
Crooks, A.T., Croitoru, A., Jenkins, A., Mahabir, R., Agouris, P. and Stefanidis A. (2016). “User-Generated Big Data and Urban Morphology,”  Built Environment, 42 (3): 396-414. (pdf)

Monday, August 15, 2016

Summer Projects

Over the summer, Arie Croitoru and myself took part in the George Mason University Aspiring Scientists Summer Internship Program. We worked with three very talented high-school students who over the course of the seven and a half week program produced some excellent research around the areas of agent-based modeling and social media analysis. An overview of their work can be seen in the posters and abstracts that the students produced at the end of the internship.

Lawrence Wang explored how social media could be used with respect to predicting election results under a project entitled "And the Winner Is? Predicting Election Results using Social Media". Below you can read Lawrence's abstract and see his poster.

"The 2012 U.S. presidential election demonstrated how Twitter can serve as a widely accessible forum of political discourse. Recently, researchers have investigated whether social media, particularly Twitter, can function as a predictive tool. In the past decade, multiple studies have claimed to successfully predict the results of elections using Twitter data. However, many of these studies fail to account for the inherent population bias present in Twitter data, leading to ungeneralizable results. In this project, I investigate the prospects of using Twitter data as an alternative to poll data for predicting the 2012 presidential election. The tweet corpus consisted of tweets published one month before the November election day. Using VADER, a sentiment analysis tool, I analyzed over 140,000 tweets for political sentiment. I attempted to circumvent the Twitter population bias by comparing age, race, and gender metrics of the Twitter population with that of the U.S. population. Furthermore, I utilized Bayesian inference with prior distributions from the results of the 2008 presidential election in order to mitigate the effects of limited tweet data in certain states. The resulting model correctly predicted the likely outcomes of 46 of the 50 states and predicted that President Obama would be reelected with a probability of 0.945. Such a model could be used to explore the forthcoming elections. " 


In a second project, Varun Talwar, explored how knowledge bases could be utilized to better contextualize social media discussions with a project entitled "Context Graphs: A Knowledge-Driven Model for Contextualizing Twitter Discourse." Below you can read Varun's project abstract and his end of project poster.

"Introduction: User posted content through online social media (SM) platforms in recent years has emerged as a rich field for narrative analysis of topics captured during the discussion discourse. In particular, collective discourse has been used to manually contextualize public perception of health related events.

Objective: As SM feeds tend to be noisy, automated detection of the context of a given SM discourse stream has proven to be a challenging task. The primary objective of this research is to explore how existing knowledge bases could be utilized to better contextualize SM discussions through topic modeling and mining. By utilizing such existing knowledge it would then be possible to explore to what extent a given discourse is related to a known or a new context, as well as compare and contrast SM discussions through their respective contexts.

Methods: In order to accomplish these goals this research proposes a novel approach for contextualizing SM discourse. In this approach, topic modeling is combined with a knowledgebase in a two-step process. First, key topics are extracted from a SM data corpus by applying a statistical topic-modeling algorithm, a process that also results in data dimensionality reduction. Once a set of salient topics are extracted, each topic is then used to mine the knowledge base for sub graphs that represent the contextual linkages between knowledge elements. Such sub-graphs can then further disambiguate the topic modeling results, and be utilized for qualifying context similarity across SM discussions.

Results: The time-series analysis of the Twitter discourse via graph-matching algorithms reveals the change in topics as evidenced by the emergence of the terms “pregnancy” and “abortion” as information about the virus propagated through the Twitter community. "




Elizabeth Hu explored the current migration crisis in Europe in a project entitled "Across the Sea: A Novel Agent-Based Model for the Migratory Patterns of the European Refugee Crisis". Below is Elizabeth's abstract, poster and an example model run.

"Since 2010, a growing number of refugees have sought asylum in European nations, fleeing violence and military conflict in their home countries. Most of the refugees originate from Syria, Iraq, Afghanistan, and African nations. The vast majority of refugees risk their lives in the popular yet perilous Mediterranean Sea Route often prone to boat accidents and subsequent deaths of migrants.  The flow of millions of refugees has introduced a humanitarian crisis not seen since World War II. European nations are struggling to cope with the influx of refugees through various border policies.

In order to explore this crisis, a geographically explicit agent-based model has been developed to study the past and future patterns of refugee flows. Traditional migration models, which represent the population as an aggregate, fail to consider individual decision-making processes based on personal status and intervening opportunities. However, the novel agent-based model developed here of migration allows population behavior to emerge as the result of individual decisions. Initial population, city, and route attributes are based upon data from the UNHCR, EU agencies, crowd-sourced databases, and news articles. The agents, refugees, select goal destinations in accordance with the Law of Intervening Opportunities. Thus, goals are prone to change with fluctuating personal needs. Agents choose routes not only based on distance, but also other relevant route attributes. The resulting migration flows generated by the model under various circumstances could provide crucial guidance for policy and humanitarian aid decisions."



The movie below gives a sense of the migration paths the refugees are taking.




Wednesday, March 23, 2016

Call For Papers: Smart Buildings and Cities


Special Issue on Smart Buildings and Cities for IEEE Pervasive Computing


Submission deadline: 1 July 2016  Extended to July 18th, 2016
Publication date: April–June 2017


One of Mark Weiser’s first envisionments of ubiquitous and pervasive computing had the smart home as its central core. Since then, researchers focused on realizing this vision have built out from the smart home to the smart city. Such environments aim to improve the transparency of information and the quality of life through access to smarter and more appropriate services.

Despite efforts to build these environments, there are still many unanswered questions: What does it mean to make a building or a city “smart”? What infrastructure is necessary to support smart environments? What is the return on investment of a smart environment?

The key to building smart environments is the fusion of multiple technologies including sensing, advanced networks, the Internet of Things, cloud computing, big data analytics, and mobile devices. This special issue aims to explore new technologies, methodologies, case studies, and applications related to smart buildings and cities. Contributions may come from diverse fields such as distributed systems, HCI, ambient intelligence, architecture, transportation and urban planning, policy development, and cyber-physical systems. Relevant topics for issue include
  • Applications, evaluations, or case studies of smart buildings/cities
  • Architectures and systems software to support smart environments
  • Big data analytics for monitoring and managing smart environments
  • Economic models for smart buildings/cities
  • Models for user interaction in smart environments
  • Formative studies regarding the design, use, and acceptance of smart services
  • Configuration and management of smart environments
  • Embedded, mobile ,and crowd sensing approaches
  • Cloud computing for smart environments
  • Domain-specific investigations (such as transportation or healthcare)
The guest editors invite original and high-quality submissions addressing all aspects of this field, as long as the connection to the focus topic is clear and emphasized.

Guest Editors
Submission Information

Tuesday, February 16, 2016

Call For Papers: Rethinking the ABCs

Readers of the blog might be interested in a workshop being organized by Daniel Brown, Eun-Kyeong Kim, Liliana Perez, and Raja Sengupta entitled:


Rethinking the ABCs: Agent-Based Models and Complexity Science in the age of Big Data, CyberGIS, and Sensor networks

September 27th, 2016 in Montreal, Canada


To quote from the call:

"A broad scope of concepts and methodologies from complexity science – including Agent-Based Models, Cellular Automata, network theory, chaos theory, and scaling relations – has contributed to a better understanding of spatial/temporal dynamics of complex geographic patterns and process.

Recent advances in computational technologies such as Big Data, Cloud Computing and CyberGIS platforms, and Sensor Networks (i.e. the Internet of Things) provides both new opportunities and raises new challenges for ABM and complexity theory research within GIScience. Challenges include parameterization of complex models with volumes of georeferenced data being generated, scale model applications to realistic simulations over broader geographic extents, explore the challenges in their deployment across large networks to take advantage of increased computational power, and validate their output using real-time data, as well as measure the impact of the simulation on knowledge, information and decision-making both locally and globally via the world wide web.

The scope of this workshop is to explore novel complexity science approaches to dynamic geographic phenomena and their applications, addressing challenges and enriching research methodologies in geography in a Big Data Era."

More information about the workshop can be found at https://sites.psu.edu/bigcomplexitygisci/

Tuesday, January 26, 2016

“Space, the Final Frontier”: How Good are Agent-Based Models at Simulating Individuals and Space in Cities?

Recently, Alison Heppenstall, Nick Malleson  and myself have just had a paper accepted in Systems entitled: “Space, the Final Frontier”: How Good are Agent-Based Models at Simulating Individuals and Space in Cities?" In the paper we critically examine how well agent-based models have  simulated a variety of urban processes. We discus what considerations are needed when choosing the appropriate level of spatial analysis and time frame to model urban phenomena and what role Big Data can play in agent-based modeling. Below you can read the abstract of the paper and see a number of example applications discussed.
Abstract: Cities are complex systems, comprising of many interacting parts. How we simulate and understand causality in urban systems is continually evolving. Over the last decade the agent-based modeling (ABM) paradigm has provided a new lens for understanding the effects of interactions of individuals and how through such interactions macro structures emerge, both in the social and physical environment of cities. However, such a paradigm has been hindered due to computational power and a lack of large fine scale datasets. Within the last few years we have witnessed a massive increase in computational processing power and storage, combined with the onset of Big Data. Today geographers find themselves in a data rich era. We now have access to a variety of data sources (e.g., social media, mobile phone data, etc.) that tells us how, and when, individuals are using urban spaces. These data raise several questions: can we effectively use them to understand and model cities as complex entities? How well have ABM approaches lent themselves to simulating the dynamics of urban processes? What has been, or will be, the influence of Big Data on increasing our ability to understand and simulate cities? What is the appropriate level of spatial analysis and time frame to model urban phenomena? Within this paper we discuss these questions using several examples of ABM applied to urban geography to begin a dialogue about the utility of ABM for urban modeling. The arguments that the paper raises are applicable across the wider research environment where researchers are considering using this approach.
Keywords: cities; agent-based modeling; big data; crime; retail; space; simulation

Figure 1. (A) System structure; (B) System hierarchy; and (C) Related subsystems/processes (adapted from Batty, 2013).



Reference cited:
Batty, M. (2013).  The New Science of Cities; MIT Press: Cambridge, MA, USA.

Full reference to the open access paper:
Heppenstall, A., Malleson, N. and Crooks A.T. (2016). “Space, the Final Frontier”: How Good are Agent-based Models at Simulating Individuals and Space in Cities?, Systems, 4(1), 9; doi: 10.3390/systems4010009 (pdf)
 

Wednesday, April 22, 2015

Leveraging Crowdsourced data for Agent-based modeling: Opportunities, Examples and Challenges

This week I am attending the AAG Annual Meeting in Chicago. While here, we organized 3 sessions entitled "Geosimulation and Big Data: A Marriage made in Heaven or Hell?" in which I presented a paper, co-authored with Sarah Wise: "Leveraging Crowdsourced data for Agent-based modeling: Opportunities, Examples and Challenges." The abstract is below:
The rise of crowdsourcing has made new kinds of data available to the  geographical community. New forms of data range in their characteristics and purpose. One example is Volunteered Geographical Information (VGI), were users purposely contribute Geographic Information (GI) as in the case of OpenStreetMap; another is Ambient Geographic Information (AGI), where the intention of contributors is not necessarily to provide GI, but GI can be derived, as from Twitter. While much progress has been made in utilizing these new sources of data in GIScience, they have only recently begun to be integrated into agent-based models (ABM). This paper will discuss the opportunities that crowdsourced data provides for ABMs, specifically focusing on how such information gives us a new lens to study the micro-interactions of individuals. Through as series of examples we will demonstrate how such data can be integrated into geographically explicit ABMs. By building on these examples we will showcase how the spatial environment and agent populations can be built using crowdsourced information, and highlight how agent behaviors can be informed and validated by such information. We further discuss the challenges associated with this program of research: using such data is not without its difficulties, including gathering or accessing the data, storing the data, analyzing the collected data, and assessing its validity. Together, this work provides a brief overview of the current state of crowdsourced data-informed ABM. 
If you like what is written above, you can have a flick through the slides from the talk or check out one of the movies:




Monday, February 09, 2015

Geosimulation and Big Data: A Marriage made in Heaven or Hell? Schedule

Do you like big data and geosimulation and wondering when to book flights or which sessions to attend at the forthcoming AAG Annual Meeting,  If so, you might like our sessions entitled "Geosimulation and Big Data: A Marriage made in Heaven or Hell? " taking place on Wednesday the 22nd of April 2015.

Abstract of the Sessions:

In recent years, human emotions, intentions, moods and behaviors have been digitised to an extent previously unimagined in the social sciences. This has been in the main due to the rise of a vast array of new data, termed 'Big Data'.  These new forms of data have the potential to reshape the future directions of social science research, in particular the methods that scientists use to model and simulate spatially explicit social systems. Given the novelty of this potential "revolution" and the surprising lack of reliable behavioral insight to arise from Big Data research, it is an opportune time to assess the progress that has been made and consider the future directions of socio-spatial modelling in a world that is becoming increasingly well described by Big Data sources.

In these sessions we will have methodological, theoretical and empirical papers that that engage with any aspect of geospatial modelling and the use of Big Data. We are particularly interested in the ways that insight into individual or group behavior can be elucidated from new data sources - including social media contributions, volunteered geographical information, mobile telephone transactions, individually-sensed data, crowd-sourced information, etc. -  and used to improve models or simulations.  Topics include, but are not limited to:
  • Using Big Data to inform individual-based models of geographical systems;
  • Translating Big Data into agent rules;
  • Elucidating behavioral information from diverse data;
  • Improving simulated agent behavior;
  • Validating agent-based models (ABM) with Big Data;
  • Ethics of data collected en masse and their use in simulation.
2192 Geosimulation and Big Data: A Marriage made in Heaven or Hell? (1)

Wednesday, 4/22/2015.
8:00 AM - 9:40 AM.
600a Classroom, University of Chicago Gleacher Center, 6th Floor.

Chair: Nick Malleson 

Abstracts:

*Atsushi Nara:
A GPGPU approach for simulating and analyzing human dynamics
*Kira KowalskaJohn Shawe-Taylor and Paul Longley:
 Data-driven modelling of police patrol activity 
*Martin Zaltz Austwick, Gustavo Romanillos Arroyo and Borka Moya-Gomez:
Simulating Rush Hour Bicycle Traffic in Madrid 
*Hai Lan  and Paul Torrens:
Voxel based Cellular Automata with massive cells for Geo-simulation: Ice dynamics simulation in Antarctic locations as example
*Philippe J. Giabbanelli, Thomas Burgoine, Pablo Monsivais and James Woodcock:
Using big data to develop individual-centric models of food behaviours

2292 Geosimulation and Big Data: A Marriage made in Heaven or Hell? (2) 

Wednesday, 4/22/2015.
10:00 AM - 11:40 AM.
600a Classroom, University of Chicago Gleacher Center, 6th Floor.

Chair: Alison Heppenstall

Abstracts:

*Kostas Cheliotis:
Coupling Public Space Simulations with Real-Time Data Streams 
*Andrew Crooks and Sarah Wise:
Leveraging Crowdsourced data for Agent-based modeling: Opportunities, Examples and Challenges 
*Ed Manley, Chen Zhong and Michael Batty:
Towards Real-Time Simulation of Transportation Disruption - Building Agent Populations from Big Mobility Data 
*Alison Heppenstall, *Nick Malleson and Andrew Evans:
Evaluating Big Data demographics for population modelling 
Muhammad Adnan, Alistair Leak and *Paul Longley:
Exploring the geo-temporal patterns of Twitter messages

2492 Geosimulation and Big Data: A Marriage made in Heaven or Hell? (3) Discussion Session

Wednesday, 4/22/2015.
1:20 PM - 3:00 PM.
600a Classroom, University of Chicago Gleacher Center, 6th Floor.

Chair: Nick Malleson

Abstracts:
 
*Paul M Torrens and Hai Lan:
Micro big data and geosimulation 
*Mark Birkin:
The Ten Commandments of Big Data 
 2:00 PM to 3:00PM: Discussion

 Organizers

  • Alison Heppenstall, School of Geography, University of Leeds
  • Nick Malleson, School of Geography, University of Leeds
  • Andrew Crooks, Department of Computational Social Science, George Mason University
  • Paul Torrens, Department of Geographical Sciences, University of Maryland
  • Ed Manley, Centre for Advanced Spatial Analysis, University College London