Showing posts with label Large Language Models. Show all posts
Showing posts with label Large Language Models. Show all posts

Wednesday, August 26, 2026

A Dual-LLM Supervised Workflow for Replicating Agent-based Models

In the past we have written about large language models (LLMs) and how these can be used in agent-based modeling. However, these previous posts only touched the surface on what is possible. One area we are currently exploring is how LLMs can aid in model replication, which is a major challenge in agent-based modeling. In the sense, there are countless agent-based models but very few are replicated and if  a model is to withstand the test of time, replication is needed. 

To this end, Boyu Wang, JoAnn Lee and myself have an extended abstract entitled: "A Dual-LLM Supervised Workflow for Replicating Agent-based Models: From NetLogo to Mesa" at the 2026 Social Simulation Conference. In this work we demonstrate how LLMs can be used to replicate an exiting model into another modeling package, Specifically, we take the Ya-TASERPS model which was initially implemented in NetLogo and re-implement it in Mesa via a  LLM workflow. 

The workflow uses two GPT-5.4 models with distinct LLM roles under continuous human mediation. One LLM role handles planning and evaluation, from checking source code, decomposing tasks into phases, and reviewing whether implementations are consistent with the source NetLogo model. The other LLM role performs the actual implementations in phases, as outlined by the first LLM role.

Each task phase defines acceptance criteria, proceeds through implementation and verification, and ends with human review before progression. The human researcher mediates both LLM roles, resolves ambiguities, and decides whether the resulting model state is acceptable. The aim is not towards full automation but to reduce prompt drift, limit uncontrolled rewrites, and keep generated code subordinate to explicit validation. 

If you want to read about our replication, why we chose the Ya-TASERPS model along with our findings, please feel free to read the paper and more information about this can be found at: https://github.com/wang-boyu/Ya-TASERPS.

The NetLogo-to-Mesa replication workflow utilizing two distinct LLM roles.

Large-run prosocial outcomes across the six user-controlled inputs in NetLogo and Mesa with the same set of 729 parameter combinations, 10 random seeds, and 1800 days per run. The two implementations preserve the same broad significance and directional patterns.

Full Reference: 

Wang, B., Lee, J. and Crooks, A.T.  (2026), A Dual-LLM Supervised Workflow for Replicating Agent-based Models: From NetLogo to Mesa. Social Simulation Conference 2026, Durham, UK. (pdf)

Thursday, April 02, 2026

Research Updates: AAG 2026


At the AAG Annual meeting this year, two of my students gave talks about their ongoing research. Ying Zhou presented her work with a talk entitled "Exploring the Relationship between Urban Morphology and People’s Emotions." In this talk, Ying showed how one could mine social media posts to gain a sense of how different emotions are spatially spread around a city using New York city as a case study. If this sounds of interest, below you can see the abstract of the talk, the research methodology and a sample of the results.  

Abstract: 
Urban morphology records physical information about spatial patterns (e.g., streets and land use) and their evolution over time, as well as human settlement information. People who live in or visit a city gain experiences through interaction with its spatial patterns, and these experiences influence people’s emotions. Therefore, it is necessary to explore the spatial relationships between urban morphology and people’s emotions. Taking New York City as a case study, this research uses social media data to obtain and locate people's emotions in different parts of the city. To extract the emotion relating to specific space, we use the RoBERTa-based model to label texts in social media with six primary emotions (i.e., happiness, sadness, fear, anger, surprise, and disgust). We then used DBSCAN to identify spatial clustering features of these emotions. Finally, we compared the clustered emotions with urban morphology (both in terms of both its form and function) and how such emotions evolve and change over a span of five years. Such analysis reveals the relationship between people’s emotions and broader setting that they inhabit (i.e., the city). Moreover, these works offer bottom-up insights into how urban morphology shapes people’s feelings, which can serve as feedback for urban planning and management.
 
Keywords: Urban Morphology, Emotion Detection, Spatial Analysis, Urban Studies.



While in another talk, Boyu Wang continues to add new functionality to the Mesa, a python agent-based modeling toolkit, this time in the form of utilizing large language models for agent-based decision making, with a talk entitled "Mesa-LLM: Generative agent-based modeling with large language models empowered agents" 

If this sounds of interest, below you can see the abstract of the talk, along with the Mesa-LLM architecture. While further details about Mesa-LLM can be found on Boyu's GitHub page: https://github.com/mesa/mesa-llm.

Abstract 

Agent‐based models (ABMs) have long been used to examine how individual behaviors give rise to aggregated social and spatial phenomena. Mesa, an open source ABM library in Python, provides modular components and browser based visualization to create and analyze agent based models in the PyData ecosystem. Agents’ behaviours in these models are often governed by rule-based decisions. The recent advancements of large language models (LLMs) have created a new paradigm, namely generative agent-based modeling, where LLMs are integrated as decision-making engines so that agents can communicate, negotiate, and decide based on natural language. In this paper, we introduce Mesa-LLM, an LLM extension to the Mesa framework. Its modular design allows users to customize reasoning, memory and planning components and plug in different LLMs (e.g., GPT, Gemini, Llama). We demonstrate Mesa-LLM through Epstein’s civil violence model. In contrast to the classical model where agents act based on calculated probabilities and pre-defined thresholds, agents through Mesa-LLM have their decisions articulated in natural language. This demonstration shows how an archetypal ABM can be enriched by language-based decision making to explore complex social dynamics such as protest escalation. Through this simple example, we highlight how incorporating LLMs into ABMs opens new possibilities for geographers to model human behavior from the bottom up by leveraging generative artificial intelligence (GenAI).
 Keywords: Agent-Based Modeling, Large Language Model, AI Agent, Python.

References 

Wang, B., Frisch, C., Nair, S., Kazil, J. and Crooks, A.T. (2026), Mesa-LLM: Generative Agent-Based Modeling with Large Language Models Empowered Agents, The Association of American Geographers (AAG) Annual Meeting, 17th –21th March, San Francisco, CA. (pdf)

Zhou, Y. and Crooks, A.T. (2026), Exploring the Relationship between Urban Morphology and People’s Emotions, The Association of American Geographers (AAG) Annual Meeting, 17th –21th March, San Francisco, CA. (pdf)

Monday, December 15, 2025

Creating and Assessing an Unconventional Global Database of Dust Storms Utilizing Generative AI

In the past we have written about how one can use social media to monitor dust storms along with how multi-modal large language models (MLLMs) can be used to analyze images. At the recent American Geophysical Union (AGU) Fall Meeting we (Sage Keidel, Stuart Evans and myself) brought these two strands of research together in a poster entitled "Creating and Assessing an Unconventional Global Database of Dust Storms Utilizing Generative AI."

In this work we showcase how MLLMs are providing new opportunities and accessible methods for information extraction from imagery data using geo-located images from Flickr which have a dust keyword tag associated with it from multiple languages (e.g., Arabic, English, Spanish).  We run these images through ChatGPT, which classifies them as dust storms or not and compare this classification with human classifed images. If this sounds of interest, below you can read the abstract, see the poster along with a selection of images that have been labeled as as dust storm or not and ChatGPTs confidence in its classification. While the dust storm database itself can be found here

Abstract:

Complete observations of dust events are difficult, as dust’s spatial and temporal variability means satellites may miss dust due to overpass time or cloud coverage, while ground stations may miss dust due to not being in the plume. As a result, an unknown number of dust events go unrecorded in traditional datasets. Dust’s importance both for atmospheric processes and as a health and travel hazard makes detecting dust events whenever possible important, and in particular, studies of the health impacts of dust are limited by detailed exposure information. 

In recent years, social media platforms have emerged as a valuable source of unconventional data to study events such as earthquakes and flooding around the world. However, one challenge with respect to using such data is classifying and labeling it (i.e., is it a dust storm or not?). While it is relatively simple to classify textural data through natural language processing, it is not the case with imagery data. Traditionally, classifying imagery data was a complex computer vision task. However, recent advancements in generative artificial intelligence (AI) especially multi-modal large language models (MLLMs) are opening up new opportunities and offering accessible methods for information extraction from imagery data. Therefore, in this study we collected geotagged Flickr images referencing dust from around the globe from multiple languages (e.g., English, Spanish, Arabic) and use generative AI (i.e., ChatGPT) to classify the images as dust storms or not. Furthermore, we compare a sample of these classified images from ChatGPT with human classified images to assess its accuracy in classification. Our results suggest that ChatGPT can relatively accurately detect dust storms from Flickr images and thus helps us create an unconventional global database of dust storm events that might otherwise go unobserved from more traditional datasets.



Workflow

Poster

Dust storm database (click here to go to it)

Full Referece: 
Keidel, S., Evans S. and Crooks, A.T. (2025), Creating and Assessing an Unconventional Global Database of Dust Storms Utilizing Generative AI, American Geophysical Union (AGU) Fall Meeting, 15th–19th December, New Orleans, LA. (pdf of poster).

Friday, August 01, 2025

LLMs and ABMs

In a previous post we talked about the potential of Generative AI for urban modeling, keeping with this theme at the 11th International Conference on Computational Social Science (IC2S2), Na Jiang, Boyu Wang and myself had a poster entitled  Agent-based Models with Large Language Models: Two Modeling Examples. 

In this poster and extended abstract we detail how LLMs can help with many aspects of agent-based modeling development. If this sounds of interest, below you can see the abstract, the poster and the full referece and link to the extended abstract .

Abstract:

Large language models (LLMs) play an important role in AI-powered code assistants such as code completion, debugging, and documentation. Such models can be further fine-tuned on smaller amount of data for specific tasks, often with the improvement of performance compared to generic LLMs. However, such fine-tuning techniques are seldomly used in generating sophisticated agent-based models (ABMs), because they are often implemented as software that demands extra standards such as the Overview, Design concepts, and Details (ODD) protocol. This research examines how we can bridge this gap by utilizing LLMs in designing or conceptualizing, building, and running agent-based models in the form of user prompts. In this work, two models are created to demonstrate the proposed method. Specifically, Sakoda's checkerboard model of social interaction is created by LLM from explicit design and description through prompts. The other model stimulates consumer preferences and restaurant visits as designed and implemented by a LLM. These models are evaluated by human experts on their code correctness and quality for both verification and validation purposes. This work serves as a first step towards fine-tuned LLMs on existing models and documentations to create high-quality and functional ABMs based on either user prompts or standard protocols, contributing to further exploration on the future of AI-assisted geospatial simulation development.

Keywords: agent-based modeling, geospatial simulations, large language models, generative AI, coding 


Full reference: 

Jiang, N., Wang, B. and Crooks, A.T. (2025), Agent-based Models with Large Language Models: Two Modeling Examples, 11th International Conference on Computational Social Science (IC2S2), 21-24th July, Norrkoping, Sweden. (extended abstract pdf) (poster pdf)

Saturday, July 05, 2025

New Editorial: Generative AI and Urban Modeling

In the current issue of Environment and Planning B, we (Boyu Wang, Na Jiang and myself) have a new editorial entitled "Generative AI and Urban Modeling". The premise of this editorial is that Generative AI (GenAI) is impacting all aspects of our daily lives and as such has we were wondering how will it impact urban modeling? 

For example, in the editorial we discuss how  GenAI could speed up the overall urban modeling process. To demonstrate this we show how ChatGPT (and its built-in coding interface Canvas) can take published papers and build agent-based models from them (one being of an abstract space and another being spatially explicit). 

However, while model building is time consuming task, another challenge modelers face is how to incorporate decision making within them. To this end we also discuss how large language models (LLMs) have the potential to help with  agent-decision making in the form of generating  agent-personas or scheduling agent activities. 

We conclude the editorial with a series of questions: how will GenAI impact urban modeling? Will it open up the field to more people without the need for strong coding skills? Will we see growth in using LLMs for generating behavior? Will GenAI lead to a new generation of modeling toolkits? While these are only a short list of questions, they also raise concerns that relate back to some of the more thorny issues of urban modeling, that of verification and validation. 

If this sounds of interest you can read the full editorial here. 

Full Referece: 

Crooks, A.T., Jiang, N. and Wang, B. (2025), Generative AI and Urban Modeling, Environment and Planning B, 52(6), 1277-1281. (pdf)

Monday, March 31, 2025

AAG 2025 Talks

As the AAG has just wrapped up I thought I would write brief (well actually quite long) post on the talks that I was involved with at the conference. These talks would not have been possible without the many great students and colleagues who I have been collaborating with over time. Below you will find a brief summary of the talks and if any sound interesting, please reach out and we can give you more details. 

First up (in order in which they were presented) was "Utilizing Streetview Images for Mapping Building Attributes with ChatGPT" with Qingqing Chen and Linda See. In this talk we discussed how multimodal Large Language Models are giving us a new way to study cities, in the sense, lowering the boundary for information extraction. Using ChatGPT and street view images from Mapillary as an example, we showed how one can extract building age, usage (e.g., commercial, mixed use, residential) and estimate building height  which could all be used to inform urban climate models which require detailed information on buildings. 

Abstract: 

With increasing rates of urbanization, many challenges are emerging regarding sustainability such as the energy usage of buildings. Coinciding with this is the growing attention of urban climate models for energy demand estimation and climate adaptation strategies. However, the applicability of these models is constrained by the lack of detailed urban surface information. Therefore, creating comprehensive datasets that capture urban surface information at a granular scale is crucial for responding to our rapidly urbanizing world. Recent advancements in Large Language Models (LLMs) have opened new opportunities in urban studies, offering accessible methods for information extraction. In this talk we explore the feasibility of ChatGPT to extract building attributes from images. Taking New York City as a case study, we collect building images from Mapillary and process them through ChatGPT by posing specific questions to extract building attributes (e.g., height, functions, age). These attributes are then compared with authoritative data. The proposed method helps address the current dearth of fine-grained surface data on urban issues, therefore enhancing the accuracy and utility of urban climate models. Overall, this study demonstrates the practical applications of ChatGPT in geographic knowledge extraction, advancing the understanding of LLMs in geographic contexts, and more broadly to the discourse on Artificial Intelligence (AI) in urban modeling and climate science. 

Keywords: Buildings, ChatGPT, Large Language Models (LLMs), Mapillary, Street View Images (SVI), GeoAI.

Example of Workflow.

Reference: 

Chen, Q., See, L. and Crooks, A.T. (2025), Utilizing Streetview Images for Mapping Building Attributes with ChatGPT, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)

This was followed by a talk by lead by Qingqing Chen entitled "Multi-sensory Experiences: The Connection Between the Smell and Vision in Understanding Urban Environments" where we explored to what extent can visual data from street view imagery be used as a proxy for capturing large-scale urban smell perceptions when compared to geosocial media. Such as what visual cues evoke specific smell perceptions.

Abstract:

Smell is a crucial transversal sense, which bridges the tangible aspects of urban environments, such as exhaust and garbage, with their intangible impacts on emotions, social interactions and well-being. Despite its crucial role in our everyday life, many urban studies primarily focus on the visual dimension, potentially introducing biases in our understanding of urban spaces. This research transcends this visual-centric bias by integrating the olfactory perceptions to investigate the nuanced relationship between smell and vision in urban environments. Specifically, we utilize advanced semantic segmentation to extract visual elements from street view imagery (i.e., Mapillay) and apply casual forest analysis to examine their causal effects on smell expectations recorded from human participants. These expectations, often tied to personal experiences and/or cultural associations, are compared with real-environment smell experiences derived from geosocial media (i.e., Twitter/X). The results show that visual cues can predict smells in straightforward urban settings, such as small parks or less densely populated areas. However, in complex urban environments, the predictive power of visual cues diminishes as diverse and overlapping scents obscure specific smells, even in visually distinct areas. These findings underscore the importance of a multisensory approach in urban studies, enhancing our understanding of the complex interplay between sensory experiences and informing urban design strategies that integrate multiple senses to create more engaging and inclusive environments. This is especially important for individuals with sensory impairments, such as anosmia or visual impairments, who rely on other senses to compensate for their perception of urban environments. 

Keywords: Multi-sensory Experiences, Smell and Vision; Semantic Segmentation, Causal Effects, Geosocial Media, Street View Imagery (SVI).

Workflow

Reference: 

Chen, Q. and Crooks, A.T. (2025), Multi-sensory Experiences: The Connection Between the Smell and Vision in Understanding Urban Environments, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)

In the Geosimulation session that we organized, we had a talk entitled "Large Language Models for Conceptualizing, Designing, and Generating Agent-based Models" where Na Jiang, Boyu Wang and myself presented our work on exploring using multimodal Large Language Models (LLMs) to create age-based models. In the sense as modelers, we spend a lot of time developing and writing code and we were curious what could be done though the use of LLMs. 

To give a sense of what is possible, below is an example of using ChatGPT for creating a model from a published paper.



Abstract:
Large language models (LLMs) play an important role in AI-powered code assistants such as code completion, debugging, and documentation. Such models can be further fine-tuned on smaller amount of data for specific tasks, often with the improvement of performance compared to generic LLMs. However, such fine-tuning techniques are seldomly used in generating sophisticated agent-based models (ABMs), because they are often implemented as software that demands extra standards such as the “Overview, Design concepts, and Details” (ODD) protocol. This research examines how we can bridge this gap by utilizing LLMs in designing or conceptualizing, building, and running agent-based models in the form of user prompts. . In this work, two models are created to demonstrate the proposed method. Specifically, Sakoda’s checkerboard model of social interaction is created by LLM from explicit design and description through prompts. The other model stimulates consumer preferences and restaurant visits as designed and implemented by a LLM. These models are evaluated by human experts on their code correctness and quality for both verification and validation purposes. This work serves as a first step towards fine-tuned LLMs on existing models and documentations to create high-quality and functional ABMs based on either user prompts or standard protocols, contributing to further exploration on the future of AI-assisted geospatial simulation development. 

Keywords: Agent-Based Modeling, Large Language Models, Geospatial Simulation

Reference: 

Jiang, N., Wang, B. and Crooks, A.T. (2025), Large Language Models for Conceptualizing, Designing, and Generating Agent-based Models, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)
Next up was Ying Zhou who presented our work entitled "Identifying Environmental Characteristics That Influence Perceived Safety in Urban Spaces." In this work we explored how using social media data can be used to study the fear and how this relates to actual crimes within New York city. Broadly speaking we find through our analysis, that fear sentiment may spread out between the neighborhoods and their surrounding areas and that neighborhoods surrounded by crime-clusters may have high sentiments of fear. 

Abstract:
One goal of creating livable cities is to enhance health and safety. While previous research in spatial analysis and urban planning has focused on correlations between physical environments and crime, typically relying on police-reported crime data from sources like the Crime Open Database (CODE), safety perception is inherently subjective and cannot be fully represented by objective crime statistics alone. Also, urban planning today has gradually shifted its focus from a top-down mechanism to a bottom-up mechanism, so understanding and fostering spaces where residents feel safe is essential. This research examines factors that contribute to residents’ perceived insecurity in New York City. In addition to spatial analysis of the open crime data, the research used social media data to acquire people’s perceptions. The result indicates that the aggregations of perceived unsafe locations overlapped with aggregations of crime data's locations, such as in Manhattan’s neighborhoods, but they do not overlap with each other entirely. By adopting Latent Dirichlet Allocation (LDA), a method of topic modeling, the research filtered and summarized the posted texts and contents related to the negative descriptions of places or spaces in the city, and then it identified the related characteristics of the environments. The characteristics are investigated by the method of local Moran’s I, which indicates their spatial autocorrelation in some neighborhoods in the city of New York. This research offers “bottom’s views” about urban safety for both urban planning and decision-makers, which contributes to people-centered consideration for future development and urban resource distribution. 

Keywords: Safety, Crime, Urban Space, Livable Cities, Social Media, Spatial Analysis.
Methodology
Reference: 
Zhou, Y. and Crooks, A.T. (2025), Identifying Environmental Characteristics That Influence Perceived Safety in Urban Spaces, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)
The last day of the conference was another busy day with two talks. First was entitled "PySGN: A Python Package for Constructing Synthetic Geo-social Networks" where Boyu Wang presented our work (with Taylor Anderson and Andreas Züfle) on a Python package that can be used to generate synthetic geo-social networks. As readers of this blog might know we have a an interest in social networks and using them in modeling and this package provides a toolkit for others to easily create their own geosocial networks (e.g., Geospatial ErdÅ‘s-Rényi, Barabási–Albert and Watts-Strogatz models). For interested readers, the source code available at: https://github.com/wang-boyu/pysgn.

Abstract:
Synthetic population has been widely used in social simulations such as traffic modeling, pedestrian movements, and the spread of infectious diseases. In recent years, much attention was focused on generating synthetic population with social networks, that captures social connections between individuals. While synthetic populations are often geographically explicit, various algorithms have been proposed to create realistic geographic social (geo-social) networks, aiming to integrate spatial information into people’s social links. We build an open-source Python package, namely PySGN, for constructing synthetic geo-social networks that incorporates position information, exhibits small-world network properties, and can be scaled to hundreds of thousands and potentially millions of nodes. We discuss different ways of parametrizing the method, by either a global average node degree, or an expected degree for each individual node. It is demonstrated through a case study with synthetic population in Buffalo, NY. By doing so, we aim to illustrate how such synthetic geo-social networks can be created, utilized, and analyzed in downstream agent-based modeling and network analysis tasks. This work is available as an open-source Python package and integrated with the PyData ecosystem (e.g., GeoPandas, NetworkX), and can be further extended with more synthetic geo-social network algorithms in the future. 
Keywords: Agent-Based Modeling, Synthetic Geo-Social Network, Python, Open-Source Software
Examples of Geosocial Networks Created in PySGN

Reference: 

Wang, B., Crooks, A.T., Anderson, T. and Züfle, A. (2025), PySGN: A Python Package for Constructing Synthetic Geo-social Networks. The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)
The final talk (well for me) was presented by Fuzin Yin who presented our work with Lucie Laurian  and Emmanuel Frimpong Boamah entitled "Analysis of Online Mutual Aid Network during Buffalo Blizzard 2022: Actors and Weak Ties." In this work we explored what kind of support was offered and requested over Facebook groups along with their network structures durring and shortly after the event utilizing machine learning. 

Abstract: 
In December 2022, Buffalo, NY encountered a once-in-a-generation blizzard that dropped over 4 feet of snow. This four-day snow event halted emergency services and left 47 dead. In the face of the devastating blizzard, Buffalonian demonstrated resilience and solidarity by establishing Facebook (FB) groups to share information and coordinate behaviors including donations, wellness checks, and snow removals. These spontaneous behaviors created an essential layer of protection when the major infrastructure was down. This research has collected data from Buffalo blizzard FB groups to analyze community-led self-help behaviors. We have used machine learning to classify FB messages into four categories (e.g., requesting help, offering help, emotional support, and other), and social network analysis to explore users’ communication patterns. Results show that out of all messages (n=9,988), 37% of them express emotional support, which is followed by messages offering help (25%). While requests for help constitute a small proportion (8%), they stimulate more replies than other categories. Network statistics suggest that the mutual aid network is low-density but with a high clustering coefficient. This implies that most group members are strangers with weak ties, but their connections are in the right place to allow efficient communication. However, users do not equally benefit where people requesting or offering help are central in online conversation while pure emotional supporters are at the periphery. We conclude that during the Buffalo blizzard 2022, online interactions translate into offline mutual assistance by establishing weak ties among disconnected users to facilitate the flow of information and resources. 
Keywords: crisis informatics, mutual aid, social network analysis, machine learning, social media

 

Results of Mutual Aid Network during Buffalo Blizzard 2022

Reference: 

Yin, F., Laurian, L., Crooks, A.T. and Boamah, E.F. (2025), Analysis of Online Mutual Aid Network during Buffalo Blizzard 2022: Actors and Weak Ties, The Association of American Geographers (AAG) Annual Meeting, 24th –28th March, Detroit, MI. (pdf)

While this is a rather longer post than normal, we hope you found it interesting and also as noted at the top of the post, if any of these talks/topics are of interest to you please feel free to reach out. 

Friday, January 31, 2025

New Directions in Mapping the Earth’s Surface with Citizen Science and Generative

In previous posts, we have written how large language models (LLMs) like ChatGPT can be used in various urban analytical applications. We have kept exploring this potential especially with respect to citizen science applications. To this end we have just published a new paper in iScience, entitled "New Directions in Mapping the Earth’s Surface with Citizen Science and Generative AI". In the paper, lead by Linda See, we discuss how multi-modal LLMs (MLLMs) which are like LMMs but can take different forms of inputs (e.g., text, images, video) and output multi-modal information (e.g., take an image and output a description) could be leveraged to enhance citizen science land cover/land use mapping campaigns. If this sounds of interest, below you can read the abstract to the paper, see some of the figures we use to build our argument, while at the bottom of the post you can see the full reference and a link to the actual paper.
Abstract: 
As more satellite imagery has become openly available, efforts in mapping the Earth’s surface have accelerated. Yet the accuracy of these maps is still limited by the lack of in-situ data needed to train machine learning algorithms. Citizen science has proven to be a valuable approach for collecting in-situ data through applications like Geo-Wiki and Picture Pile, but better approaches for optimizing volunteer time are still required. Although machine learning is being used in some citizen science projects, advances in generative Artificial Intelligence (AI) are yet to be fully exploited. This paper discusses how generative AI could be harnessed for land cover/land use mapping by enhancing citizen science approaches with multi-modal large language models (MLLMs), including improvements to the spatial awareness of AI.
Visual interpretation tasks undertaken by ChatGPT for (a) a wetland/mangrove landscape in South America (b) an agricultural area in central Europe.
Visual interpretation tasks undertaken by ChatGPT for identification of natural and non-natural ecosystems where ChatGPT misclassified the images as non-natural for locations in (a) Chad and (b) Austria. In (c), the image from Colombia was classified as unsure by validators but natural by ChatGPT.
Integrating multi-modal Large Language Models (MLLMs) in a citizen science visual interpretation workflow.
Full reference : 
See, L., Chen, Q., Crooks, A., Bayas, J.C.L., Fraisl, D., Fritz, S., Georgieva, I., Hager, G., Hofer, M., and Lesiv, M., Malek, Ž., Milenković, M., Moorthy, I., Orduña-Cabrera, F., Pérez-Guzmán, K., Schepaschenko, D., Shchepashchenko, M., Steinhauser, J.and McCallum, I. (2025), New Directions in Mapping the Earth’s Surface with Citizen Science and Generative AI, iScience, doi: https://doi.org/10.1016/j.isci.2025.111919. (pdf)

Monday, February 19, 2024

Exploring the New Frontier of Information Extraction through Large Language Models in Urban Analytics

Over the last year or so there has been a lot of hype about artificial intelligence (AI) and Large Language Models (LLMs) in particular, such as Generative Pre-trained Transformers (GPT) like ChatGPT. In a recent editorial in Environment and Planning B written by Qingqing Chen and myself we discussed how LLMs could be used for lower the barrier for researchers wishing to study urban problems through the lens of urban analytics. For example, analyzing street view images in the past required training and segmentation of such data which a time consuming and a rather technical task. But what can be done using ChatGPT? To test this we provided ChatGPT some images from Flickr and Mapillary: 

Examples of using ChatGPT for extracting information from imagery.

And then asked it some questions and we were quite amazed by the answers:  

Examples questions and responses when using ChatGPT for extracting information from imagery.

If this sounds of interest I encourage you to read the editorial and think how you could leverage LLMs for your own research. 

Full Reference: 

Crooks A.T. and Chen, Q (2024), Exploring the New Frontier of Information Extraction through Large Language Models in Urban Analytics, Environment and Planning B. Available at https://doi.org/10.1177/23998083241235495. (pdf)