BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//talks.osgeo.org//foss4g-2023-academic-track//talk//SW3M3
 V
BEGIN:VTIMEZONE
TZID:CET
BEGIN:STANDARD
DTSTART:20001029T040000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10
TZNAME:CET
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20000326T030000
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=3
TZNAME:CEST
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
UID:pretalx-foss4g-2023-academic-track-SW3M3V@talks.osgeo.org
DTSTART;TZID=CET:20230630T143000
DTEND;TZID=CET:20230630T150000
DESCRIPTION:### Introduction\nThe analysis of georeferenced social media (S
 M) data holds broad potential for informing municipal policy-making. Local
  adaptation to climate change and disaster resilience\, transforming city 
 centers\, gentrification\, and demographic change are significant challeng
 es for municipalities. \nIn light of these pressing topics\, a growing awa
 reness for data-driven decision making has fostered geospatial interfaces 
 that allow practitioners to interactively explore data source. \nParticula
 rly SM offers the potential of a live feed and continuous reflection of ev
 ents at scale. Although many studies have an urgent need for a purpose-dri
 ven\, customized visualization of spatial data\, little emphasis has been 
 put on how to display these data.\nMany studies on map-based visualization
  in SM use traditional cartographic methods\, such as pins or choropleth m
 aps\, with varying color scales or heatmaps to represent absolute or relat
 ive values. However\, SM data presents challenges that require more sophis
 ticated statistical metrics and flexible visualization techniques. We asse
 ss the signed chi metric\, specifically designed for mapping via binning\,
  and expand its use in a Bonn case study using an on-the-fly hexagonal bin
 ning method for frontend applications like dashboards. We then evaluate th
 e advantages and disadvantages of the various proposed metrics and visuali
 zations in terms of their practical applications.\n\n### Problem Statement
 \nAs the overview by (Teles da Mota & Pickering 2020) has shown\, research
  involving geo-SM from different platforms has become increasingly popular
  but bears specific problems inherent to the characteristics of volunteere
 d geographic information (VGI) – volume\, veracity\, velocity\, variety 
 are just broad categories used to characterize these.\n\nFirstly\, Access 
 to SM databases\, such as Meta or Twitter\, is usually limited to capital 
 intensive partner companies. Instagram's public-facing API is largely undo
 cumented and opaque to end-users\, causing uncertainty about data selectio
 n criteria (Dunkel 2023). Hence\, the lack of knowledge about data context
  and possible biases can affect the representativeness of the data subset.
  \n\nSecond\, "super users" sharing repeated content may create noise and 
 skew analysis outcomes if absolute values are solely considered.\n\nThird\
 , as Teles da Mota & Pickering (2020) point out\, research has been conduc
 ted mainly for large areas ranging from national parks to entire countries
  or seldomly even the whole world (cf. Dunkel et al. 2023). Studies workin
 g with data on the municipal level where individual locations and differen
 ces of only a few meters play a significant role\, are usually not focusin
 g on methodological cartographic issues or appropriate metrics but rather 
 on effectively communicating core research results. Due to this lack of re
 ference material for the municipal level\, a research gap of proper visual
 ization methods is identified.\n\nLastly\, VGI\, as practiced by Instagram
 \, poses a unique problem for researchers. Users are allowed to create pub
 lic "Instagram Locations" and tag their posts with a coordinate of their c
 hoice\, which can then be referenced by other users as well. However\, the
  user is not obligated to provide a clear definition of what exactly is me
 ant by the location they choose\, creating ambiguity. For instance\, the [
 "Bonn" location](https://www.instagram.com/explore/locations/107481562)'s 
 coordinates (50.7333\, 7.1) are situated in the city's center. What it act
 ually refers to is entirely subject to the interpretion of the user. It co
 uld refer to different extents of the city center\, the official administr
 ative boundaries of Bonn or anything loosely associated with Bonn\, includ
 ing cultural references or events. This ambiguity which Meta is aware of (
 Delvi et al. 2014) can be observed on different zoom levels such as city d
 istricts\, cities\, countries or continents throughout Instagram data and 
 poses an enormous challenge to researchers working with city-scale areas o
 f interest.\n\n### Research Interest\nIn order to deal with these challeng
 es\, a thorough data cleaning is insufficient. We propose an application-o
 riented system of metrics for data processing and visualization depending 
 on the user’s needs\, by comparing possible application scenarios as wel
 l as limitations based on a case study for the city of Bonn with Instagram
  data from 2010 - 2022:\n1. Absolute values – absolute number of observe
 d posts per location or bin\n2. Relative values – relation between obser
 ved and expected posts per location or bin\n3. Signed chi – statistic va
 lue indicating significance and direction per location or bin\n\nThe *obse
 rved* value usually refers to a quantity found at a specific bin\, using a
  specific query such as a thematic filter. In contrast\, the *expected* va
 lue often refers to an average quantity of a generic query\, such as the a
 verage of all SM posts in Bonn\, and it is used to identify over- or under
 represented spatial patterns at local bins (Visvalingam 1978). However\, w
 hat is considered as the observed value for normalization is up to the ana
 lyst (Wood et al. 2007). One could also compare average thematic posts in 
 all German cities (the expected value) to those found in Bonn\, as a means
  to concentrate on the difference of the subject under analysts (posts in 
 the city of Bonn). Or\, another option could be to use discrete periods of
  historical time intervals as the expected value\, and compare to the rece
 nt posts quantities to identify recent and unusual spatial posting behavio
 r trends. \n\nWe evaluate these metrics through a hexagonal on-the-fly bin
 ning approach with different color scaling and propose easily customizable
  scripts for the [leaflet-d3 plugin](https://github.com/bluehalo/leaflet-d
 3). We provide all our scripts for reproduction with explanations and usag
 e recommendations as well as a demo dashboard in a public GitHub repositor
 y.\n\nOur findings suggest that all of the investigated metrics can offer 
 insight into data\, but their appropriate use highly depends on the resear
 ch question at hand. When using the dashboard frontend\, outliers should b
 e highlighted\, non-significant values reduced in opacity\, or intra-datas
 et validations being carried out through automatic comparisons across metr
 ics and filters. Overall\, the absolute metric is to be used sparingly. Th
 e relative metric generates only a very narrow gain in knowledge whereas t
 he signed chi metric yields the best overall results and deals very well w
 ith the above issues.
DTSTAMP:20260714T132857Z
LOCATION:UBT E / N209 - Floor 3
SUMMARY:An application-oriented implementation of hexagonal on-the-fly binn
 ing metrics for city-scale georeferenced social media data - Dominik Weckm
 üller
URL:https://talks.osgeo.org/foss4g-2023-academic-track/talk/SW3M3V/
END:VEVENT
END:VCALENDAR
