DATA STORYTELLING TOOLKIT

Data Storytelling for Digital Research Infrastructure Project

Igor Tkalec, Social Data Institute (SODA), UCL
Reem Khurshid, Geographic Data Service (GeoDS), UCL


The project

Data Storytelling for Digital Research Infrastructure is a UKRI-funded project (September 2024–March 2027) that addresses skills as one of the strategic priorities of UKRI Digital Research Infrastructure (DRI). Through in-person workshops conducted across the UK on Concept and Practice of Data Storytelling and Python for Data Storytelling, the project equips (early career) researchers with skills in data storytelling communication and data visualisation.

In brief, data stories enmesh narratives, data and supporting visuals in order to inspire and creatively underpin a call for action by the story’s targeted audience. The power and impact of data stories draw on the general tendency and preference of people to understand and process information through stories/narratives — this effectively supplements and enhances prevalent and conventional ways to present data insight rooted in numbers and statistics.

The toolkit

Note: The toolkit is a work in progress undertaking and may be periodically updated and / or enriched. It is vital to note that the toolkit is a not the definitive guide to data storytelling.

This toolkit is an open-access supplementary learning material that can serve as a companion for ongoing data storytelling efforts for researchers who have experience with data storytelling and / or for those who attended the aforementioned workshops. The toolkit outlines and exemplifies main building blocks for data storytelling including:
  [1] creating an audience profile for a data story
  [2] main elements of a data story
  [3] usefulness of outliers for inspiring/crafting data stories

Parts [1] and [2] of the toolkit are based on Dykes (2020).

In demonstrating the main elements of a data story, the toolkit utilises a data story example from the Data After Dark project led by UCL Social Data Institute and UCL Urban Laboratory. The project was supported by the Mayor of London. It received support, advice and funding from UCL Innovation & Enterprise. Driven by both quantitative and qualitative data and evidence, the project sheds new light on night workers’ experiences and working conditions, providing actionable insights to support policymaking.


Click on the icons/boxes to learn more and explore the example.

[1] Know your audience…

Audience profile traits to consider when developing a data story (ideally prior to writing it).

High familiarity: data story may demonstrate an immediate focus on explanation of data insight.

Low familiarity: in explanation of the data insight, data story ought to provide more context and conceptual information to underpin the explanation.

It is also important to consider the level of audience’s data literacy and adjust the explanation of data insight accordingly. The level of data literacy may also affect the depth of data-technical discussion as well as the vocabulary used throughout the story.

A storyteller may pay attention to the alignment between the explanation of a data insight and what matters to the audience — what are the their general preferences; what are their preferences in regard to the issue / topic? Have they taken any (recent) action in relation to the issue / topic?

Anticipate what the the audience may expect from you as a storyteller, and what they may expect from engaging with the issue / problem in question that you are addressing. Arguably, data stories may have a greater impact on the audience (in terms of call for action and change) if they reveal an unexpected insight deriving from the data.

Considerations around audiences’ preferences and expectations can help steer the narrative and explanative focus of a data story.

It is helpful for a storyteller to recognise best timing (in terms of the audience’s agenda) to argue for the main point conveyed through a data story.

Therefore, becoming familiar (to the greatest feasible extent) with orgnisational or institutional contexts (e.g., policy cycles) and contemporary external factors (e.g., elections, crises, relevant and / or recent incidents, upcoming events) around the audience is helpful.

This may boost general momentum for the issue / topic, relevance, persuasion potential and thus the impact of the data story.

When thinking about and developing a data story, a storyteller may consider audience’s organisation/institution.

Here, factors such as organisational legacy, mission statement, core values, short- and long-term objectives, and insight from official reports may prove helpful and further enhance the audience profile. Drawing on past experiences / engagement with the organisation or institution, personal networks, and trust-building could be useful.

This may also help steer the overall narrative, and the way a storyteller communicates and explains the main data insight, to achieve greater impact on the audience.


[2] Start crafting a data story…

The six elements of a data story (n.b. not steps in its development) (Dykes, 2020) are outlined and demonstrated through the data story output from the Data After Dark project.

The data foundation is the factual, credible and truthful base of a data story. Any type of data (quantitative, qualitative) could be used to underpin a data story. The data foundation can be relatively simple (e.g., a few descriptive observations from the data) or relatively complex (e.g., results / insight from statistical or machine learning modelling).

The main message is the central insight (rooted in data) of a data story, which entails a message the storyteller conveys / communicates to the audience. It ensures that a data story has a clear purpose, and serves as the basis for a call for action that may inspire change (once the audience acts upon it). The story may be more impactful if the insight is unexpected / counter-intuitive (from the audience’s perspective).

With a data story, a storyteller aims to explain (rather than just describe) the central insight. Put differently, we aim to fully communicate a message to the audience rather than only inform them of it.

The sequencing of information is the way that evidence that supports and explains the central insight is outlined and organised. Here it may be effective to undertake a gradual, build-up approach when presenting evidence — e.g., starting from the most pertinent piece of evidence to the least pertinent (or vice versa).

In order to capture the audience’s attention and increase their engagement with the data story, a storyteller may use dramatic elements such as plot development, personalisation, development of characters (protagonist-antagonist). This will help communicate urgency around the topic or generate an emotional appeal, and may thus amplify the story’s impact.

Visuals help enlighten the data and engage or entertain the audience. They reinforce and / or illustrate what a storyteller argues in the narrative. Effective and powerful visuals are usually simple and intuitive to read and understand. Each data visual may ideally portray and convey a message to your audience.


[3] Making use of outliers…

Outliers / anomalies are data observations that deviate from other observations (which are within the boundary of what is defined as 'normality'). This is a brief guide for outlier detection. Outliers can sometimes serve as starting points for data stories. They are helpful in the development of data stories as they, by default, represent something unexpected, unconventional and potentially intriguing — this resonates with the main rationale for writing data stories. Outliers may thus be deemed as the core of a 'disbalance' (if we assume that non-deviant observations represent a 'balance') that may provide grounding to the central data insight and foster audience engagement. After detecting outlier(s), a storyteller ought to conduct an in-depth inspection / investigation of the outlier (e.g., root and / or context of deviancy, expansion of the data foundation / evidence, finding an angle for a narrative, etc.). Following the logic of the Start crafting a data story elements can be helpful for an in-depth outlier inspection. The below draws on Lysy & Crain (2016).

When searching for outliers, stanard tabular datasets may be explored across rows (e.g., individual variables [univariate setting]); across columns (e.g., values over time for a single country); or in a bi/multivariate setting (e.g., correlations between variables).

To find an outlier, a storyteller ought to define boundaries for “normality”. Observations that are not within these boundaries may be considered as outliers. Visual inspections (histograms, boxplots, scatterplots) together with statistical criteria (e.g., Z-scores, standard deviations, interquartile range, outlier tests) may be helpful for defining the boundaries for 'normality'.

A) Point (global) outlier: a single obs. considered as an anomaly in relation to the rest of obs.

B) Contextual outlier: an obs. is considered an anomaly in a specific context and not otherwise.

C) Collective outlier: a collection of obs. which are individually not anomalies are considered as anomalies in relation to the rest of obs.

References

Dykes, B. (2020). Effective data storytelling: How to drive change with data, narrative and visuals. Hoboken: John Wiley and Sons, Inc.

Lysy, C., & Crain, D. (2016). Outlier Analyses: Step-by-Step Guide (Version 1.0). IDEA DataCenter.

Singh, K., & Upadhyaya, D. S. (2012). Outlier Detection: Applications And Techniques. International Journal of Computer Science Issues, 9(1), 301–323.

Further reading:

Nussbaumer Knaflic, C. (2015). Storytelling with Data. Hoboken: John Wiley & Sons, Inc.

NussbaumerKnaflic, C. (2020). Storytelling with data: Let’s practice! Hoboken: John Wiley & Sons, Inc.