Cross-modal Transfer Between Vision and Language for Protest Detection

Ria Dass Raj; Kajsa Andreasson; Tobias Norlund; Richard Johansson; Aron Lagerberg

doi:10.18653/v1/2022.case-1.8

Cross-modal Transfer Between Vision and Language for Protest Detection
Paper in proceeding, 2022

Most of today’s systems for socio-political event detection are text-based, while an increasing amount of information published on the web is multi-modal. We seek to bridge this gap by proposing a method that utilizes existing annotated unimodal data to perform event detection in another data modality, zero-shot. Specifically, we focus on protest detection in text and images, and show that a pretrained vision-and-language alignment model (CLIP) can be leveraged towards this end. In particular, our results suggest that annotated protest text data can act supplementarily for detecting protests in images, but significant transfer is demonstrated in the opposite direction as well.

Author

Ria Dass Raj

Student at Chalmers

Recorded Future

Kajsa Andreasson

Recorded Future

Student at Chalmers

Tobias Norlund

Recorded Future

Chalmers, Computer Science and Engineering (Chalmers), Data Science and AI

Other publications Research

Richard Johansson

University of Gothenburg

Other publications Research

Aron Lagerberg

Recorded Future

Other publications Research

CASE 2022 - 5th Workshop on Challenges and Applications of Automated Extraction of Socio-Political Events from Text, Proceedings of the Workshop

56-60
978-1-959429-05-0 (ISBN)

5th Workshop on Challenges and Applications of Automated Extraction of Socio-political Events from Text (CASE)
Abu Dhabi, United Arab Emirates,

Subject Categories (SSIF 2011)

Other Computer and Information Science

Language Technology (Computational Linguistics)

Computer Vision and Robotics (Autonomous Systems)

DOI

10.18653/v1/2022.case-1.8

Publication data connected to DOI

More information

Latest update

6/24/2026

Cross-modal Transfer Between Vision and Language for Protest Detection Paper in proceeding, 2022

Author

Ria Dass Raj

Kajsa Andreasson

Tobias Norlund

Richard Johansson

Aron Lagerberg

CASE 2022 - 5th Workshop on Challenges and Applications of Automated Extraction of Socio-Political Events from Text, Proceedings of the Workshop

Subject Categories (SSIF 2011)

DOI

More information

Latest update

Cross-modal Transfer Between Vision and Language for Protest Detection
Paper in proceeding, 2022