Research Project:
Building the Future: Excelling in Computational and Quantitative Social Sciences in Turkey

Loading...
Project Logo

Contributors

Funders

ID

EC.00135

Authors

Person
Yörük, Erdem
Faculty Member

Publications

Placeholder
Publication
Validating digital traces with survey data: the use case of religiosity
(Association for Computing Machinery, 2024) Atsızelti, Şükrü; Duruşan, Fırat; Etgü, Tolga; Yardı, Melih Can; Yörük, Erdem; Gürerk, Oğuz; Kina, Mehmet Fuat; Nişancı, Zübeyir; Hürriyetoglu, Ali; Akbulut, Yusuf; Bacaksızlar Turbic, Gizem; Department of Sociology; Department of Mathematics; CCSS (Center for Computational Social Sciences); Yes; College of Social Sciences and Humanities; College of Sciences; Research Center
This paper tests the validity of a digital trace database (Politus) obtained from Twitter, with a recently conducted representative social survey, focusing on the use case of religiosity in Turkey. Religiosity scores in the research are extracted using supervised machine learning under the Politus project. The validation analysis depends on two steps. First, we compare the performances of two alternative tweet-To-user transformation strategies, and second, test for the impact of resampling via the MRP technique. Estimates of the Politus are examined at both aggregate and region-level. The results are intriguing for future research on measuring public opinion via social media data.
Thumbnail Image
PublicationOpen Access
A task set proposal for automatic protest information collection across multiple countries
(Springer, 2019) Duruşan, Fırat; Gürel, Burak; Hürriyetoğlu, Ali; Mutlu, Osman; Yoltar, Çağrı; Yörük, Erdem; Yüret, Deniz; Department of Sociology; Department of Computer Engineering; Graduate School of Sciences and Engineering; Graduate School of Social Sciences and Humanities; Yes; College of Engineering; College of Social Sciences and Humanities; GRADUATE SCHOOL OF SCIENCES AND ENGINEERING; GRADUATE SCHOOL OF SOCIAL SCIENCES AND HUMANITIES
We propose a coherent set of tasks for protest information collection in the context of generalizable natural language processing. The tasks are news article classification, event sentence detection, and event extraction. Having tools for collecting event information from data produced in multiple countries enables comparative sociology and politics studies. We have annotated news articles in English from a source and a target country in order to be able to measure the performance of the tools developed using data from one country on data from a different country. Our preliminary experiments have shown that the performance of the tools developed using English texts from India drops to a level that are not usable when they are applied on English texts from China. We think our setting addresses the challenge of building generalizable NLP tools that perform well independent of the source of the text and will accelerate progress in line of developing generalizable NLP systems.
Thumbnail Image
PublicationOpen Access
Overview of CLEF 2019 lab protestnews: extracting protests from news in a cross-context setting
(Springer, 2019) Akdemir, Arda; Gürel, Burak; Hürriyetoğlu, Ali; Mutlu, Osman; Yoltar, Çağrı; Yörük, Erdem; Yüret, Deniz; Department of Sociology; Department of Computer Engineering; Graduate School of Sciences and Engineering; Graduate School of Social Sciences and Humanities; Yes; College of Engineering; College of Social Sciences and Humanities; GRADUATE SCHOOL OF SCIENCES AND ENGINEERING; GRADUATE SCHOOL OF SOCIAL SCIENCES AND HUMANITIES
We present an overview of the CLEF-2019 Lab ProtestNews on Extracting Protests from News in the context of generalizable natural language processing. The lab consists of document, sentence, and token level information classification and extraction tasks that were referred as task 1, task 2, and task 3 respectively in the scope of this lab. The tasks required the participants to identify protest relevant information from English local news at one or more aforementioned levels in a cross-context setting, which is cross-country in the scope of this lab. The training and development data were collected from India and test data was collected from India and China. The lab attracted 58 teams to participate in the lab. 12 and 9 of these teams submitted results and working notes respectively. We have observed neural networks yield the best results and the performance drops significantly for majority of the submissions in the cross-country setting, which is China.

Organizational Units

Description

Keywords