Methodological proposal to identify the nationality of Twitter users through random forests

We disclose a methodology to determine the participants in discussions and their contribu tions in social networks with a local relationship (e.g., nationality), providing certain levels of trust and efficiency in the process. The dynamic is a challenge that has demanded studies and some approximati...

Descripción completa

Autores Principales: Quijano, Damián, Gil Herrera, Richard
Formato: Artículo
Idioma: Inglés
Publicado: Plos One 2023
Materias:
Acceso en línea: http://repositorio2.udelas.ac.pa/handle/123456789/1149
Sumario: We disclose a methodology to determine the participants in discussions and their contribu tions in social networks with a local relationship (e.g., nationality), providing certain levels of trust and efficiency in the process. The dynamic is a challenge that has demanded studies and some approximations to recent solutions. The study addressed the problem of identify ing the nationality of users in the Twitter social network before an opinion request (of a politi cal nature and social participation). The employed methodology classifies, via machine learning, the Twitter users’ nationality to carry out opinion studies in three Central American countries. The Random Forests algorithm is used to generate classification models with small training samples, using exclusively numerical characteristics based on the number of times that different interactions among users occur. When averaging the proportions achieved by inferences of the ratio of nationals of each country, in the initial data, an average of 77.40% was calculated, compared to 91.60% averaged after applying the automatic clas sification model, an average increase of 14.20%. In conclusion, it can be seen that the sug gested set of method provides a reasonable approach and efficiency in the face of opinion problems.