TY - JOUR
T1 - Latent human traits in the language of social media
T2 - An open-vocabulary approach
AU - Kulkarni, Vivek
AU - Kern, Margaret L.
AU - Stillwell, David
AU - Kosinski, Michal
AU - Matz, Sandra
AU - Ungar, Lyle
AU - Skiena, Steven
AU - Schwartz, H. Andrew
N1 - Publisher Copyright:
© 2018 Kulkarni et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
PY - 2018/11
Y1 - 2018/11
N2 - Over the past century, personality theory and research has successfully identified core sets of characteristics that consistently describe and explain fundamental differences in the way people think, feel and behave. Such characteristics were derived through theory, dictionary analyses, and survey research using explicit self-reports. The availability of social media data spanning millions of users now makes it possible to automatically derive characteristics from behavioral data—language use—at large scale. Taking advantage of linguistic information available through Facebook, we study the process of inferring a new set of potential human traits based on unprompted language use. We subject these new traits to a comprehensive set of evaluations and compare them with a popular five factor model of personality. We find that our language-based trait construct is often more generalizable in that it often predicts non-questionnaire-based outcomes better than questionnaire-based traits (e.g. entities someone likes, income and intelligence quotient), while the factors remain nearly as stable as traditional factors. Our approach suggests a value in new constructs of personality derived from everyday human language use.
AB - Over the past century, personality theory and research has successfully identified core sets of characteristics that consistently describe and explain fundamental differences in the way people think, feel and behave. Such characteristics were derived through theory, dictionary analyses, and survey research using explicit self-reports. The availability of social media data spanning millions of users now makes it possible to automatically derive characteristics from behavioral data—language use—at large scale. Taking advantage of linguistic information available through Facebook, we study the process of inferring a new set of potential human traits based on unprompted language use. We subject these new traits to a comprehensive set of evaluations and compare them with a popular five factor model of personality. We find that our language-based trait construct is often more generalizable in that it often predicts non-questionnaire-based outcomes better than questionnaire-based traits (e.g. entities someone likes, income and intelligence quotient), while the factors remain nearly as stable as traditional factors. Our approach suggests a value in new constructs of personality derived from everyday human language use.
UR - https://www.scopus.com/pages/publications/85057458065
U2 - 10.1371/journal.pone.0201703
DO - 10.1371/journal.pone.0201703
M3 - Article
C2 - 30485276
AN - SCOPUS:85057458065
SN - 1932-6203
VL - 13
JO - PLoS ONE
JF - PLoS ONE
IS - 11
M1 - e0201703
ER -