ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

A standardized personality lexicon for enhancing personalized human-machine interaction.

Jin T et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cognitive psychology

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Sci Data . 2026 Mar 3;13:579. doi: 10.1038/s41597-026-06783-6 Search in PMC Search in PubMed View in NLM Catalog Add to search A standardized personality lexicon for enhancing personalized human-machine interaction Tao Jin Tao Jin 1 West China Biomedical Big Data Center, West China Hospital of Sichuan University, Chengdu, China Find articles by Tao Jin 1 , Hui Cai Hui Cai 2 College of Mathematics, Sichuan University, Chengdu, China Find articles by Hui Cai 2 , Xinyi Shi Xinyi Shi 3 Department of Population Health, New York University Grossman School of Medicine, New York, USA Find articles by Xinyi Shi 3 , Xiaomin Kou Xiaomin Kou 4 Mental Health Center, National Medical Center for Mental Diseases, West China Hospital of Sichuan University, Chengdu, China Find articles by Xiaomin Kou 4 , Xialian Hu Xialian Hu 5 West China School of Medicine, Sichuan University, Chengdu, China Find articles by Xialian Hu 5 , Hua Zhong Hua Zhong 6 West China School of Nursing, Sichuan University, Chengdu, China Find articles by Hua Zhong 6 , Yan Yang Yan Yang 1 West China Biomedical Big Data Center, West China Hospital of Sichuan University, Chengdu, China Find articles by Yan Yang 1 , Jingwen Jiang Jingwen Jiang 1 West China Biomedical Big Data Center, West China Hospital of Sichuan University, Chengdu, China Find articles by Jingwen Jiang 1 , Yuchen Li Yuchen Li 4 Mental Health Center, National Medical Center for Mental Diseases, West China Hospital of Sichuan University, Chengdu, China Find articles by Yuchen Li 4, # , Wei Zhang Wei Zhang 1 West China Biomedical Big Data Center, West China Hospital of Sichuan University, Chengdu, China Find articles by Wei Zhang 1, ✉, # Author information Article notes Copyright and License information 1 West China Biomedical Big Data Center, West China Hospital of Sichuan University, Chengdu, China 2 College of Mathematics, Sichuan University, Chengdu, China 3 Department of Population Health, New York University Grossman School of Medicine, New York, USA 4 Mental Health Center, National Medical Center for Mental Diseases, West China Hospital of Sichuan University, Chengdu, China 5 West China School of Medicine, Sichuan University, Chengdu, China 6 West China School of Nursing, Sichuan University, Chengdu, China ✉ Corresponding author. # Contributed equally. Received 2025 Jul 31; Accepted 2026 Jan 31; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13065751  PMID: 41775737 Abstract Personality, as a stable and coherent set of behavioral and cognitive patterns, significantly influences linguistic expression, emotional regulation, and cognitive functioning. The Big Five personality traits—neuroticism, extraversion, openness, agreeableness, and conscientiousness— are especially relevant for understanding to language use and social interaction, making them foundational for developing of personality-informed natural language processing (NLP) systems. Despite this, existing personality lexicons often lacks rigorous validation, show weak alignment with linguistic features and personality traits, and fail to adapt to dynamic language environments such as social media. This study presents the construction and empirical validation of a personality lexicon derived from established psychological scales, dictionaries, and literature. Validation using real-world participant data yielded high hit rates across all Big Five dimensions (all > 0.70; mean = 0.787) and their 30 corresponding facets (all > 0.60; mean = 0.768). This lexicon provides a robust foundation for advancing computational personality assessment and supports applications in personalized NLP, large language models, and mental health prediction. Subject terms: Psychology, Communication Background & Summary Personality is a relatively stable and coherent set of patterns of thoughts and behaviors shaped by the interaction of genetics and environment factors 1 . Over the past century, it has been extensively studied, leading to the development of numerous theories and models. Among these, the Big Five Personality traits model, developed from the lexical hypothesis, has emerged as one of the most robust and widely accepted frameworks. This model has five core traits: neuroticism, extraversion, openness, agreeableness, and conscientiousness. These traits have been linked to linguistic abilities, emotional regulation, social behavior, and, notably, cognitive functions 1 – 5 , particularly information processing and logical reasoning 6 , 7 . Therefore, understanding personality is critical to advancing automatic personality recognition technologies, and enhancing the accuracy of interpreting individual differences in language use, behavior, cognition, and emotion in natural language processing 8 . For example, embedding personality-based corpora into chatbot systems enables the dynamic analysis of users’ inputs—including their thoughts, emotions, and preferences—allowing the system to adapt its responses accordingly 9 . This development demonstrates the importance of a personality-informed model in intelligent human-robot interaction. Notably, the lexical hypothesis posits that personality differences are encoded in language 10 , establishing a theoretical bridge between linguistic patterns and trait psychology. However, even when individuals from diverse cultural backgrounds achieve similar scores on standardized personality tests, their personality traits may still exhibit significant variations in actual behavior and linguistic expression 11 . Existing research indicates that the newly added dimension of the HEXACO model—Honesty-Humility—better captures personality traits relevant to specific cultural contexts, reflecting the cross-cultural diversity of personality theories 12 . As the primary medium for personality expression, language reveals how identical traits are encoded and presented across cultures 13 . Therefore, it is necessary to construct a cross-cultural personality lexicon, particularly within the Chinese context. In this way, building personality lexicons with varying levels of granularity is valuable for enhancing the performance of large language models and advancing personalized human-computer interaction. While prior studies confirm correlations between LIWC categories and personality traits 14 , their explanatory power remains limited (5.1–38.5% variance explained). Current Chinese lexicons like Wang et al .‘s dataset 15 suffer from unvalidated annotations (F1 = 0.33 for facets), and social-media-adaptive tools like TextMind 16 lack scalability. To address these gaps, this study: integrates multi-source seed words from standardized scales (e.g., IPIP-NEO), lexical databases (e.g., LIWC), and literature related to personality-related word; implements two-stage validation with real-world participant ratings; delivers a granularly annotated lexicon covering 5 dimensions and 30 facets. In our research, we developed a reliable large, high-quality personality lexicon that covers both coarse- and fine-grained personality categories and can be used for comprehensive personality perception tasks. We also detailed a systematic construction and validation process, demonstrating its robustness across the five dimensions and 30 facets using real participant data. This lexicon enables three novel applications: (1) Dynamic personality adaptation in chatbots using valence-facet mapping 17 ; (2) Mental health screening via social media text mining 18 (e.g., detecting high Neuroticism markers); (3) LLM personalization through trait-aligned prompt engineering 19 . Methods Seed word extraction Scale-derived words: selected from validated instruments (e.g., NEO-PI-R, BFI-2) based on Cronbach’s α > 0.85. The seed words in this study primarily came from three areas: personality scales, word dictionaries, and research literature, all selected through a comprehensive literature review. For personality scales, relevant studies were identified through searches of PsycINFO, PubMed, Web of Science, Scopus, and JSTOR, using the following keywords: (“Personality Traits” OR “Big Five” OR “NEO” OR “Five Factor Model” OR “neuroticism” OR “conscientiousness” OR “openness” OR “extraversion” OR “agreeableness”) AND (“questionnaire” OR “personality inventory” OR “scale” OR “personality test”). Scales included in the final selection were well-established, widely used, and demonstrated validity and reliability. Lexicon sources: for word dictionaries or lexicon, searches were conducted using the keywords: (“Personality Traits” OR “Big Five” OR “NEO” OR “Five Factor Model” OR “neuroticism” OR “conscientiousness” OR “openness” OR “extraversion” OR “agreeableness”) AND (“word dictionary” OR “word lexicon” OR “word library” OR “linguistic inquiry” OR “LIWC”). Only dictionaries that were highly cited and commonly used in personality research were included. Literature sources: for research literature, search terms included “Big Five Personality Traits”, “Big Five”, “NEO”, “Five Factor Model”, “neuroticism”, “conscientiousness”, “openness”, “extraversion”, “agreeableness”, “behavior*“, “language*“, and “textual*“. Studies were screened based on the following inclusion criteria: (1) published within the past 5–10 years; (2) to ensure the quality and credibility of the research literature, English-language publications with at least 50 citations at the time when search was performed (2024/3-2024/6) or published in high-quality, influential, and peer-reviewed journals (a list of eligible journals is provided in Supplementary Materials 1 ); (3) Chinese-language studies published in core Chinese academic journals. Exclusion criteria were: (1) inaccessible the full text, such as databases that are not open to the public; (2) studies that did not report words or linguistic features related to personality traits. Full texts of the selected literature were reviewed by trained psychology researchers, and personality-related terms with statistically significant associations (P < 0.05) were extracted. Seed word extraction results Scales Seed words were first extracted from widely used personality assessment instruments. These included the NEO series (Neuroticism-Extraversion-Openness Personality Inventory) 20 and the BFI (Big Five Inventory) series, which are standard tools for assessing the five major personality traits: Neuroticism, Extraversion, Openness, Agreeableness, and Conscientiousness. The following scales were utilized in this study to extract relevant words: IPIP-NEO-120: This is a shortened version of the NEO scale, containing 120 items, making it suitable for quick assessments. Each personality dimension is measured by 24 items, effectively capturing an individual’s personality traits 21 , 22 . IPIP-NEO-300: This version includes 300 items, providing a more detailed and comprehensive personality assessment. Each dimension is measured by 60 items, making it ideal for in-depth psychological research and personality analysis 23 . NEO-PI-R: The most comprehensive version of the NEO scale, with 240 items, evaluates the five major personality traits and their six sub-dimensions. This detailed breakdown aids in a deeper understanding of individual differences 24 – 27 . The BFI (Big Five Inventory) series also evaluates the five major personality dimensions but offers a more concise set of options compared to the NEO series. This study primarily used the BFI-44 scale, which contains 44 items, with approximately 9 items assessing each personality trait. It is the most widely used version of the BFI series 28 . FFMRF (Five-Factor Model Rating Form): Developed by Thomas A. Widiger and other psychologists in the late 1990s, this tool is based on the Five-Factor Model of personality. It uses simple adjectives or phrases to assess the five major personality traits. Due to its high reliability and validity, it has been widely applied across various cultures and languages, especially for quick screening and preliminary personality assessments in large-scale surveys or online research 29 – 31 . CBF-PI (California Brief Personality Inventory): Developed by a psychology research team from the University of California in the early 2000s, this inventory provides a concise and rapid assessment of the five major personality traits. It has demonstrated good internal consistency and construct validity in multiple studies, simplifying traditional long-form assessments and making personality evaluations more efficient 32 . HEXACO: Created by Michael C. Ashton and Kibeom Lee in 2001, the HEXACO personality model offers a six-factor framework, which includes Honesty-Humility, Emotionality, Extraversion, Agreeableness, Conscientiousness, and Openness to Experience. This model extends the traditional Five-Factor Model by adding the Honesty-Humility dimension and has been validated in numerous cross-cultural studies, showing high reliability and validity 33 – 35 . Chinese Adjective Big Five Personality Inventory: Developed in the late 1990s by a team of Chinese psychologists, this inventory is tailored to the Chinese linguistic and cultural context. It uses a series of adjectives to assess the Big Five personality traits and has demonstrated good psychometric properties, making it suitable for personality research within the Chinese cultural setting 36 . Cattell’s 16 Personality Factors (16PF): Developed by psychologist Raymond Cattell, the 16PF scale assesses 16 distinct personality factors, offering a comprehensive yet intricate personality description. This tool is widely used in career planning and psychological evaluations 37 , 38 . Eysenck Personality Questionnaire (EPQ): Created by Hans Eysenck in 1970, the EPQ focuses on assessing three major dimensions: Neuroticism, Extraversion, and Psychoticism. It provides insight into how individuals interact with the external world and their emotional stability. The EPQ has been widely applied and validated globally, demonstrating high reliability and validity 39 – 41 . Representative seed words extracted from these instruments are listed in Table 1 . Table 1. Examples of scale seed word extraction. Source Scale Words Original Item Description IPIP-NEO-120 Anxious Worries about many things IPIP-NEO-300 Friendly Easily makes friends BFI-44 Energetic Is full of energy FFMRF Five-Factor Word List Angry, distressed Angry hostility CBF-PI Persistent, determined, persevering Once I set a goal, I work hard to achieve it HEXACO Uninterested in aesthetics, indifferent I find visiting art galleries boring. Chinese Adjective Big Five Personality Scale Silent Silent Cattell 16 Personality Factors (16PF) Boring, humorous I am not good at telling jokes or funny stories Eysenck Personality Questionnaire (EPQ) Vulnerable Are your feelings easily hurt? Open in a new tab Lexicon databases Seed words were also drawn from two well-established lexical resources: the LIWC lexicon and the 2818 Personality Trait Descriptor Lexicon. The LIWC (Linguistic Inquiry and Word Count) lexicon, a widely recognized tool in psycholinguistic research, categorizes words into psychologically meaningful domains such as emotion, cognition, and social interaction 14 , 42 , 43 . From this lexicon, words with strong personality relevance were identified with the input of professional psychologists. The 2818 Personality Trait Descriptor Lexicon includes adjectives commonly used to describe personality traits, each evaluated for familiarity, frequency, and relevance in prior psychological studies. Its structured, open-access format provides a robust foundation for personality-related lexical analysis 44 . Examples from these databases are shown in Table 2 . Table 2. Example of dictionary seed word extraction. Words Source Dictionary/Corpus Dimensions Unbelievable LIWC Function; Negative emotion; Affect Unpopular LIWC Cognitive process; Discrepancy Indecisive LIWC Affect; Negative emotion; Anxiety Trust LIWC Affect; Positive emotion Adventurous LIWC Affect; Positive emotion Smart and Capable Personality trait descriptor / Abnormal Personality trait descriptor / Absent-minded Personality trait descriptor / Restrained Personality trait descriptor / Abstract Personality trait descriptor / Open in a new tab Research literature Seed words were extracted from published research literature 14 , 45 – 60 . Examples of these extracted seed words are shown in Table 3 . Table 3. Example of seed word extraction from research literature. Words Source Literature Gambling Brunborg, G. S., Hanss, D., Mentzoni, R. A., Molde, H. & Pallesen, S. Problem gambling and the five-factor model of personality: a large population-based study. Addiction 111, 1428–1435, 10.1111/add.13388 (2016). Depression Kern, M. L., Eichstaedt, J. C., Schwartz, H. A., Dziurzynski, L., Ungar, L. H., Stillwell, D. J., Kosinski, M., Ramones, S. M. & Seligman, M. E. The online social self: an open vocabulary approach to personality. Assessment 21, 158–169, 10.1177/1073191113514104 (2014). Perfectionism Burcaş, S. & Creţu, R. Z. Perfectionism and neuroticism: Evidence for a common genetic and environmental etiology. J. Pers. 89, 819–830, 10.1111/jopy.12617 (2021). Low Pressure Kern, M. L., Eichstaedt, J. C., Schwartz, H. A., Dziurzynski, L., Ungar, L. H., Stillwell, D. J., Kosinski, M., Ramones, S. M. & Seligman, M. E. The online social self: an open vocabulary approach to personality. Assessment 21, 158–169, 10.1177/1073191113514104 (2014). Contemplative Chapman, B. P. & Goldberg, L. R. Act-frequency signatures of the Big Five. Pers. Individ. Differ. 116, 201–205, 10.1016/j.paid.2017.04.049 (2017). Excellent Eichstaedt, J. C., Kern, M. L., Yaden, D. B., Schwartz, H. A., Giorgi, S., Park, G. & Ungar, L. H. Closed- and open-vocabulary approaches to text analysis: a review, quantitative comparison, and recommendations. Psychol. Methods 26, 677–698, 10.1037/met0000349 (2021). Health Giorgi, S., Le Nguyen, K., Eichstaedt, J. C., Kern, M. L., Yaden, D. B., Kosinski, M., Seligman, M. E. P., Ungar, L. H., Schwartz, H. A. & Park, G. Regional personality assessment through social media language. J. Pers. 90, 405–425, 10.1111/jopy.12674 (2022). Loneliness Kern, M. L., Eichstaedt, J. C., Schwartz, H. A., Dziurzynski, L., Ungar, L. H., Stillwell, D. J., Kosinski, M., Ramones, S. M. & Seligman, M. E. The online social self: an open vocabulary approach to personality. Assessment 21, 158–169, 10.1177/1073191113514104 (2014). Stifling Kern, M. L., Eichstaedt, J. C., Schwartz, H. A., Dziurzynski, L., Ungar, L. H., Stillwell, D. J., Kosinski, M., Ramones, S. M. & Seligman, M. E. The online social self: an open vocabulary approach to personality. Assessment 21, 158–169, 10.1177/1073191113514104 (2014). Prayer Koutsoumpis, A., Oostrom, J. K., Holtrop, D., van Breda, W., Ghassemi, S. & de Vries, R. E. The kernel of truth in text-based personality assessment: A meta-analysis of the relations between the Big Five and the Linguistic Inquiry and Word Count (LIWC). Psychol. Bull. 148, 843–868, 10.1037/bul0000381 (2022). Open in a new tab Lexicon annotation Seed words from Chinese sources were retained in their original form and English terms were translated according to the Oxford Tenth Edition Dictionary. Labeling was performed in two stages. Trait mapping: Two psychological researchers independently mapped words to dimensions/facets using IPIP-NEO-120 facet definitions (Table S1 (see Supplementary Information document)); five dimensions, each comprising six facets) 61 , 62 , as specified by the original scales or, for lexicon and literature sources, based on those reported in the corresponding studies. Valence coding: each term was annotated for its emotional valence—classified as positive, negative, or neutral. Emotional valence was defined as the affective sentiment associated with a given word. A positive valence indicated that the word conveyed a favorable emotional sentiment, a negative valence reflected an unfavorable sentiment, and a neutral valence indicated the absence of strong emotional sentiment. Lexicon annotation results After consolidating the collected words and removing duplicates, a total of 6084 seed words were identified. Figure 1 shows a word cloud illustrating representative descriptors for each of the five Big Five personality dimensions. Fig. 1. Open in a new tab Examples of personality word in the five dimensions. The detailed counts for each dimension and facet are provided in Table 4 . Examples of lexicon annotations are shown in Table 5 . Table 4. Summary of counts of word annotation results. Dimension Dimensions Word Count Facet Facets Word Count Neuroticism 1158 Anxiety 198 Anger 232 Depression 155 Self-conscientiousness 254 Immoderation 123 Vulnerability 196 Extraversion 1674 Friendliness 463 Gregariousness 198 Assertiveness 311 Activity Level 246 Excitement-seeking 181 Cheerfulness 274 Openness 1187 Imagination 129 Artistic interests 108 Emotionality 148 Adventurousness 193 Intellect 281 Liberalism 329 Agreeableness 1764 Trust 117 Morality 259 Altruism 299 Cooperation 437 Modesty 216 Sympathy 436 Conscientiousness 1332 Self-efficacy 164 Orderliness 200 Dutifulness 175 Achievement striving 215 Self-discipline 264 Cautiousness 314 Open in a new tab Table 5. Example of seed word annotation. Words Source Dimensions Facets Emotional valence Creative BFI-44 Openness Ideas Positive valence Steady FFMRF Neuroticism Anger and Hostility Positive valence Pessimistic FFMRF Neuroticism Depression Negative valence Vulgar HEXACO Openness Art Negative valence Meddlesome LIWC Conscientiousness Self-discipline Negative valence Clumsy EPQ Agreeableness Compliance Negative valence Open in a new tab Data Curation The data curation process consisted of two phases. The first phase involved a pilot study to test the functionality of the curation tool and to preliminarily assess the feasibility of using the Big Five Personality Questionnaire and the personality lexicon to evaluate annotation reliability. Based on the validation results from this pilot phase, formal data collection was then conducted. Sample source Participants were recruited using a convenience sampling approach through online advertisements, and the questionnaire was distributed via the Wenjuanxing platform (a Chinese online survey distribution tool). Inclusion criteria were: 1) adults aged 18 to 65 years; and 2) individuals who provided informed consent and were able to independently complete the questionnaire. Exclusion criteria were: 1) a self-reported history of a professionally diagnosed severe psychological or neurological disorder; and 2) individuals unable to comprehend or follow the study instructions; and 3) minors. The study protocol was approved centrally by the Ethics Review Committee of West China Hospital of Sichuan University (process number 20232426). Informed consent was obtained from all participants prior to their participation in the study. Data curation tools Demographic Variables Demographic variables included gender (male, female), age (18–24 years, 25–35 years, 36–50 years, or over 50 years), employment status (employed or unemployed), education level (junior high school, high school, vocational school, college, bachelor’s degree, or graduate degree). Big five personality questionnaire The Chinese version of the International Personality Item Pool–NEO 120-item (IPIP-NEO-120) questionnaire was used to assess participants’ Big Five personality traits, with Cronbach’s α coefficients indicated good internal consistency at both the domain and facet levels ( α = 0.63–0.88) 61 , 63 . Based on their scores, participants were categorized as “High,” “Average,” or “Low” for each personality dimension and facet. “High” for scores in the top 30th percentile, “Low” for scores in the bottom 30th percentile, and “Average” for scores in the middle 40% (31st–69th percentile) 61 . Lexicon rating task Two versions of the lexicon rating task were developed for the pilot and formal phases of the study. In the pilot phase, 241 seed words labeled across different personality dimensions were randomly selected from validated personality questionnaires with clearly defined traits. These words were divided into six sets of 40, and each participant was randomly assigned one set. Participants rated each word using a five-point Likert scale, where higher scores indicated greater personal relevance (0 = Very inaccurate to 4 = Very accurate). In the formal phase, 3,600 words without clearly defined dimensions or facets in existing dictionaries or literature were selected from the seed lexicon. Words from validated personality questionnaires were excluded. These words were divided into 60 sets of 60, based on response time data from the pilot phase. Each participant was randomly assigned one set and rated the words using the same five-point Likert scale. Data curation results Pilot study A total of 50 participants were included in the pilot study, the majority of whom were female. Approximately half of the sample were young adults, with more than 50% having jobs and having completed undergraduate or higher education. Detailed demographic characteristics are shown in Table 6 . Table 6. Characterization of the general demographic information in the pre-curation (n = 50). Variable Name Frequency Percentage Gender Female 32 64% Male 18 36% Age Group 18–35 years 35 70% 36–50 years 4 8% Over 50 years 11 22% Employment Status Employed 31 62% Unemployed 19 38% Educational Level Junior High School 2 4% High School 1 2% Vocational School 4 8% College 5 10% Bachelor’s Degree 21 42% Graduate Degree 17 34% Open in a new tab Most participants scored within the moderate range on the Big Five personality traits. However, Neuroticism showed greater variability, with more participants with either high or low levels of Neuroticism compared to the other dimensions. Distributions of personality traits are detailed in Table 7 . Table 7. Big Five personality traits in the pre-curation (n = 50). Big Five Personality Dimension level Frequency Percentage Agreeableness 93.90 ± 11.40 a Low 1 2% Medium 33 66% High 16 32% Extraversion 82.80 ± 11.50 a Low 2 4% Medium 42 84% High 6 12% Openness 88.8 ± 11.15 a Low 3 6% Medium 41 82% High 6 12% Conscientiousness 89.52 ± 12.53 a Low 5 10% Medium 37 74% High 8 16% Neuroticism 65.70 ± 13.66 a Low 5 10% Medium 40 80% High 5 10% Open in a new tab α: mean (std). Formal data curation The formal data curation included a total of 329 participants, the majority of whom were female. Most participants were young adults, and a large portion of them were unemployed. Nearly all participants had completed undergraduate or higher education. Detailed demographic characteristics can be found in Table 8 . Table 8. Characterization of the general demographic information in formal data curation (n = 329). Variable Name Frequency Percentage Gender Female 226 68.69% Male 103 31.31% Age Group 18–24 years 215 65.35% 25–35 years 95 28.87% 36–50 years 12 3.65% Over 50 years 7 2.13% Employment Status Employed 245 74.47% Unemployed 84 25.53% Educational Level Vocational School 4 1.22% Junior High School 8 2.43% High School 8 2.43% College 16 4.86% Bachelor’s Degree 139 42.25% Graduate Degree 154 46.81% Open in a new tab In terms of personality traits, the highest average score was observed in the Agreeableness dimension (91.86 ± 10.82), while the lowest average score was in the Neuroticism dimension (65.61 ± 16.38). The majority of participants fell within the moderate range for the Big Five personality traits. Trait Distribution details are provided in Table 9 . Table 9. Big Five Traits (n = 329). Dimension Low(%) Medium(%) High(%) M ± SD Agreeableness 2.7 70.8 26.5 91.86 ± 10.82 Extraversion 2.6 79.9 17.5 85.33 ± 11.12 Openness 4.0 80.0 16.0 91.42 ± 9.72 Conscientiousness 7.0 75.0 18.0 90.10 ± 12.02 Neuroticism 16.4 67.6 16.0 65.61 ± 16.38 Open in a new tab Data Records Datasets. The full dataset, as well as all raw material used for this study is publicly available at: 10.6084/m9.figshare.29596547.v3 64 . The main dataset is named personality_lexicon_Chinese_version.xlsx (identical to personality_lexicon_English_version.xlsx). This file contains all annotated personality adjectives, including the source, dimension, facet, and emotional polarity of each personality adjective. The column “id” contains a consecutive number identifying each individual lexical word. The dimension and facet indicate to which dimension and facet of personality each adjective belongs. In addition, the source and emotional valence (positive, negative, neutral) of the adjectives are provided. Other raw material includes the data collected from both the pilot study and the formal study. The pilot data are stored in pilot_study_Chinese_version.xlsx (identical to pilot_study_English_version.xlsx), and the formal data are stored in formal_data_revised_Chinese_version.xlsx. (identical to formal_data _revised_English_version.xlsx) Notably, the file formal_data_Chinese_version.xlsx contains two sheets: the first sheet includes the 241 personality adjectives used in the pilot study, and the second sheet includes the 3600 personality adjectives used in the formal study. In consideration of the differences between Chinese and English, we have provided two versions of the data, allowing future researchers to choose the language version that best meets their research needs. Both files contain demographic information, Big five inventory (IPIP-NEO-120) data, and rating scores for each study participant. Within the demographic information, the column “id” contains a consecutive number identifying each individual study participant; other details are described in the Methods. The big five inventory data include responses to 120 items as well as scores across the five dimensions of personality. All remaining columns represent the rating scores assigned to each lexical word by participants. To facilitate participants’ understanding of each adjective, each lexical word was embedded into a complete sentence during the rating process. Technical Validation Validation process Data from both the pilot and formal phases underwent the same two-step validation process. Big five personality questionnaire (IPIP-NEO-120) validation Reliability (the extent to which the questionnaire provides consistent or reliable measurements of a particular construct) was assessed using Cronbach’s alpha to measure internal consistency. Validity (the extent to which the questionnaire reflects the construct it is intended to measure) was evaluated using the Kaiser-Meyer-Olkin (KMO) measure of sampling adequacy and Bartlett’s test of sphericity. Following standard psychometric practice 65 , these two metrics pertain to the validity and reliability assessment of the personality assessment instrument, aiming to confirm that trait scores derived from the IPIP-NEO-120 possess psychometric reliability, thereby serving as benchmark labels for subsequent dictionary construction and evaluation. Lexicon validation To evaluate the accuracy of personality-related word annotations, a two-level validation procedure was conducted at both the Big Five dimension and facet levels. Only words labeled with either positive or negative emotional valence were included; neutral terms were excluded from analysis. Participants were stratified into high/medium and low trait groups based on their scores on the IPIP-NEO-120 questionnaire. For each word, average ratings were computed separately within each trait group, requiring a minimum of three participant ratings per group to ensure reliability. Validation was based on whether participants’ ratings aligned with their corresponding trait levels. For all dimensions and facets except Neuroticism, a word was considered a hit for high/medium trait participants if positively valenced words received high ratings (≥3) or negatively valenced words received low ratings (<3), excluding neutrals to enhance discriminative power. Conversely, for low trait participants, a hit occurred when positively valenced words received low ratings (<3) or negatively valenced words received high ratings (≥3). For example, if a participant had high agreeableness and also rated positively valenced words highly (≥3), it was considered a hit; conversely, if a participant had low agreeableness and rated negatively valenced words highly (≥3), it was considered a hit. For Neuroticism—conceptually associated with negative valence—this logic was reversed: participants with high/medium Neuroticism were expected to rate negatively valenced words highly (≥3) or positively valenced words low (<3), while those with low Neuroticism were expected to do the opposite. For example, if a participant had high neuroticism and also rated negatively valenced words highly (≥3), it was considered a hit; conversely, if a participant had low neuroticism and rated positively valenced words highly (≥3), it was considered a hit. Each word thus contributed two classification instances: one from the high/medium group (yielding either a true positive [TP] or false negative [FN]) and one from the low group (yielding either a true negative [TN] or false positive [FP]) 66 . Based on previous studies 67 , 68 , our study was more concerned with overall accuracy and therefore hit rate was used as an assessment metric. This allowed for the construction of confusion matrices and the calculation of hit rates for each dimension and facet. Hit rate calculation: Hit rate = TN + TP TP + FN + FP + TN *Where: TP: High or medium-trait user rates positive word ≥3 TN: Low-trait user rates positive word <3 Reverse logic for Neuroticism* The confusion matrix for word validation is shown in Table 10 . Table 10. Confusion matrix for word validation. Big Five Personality Test Levels Dimension/Facet level is medium or high Dimension/Facet level is low Word Score Positive valence word score ≥ 3 Negative valence word score < 3 True Positive (TP) False Positive (FP) Positive valence word score < 3 Negative valence word score ≥ 3 False Negative (FN) True Negative (TN) Open in a new tab The full process of constructing the personality lexicon is shown in Fig. 2 . Fig. 2. Open in a new tab Personality lexicon construction process. Validation results Pilot study In the pilot study, IPIP-NEO-120 questionnaire demonstrated good psychometric properties, with high reliability (Cronbach’s alpha = 0.876) and satisfactory validity (KMO = 0.794), confirming its suitability as a gold-standard measure. In the lexicon validation, all Big Five dimensions achieved hit rates above 0.7, indicating strong annotation accuracy. However, due to the limited sample size, rank-order consistency could not be reliably assessed for each dimension and facet. Detailed validation results are shown in Table 11 . Table 11. Validation results of personality lexicon dimensions in the pre-curation (n = 50). Questionnaire/Sub-questionnaire Hit Rate Overall 0.817 Openness 0.765 Conscientiousness 0.860 Extraversion 0.885 Agreeableness 0.718 Neuroticism 0.857 Open in a new tab Formal data curation In the formal data curation, IPIP-NEO-120 questionnaire consistently demonstrated good psychometric properties, with both reliability and validity scores exceeding 0.80 (Table 12 ) and the hit rates for each of the Big Five dimensions were all above 0.70, with an overall average hit rate of 0.787. At the facet level, all hit rates exceeded0.60, with an overall average hit rate of 0.768 (Table 13 ). These results indicate that the questionnaire reliably captured participants’ personality traits and that the lexicon annotations were accurate. Table 12. Reliability and validity of the Big Five Personality Inventory in formal data curation (n = 329). Questionnaire/Sub-questionnaire Reliability Validity Overall 0.820 0.843 Openness 0.794 0.775 Conscientiousness 0.883 0.881 Extraversion 0.825 0.839 Agreeableness 0.845 0.826 Neuroticism 0.923 0.911 Open in a new tab Table 13. Lexicon Validation Performance. Dimension Hit Rate Strongest Facet (HR) Weakest Facet (HR) Agreeableness 0.823 Morality (0.874) Trust (0.656) Extraversion 0.776 Cheerfulness (0.860) Gregariousness (0.700) Openness 0.734 Emotionality (0.825) Artistic interests (0.652) Conscientiousness 0.770 Dutifulness (0.814) Self-efficacy (0.660) Neuroticism 0.832 Immoderation (0.881) Vulnerability (0.711) Overall 0.787 — — Open in a new tab As shown in Table 14 , we compared the Chinese personality lexicon proposed in this study with existing cross-lingual personality lexicon. Table 14. Comparison of the proposed Chinese personality lexicon with existing personality lexicon across different languages. Resource name Language(s) Personality model Number of entries Granularity Validation approach Public This work Chinese; English Big Five 6,084 words facet-level, dimension-level, word-level 329 real-sample ratings Yes LIWC 2015 14 English; Chinese / ~6,400 entries word-level, dimension-level expert ratings, partial real-sample ratings Commercial BigFive: a Chinese textual dataset 15 Chinese Big Five 13,478 phrases facet-level, dimension-level, phrase level 3 expert ratings No Personality trait descriptor 44 English / 2,818 words word-level 1572 real-sample ratings Yes Open in a new tab Compared to existing resources, our personality lexicon is the first publicly available, facet-level Big Five personality lexicon for Chinese—grounded in psychometric theory and validated through large-scale empirical data. In contrast, most existing large-scale personality lexicons across languages are limited to simple word lists, often lacking fine-grained facet-level categorization and robust empirical validation. Usage Notes Researchers who want to conduct further analyses with the Personality Lexicon can download the full dataset “personality_lexicon.xlsx”, which is described in Data Records and is available at 10.6084/m9.figshare.29596547.v3 64 . In addition to the processed dataset, we also provide all raw datasets. Supplementary information 41597_2026_6783_MOESM1_ESM.xlsx (20.3KB, xlsx) Table S1 Definition of the Big Five personality dimensions and facets Acknowledgements This research was supported by Science and Technology Department of Sichuan Province (No. 2020YFS0582). Author contributions T.J., J.J. and W.Z. conceptualized the study and originated the dataset creation. T.J., X.K., Y.Y., H.Z., X.H., J.J. and Y.L. participated in the collection, annotation, and validation of the dataset. The manuscript was drafted by T.J. and edited by X.S., H.C. and Y.L. All authors have read and approved the final manuscript. Data availability The data supporting the findings of this study are openly available in figshare at 10.6084/m9.figshare.29596547.v3 64 . Code availability No code was generated or used during the study. Competing interests The authors declare no competing interests. Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. These authors contributed equally: Yuchen Li, Wei Zhang. Supplementary information The online version contains supplementary material available at 10.1038/s41597-026-06783-6. References 1. Majumder, N., Poria, S., Gelbukh, A. & Cambria, E. Deep learning-based document modeling for personality detection from text. IEEE Intell. Syst. 32 , 74–79 (2017). [ Google Scholar ] 2. Wang, D. F. & Cui, H. The structure of culture, language, and personality. J. Peking Univ. (Humanit. Soc. Sci.) (4), 38–46 (2000). 3. Lee, C. H. et al . The relations between personality and language use. J. Gen. Psychol. 134 , 405–413, 10.3200/GENP.134.4.405-414 (2007). [ DOI ] [ PubMed ] [ Google Scholar ] 4. Back, M. D. et al . Facebook profiles reflect actual personality, not self-idealization. Psychol. Sci. 21 , 372–374, 10.1177/0956797609360756 (2010). [ DOI ] [ PubMed ] [ Google Scholar ] 5. Dougherty, L. R. & Guillemette, L. M. Linking personality and cognition: a meta-analysis. Philos. Trans. R. Soc. B Biol. Sci. 373 , 20170282, 10.1098/rstb.2017.0282 (2018). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Baumert, A. & Schmitt, M. Personality and information processing. Eur. J. Pers. 26 , 1 10.1002/per.1850 (2012). [ Google Scholar ] 7. Fumero, A., Santamaría, C. & Johnson-Laird, P. The effect of personality on reasoning. Nat. Preced . 3 , 10.1038/npre.2008.2099.1 (2008). 8. Ait Baha, T. et al . The power of personalization: A systematic review of personality-adaptive chatbots. SN Comput. Sci. 4 , 661, 10.1007/s42979-023-02092-6 (2023). [ Google Scholar ] 9. Jiang, H., Guo, A. & Ma, J. Personality-aware chatbot: An emerging area in conversational agents. 10.13140/RG.2.2.15957.86249 (2020). 10. Boyd, R. L. & Pennebaker, J. W. Language-based personality: A new approach to personality in a digital world. Curr. Opin. Behav. Sci. 18 , 63–68, 10.1016/j.cobeha.2017.07.017 (2017). [ Google Scholar ] 11. Rocha, P. Cultural correlates of personality: A global perspective with insights from 22 nations. Cross-Cult. Res. 59 , 216–243, 10.1177/10693971241264363 (2024). [ Google Scholar ] 12. Smaldino, P. E., Lukaszewski, A., von Rueden, C. & Gurven, M. Niche diversity can explain cross-cultural differences in personality structure. Nat. Hum. Behav. 3 , 1276–1283, 10.1038/s41562-019-0730-3 (2019). [ DOI ] [ PubMed ] [ Google Scholar ] 13. Nezlek, J. B., Schütz, A., Schröder-Abé, M. & Smith, C. V. A cross-cultural study of relationships between daily social interaction and the five-factor model of personality. J. Pers. 79 , 811–840, 10.1111/j.1467-6494.2011.00706.x (2011). [ DOI ] [ PubMed ] [ Google Scholar ] 14. Koutsoumpis, A. et al . The kernel of truth in text-based personality assessment: a meta-analysis of the relations between the Big Five and the Linguistic Inquiry and Word Count (LIWC). Psychol. Bull. 148 , 843–868, 10.1037/bul0000381 (2022). [ Google Scholar ] 15. Wang, Q., Liu, A., Yan, K., Hou, J. & Li, W. BigFive: a Chinese textual dataset supporting psychology knowledge graph construction. In Proc. IEEE Int. Conf. Knowl. Graph (ICKG) 77–83, 10.1109/ICKG59574.2023.00015 (IEEE, 2023). 16. Yuan, C., Hong, Y. & Wu, J. Personality expression and recognition in Chinese language usage. User Model. User-Adapt. Interact. 31 , 121–147, 10.1007/s11257-020-09276-2 (2021). [ Google Scholar ] 17. Kovacevic, N., Boschung, T., Holz, C., Gross, M. & Wampfler, R. Chatbots with attitude: Enhancing chatbot interactions through dynamic personality infusion. Proc. 6th ACM Conf. Conversational User Interfaces , 1–16 (2024). 18. He, X. & de Melo, G. Personality predictive lexical cues and their correlations. Proc. Recent Adv. Nat. Lang. Process ., 438–447 (2021). 19. Allbert, R., Wiles, J. K. & Grankovsky, V. Identifying and manipulating personality traits in LLMs through activation engineering. arXiv:2412.10427 (2024). 20. Kajonius, P. J. & Johnson, J. A. Assessing the structure of the Five Factor Model of personality (IPIP-NEO-120) in the public domain. Eur. J. Psychol. 15 , 260–275, 10.5964/ejop.v15i2.1671 (2019). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: development of the IPIP-NEO-120. J. Res. Pers . 51 , 78–89, 10.1016/j.jrp.2014.05.003 (2014). 22. Maples, J. L. et al . A test of the International Personality Item Pool representation of the revised NEO personality inventory and development of a 120-item IPIP-based measure of the five-factor model. Psychol. Assess. 26 , 1070–1084, 10.1037/pas0000004 (2014). [ DOI ] [ PubMed ] [ Google Scholar ] 23. Power, R. A. & Pluess, M. Heritability estimates of the Big Five personality traits based on common genetic variants. Transl. Psychiatry 5 , e604, 10.1038/tp.2015.96 (2015). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Wright, A. G. C. & Simms, L. J. On the structure of personality disorder traits: conjoint analyses of the CAT-PD, PID-5, and NEO-PI-3 trait models. Personal. Disord. 5 , 43–54, 10.1037/per0000037 (2014). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Bleidorn, W. et al . The healthy personality from a basic trait perspective. J. Pers. Soc. Psychol. 118 , 1207–1225, 10.1037/pspp0000231 (2020). [ DOI ] [ PubMed ] [ Google Scholar ] 26. Anglim, J. et al . Predicting psychological and subjective well-being from personality: a meta-analysis. Psychol. Bull. 146 , 279–323, 10.1037/bul0000226 (2020). [ DOI ] [ PubMed ] [ Google Scholar ] 27. Park, G. et al . Automatic personality assessment through social media language. J. Pers. Soc. Psychol. 108 , 934–952, 10.1037/pspp0000020 (2015). [ DOI ] [ PubMed ] [ Google Scholar ] 28. Alansari, B. The Big Five Inventory (BFI): reliability and validity of its Arabic translation in non clinical sample. Eur. Psychiatry 33 , S209–S210, 10.1016/j.eurpsy.2016.01.500 (2016). [ Google Scholar ] 29. Lamkin, J., Maples‐Keller, J. L. & Miller, J. D. How likable are personality disorder and general personality traits to those who possess them? J. Pers. 86 , 173–185, 10.1111/jopy.12302 (2018). [ DOI ] [ PubMed ] [ Google Scholar ] 30. Hopwood, M., Lea, T. & Aggleton, P. Multiple strategies are required to address the information and support needs of gay and bisexual men with hepatitis C in Australia. J. Public Health 38 , 156–162, 10.1093/pubmed/fdv002 (2016). [ DOI ] [ PubMed ] [ Google Scholar ] 31. Assary, E. et al . Genetic architecture of environmental sensitivity reflects multiple heritable components: a twin study with adolescents. Mol. Psychiatry 26 , 4896–4904, 10.1038/s41380-020-0783-8 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Chen, J., Xu, J. & Li, H. The evolution and comparison of personality tests based on the five-factor approach. Adv. Psychol. Sci. 23 , 460, 10.3724/SP.J.1042.2015.00460 (2015). [ Google Scholar ] 33. Ashton, M. C. & Lee, K. The HEXACO-60: a short measure of the major dimensions of personality. J. Pers. Assess. 91 , 340–345, 10.1080/00223890902935878 (2009). [ DOI ] [ PubMed ] [ Google Scholar ] 34. Hilbig, B. E., Glöckner, A. & Zettler, I. Personality and prosocial behavior: linking basic traits and social value orientations. J. Pers. Soc. Psychol. 107 , 529–539, 10.1037/a0036074 (2014). [ DOI ] [ PubMed ] [ Google Scholar ] 35. Lee, K. & Ashton, M. C. Psychometric properties of the HEXACO-100. Assessment 25 , 543–556, 10.1177/1073191116659134 (2018). [ DOI ] [ PubMed ] [ Google Scholar ] 36. Ge, P. & Dai, X. Preliminary development of the Chinese adjective Big Five personality scale IV: development of the short form. Chin. J. Clin. Psychol. 26 , 642–646, 10.16128/j.cnki.1005-3611.2018.04.003 (2018). [ Google Scholar ] 37. Abdurahman, S. et al . A deep learning approach to personality assessment: generalizing across items and expanding the reach of survey-based research. J. Pers. Soc. Psychol. 126 , 312–331, 10.1037/pspp0000480 (2024). [ DOI ] [ PubMed ] [ Google Scholar ] 38. Kaiser, T., Del Giudice, M. & Booth, T. Global sex differences in personality: replication with an open online dataset. J. Pers. 88 , 415–429, 10.1111/jopy.12500 (2020). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 39. Wei, S. et al . Abnormal default-mode network homogeneity and its correlations with personality in drug-naive somatization disorder at rest. J. Affect. Disord. 193 , 81–88, 10.1016/j.jad.2015.12.052 (2016). [ DOI ] [ PubMed ] [ Google Scholar ] 40. Docherty, A. R. et al . SNP-based heritability estimates of the personality dimensions and polygenic prediction of both neuroticism and major depression: findings from CONVERGE. Transl. Psychiatry 6 , e926, 10.1038/tp.2016.177 (2016). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 41. Koutra, K. et al . Maternal depression and personality traits in association with child neuropsychological and behavioral development in preschool years: mother-child cohort (Rhea Study) in Crete, Greece. J. Affect. Disord. 217 , 89–98, 10.1016/j.jad.2017.04.002 (2017). [ DOI ] [ PubMed ] [ Google Scholar ] 42. Yarkoni, T. Personality in 100,000 words: A large-scale analysis of personality and word use among bloggers. J. Res. Pers. 44 (3), 363–373, 10.1016/j.jrp.2010.04.001 (2010). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 43. Olga, B., Polina, P. & Roman, T. Dark personalities on Facebook: Harmful online behaviors and language. Comput. Hum. Behav. 78 , 151–159, 10.1016/j.chb.2017.09.032 (2018). [ Google Scholar ] 44. Condon, D. M., Coughlin, J. & Weston, S. J. Personality trait descriptors: 2,818 trait descriptive adjectives characterized by familiarity, frequency of use, and prior use in psycholexical research. J. Open Psychol. Data 10 , 1, 10.5334/jopd.57 (2022). [ Google Scholar ] 45. Thalmayer, A. G., Saucier, G., Ole-Kotikash, L. & Payne, D. Personality structure in east and west Africa: Lexical studies of personality in Maa and Supyire-Senufo. J. Pers. Soc. Psychol. 119 , 1132–1152, 10.1037/pspp0000264 (2020). [ DOI ] [ PubMed ] [ Google Scholar ] 46. Burcaş, S. & Creţu, R. Z. Perfectionism and neuroticism: Evidence for a common genetic and environmental etiology. J. Pers. 89 , 819–830, 10.1111/jopy.12617 (2021). [ DOI ] [ PubMed ] [ Google Scholar ] 47. Chapman, B. P. & Goldberg, L. R. Act-frequency signatures of the Big Five. Pers. Individ. Differ. 116 , 201–205, 10.1016/j.paid.2017.04.049 (2017). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 48. Cutler, A. & Condon, D. M. Deep lexical hypothesis: Identifying personality structure in natural language. J. Pers. Soc. Psychol. 125 , 173–197, 10.1037/pspp0000444 (2023). [ DOI ] [ PubMed ] [ Google Scholar ] 49. De Raad, B. et al . Towards a pan–cultural personality structure: Input from 11 psycholexical studies. Eur. J. Pers. 28 , 497–510, 10.1002/per.1953 (2014). [ Google Scholar ] 50. Leising, D., Scharloth, J., Lohse, O. & Wood, D. What types of terms do people use when describing an individual’s personality? Psychol. Sci. 25 , 1787–1794, 10.1177/0956797614541285 (2014). [ DOI ] [ PubMed ] [ Google Scholar ] 51. Giorgi, S. et al . Regional personality assessment through social media language. J. Pers. 90 , 405–425, 10.1111/jopy.12674 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 52. Smith, M. M. et al . Perfectionism and the five-factor model of personality: A meta-analytic review. Pers. Soc. Psychol. Rev. 23 , 367–390, 10.1177/1088868318814973 (2019). [ DOI ] [ PubMed ] [ Google Scholar ] 53. Smith, M. M. et al . Are perfectionism dimensions vulnerability factors for depressive symptoms after controlling for neuroticism? A meta–analysis of 10 longitudinal studies. Eur. J. Pers. 30 , 201–212, 10.1002/per.2034 (2016). [ Google Scholar ] 54. Thalmayer, A. G., Job, S., Shino, E. N., Robinson, S. L. & Saucier, G. ǂŪsigu: A mixed-method lexical study of character description in Khoekhoegowab. J. Pers. Soc. Psychol. 121 , 1258–1283, 10.1037/pspp0000372 (2021). [ DOI ] [ PubMed ] [ Google Scholar ] 55. Wei, H. et al . Beyond the words: Predicting user personality from heterogeneous information. Proc. ACM Int. Conf. Web Search Data Mining , 355–364, 10.1145/3018661.3018717 (2017). 56. Weston, S. J., Gladstone, J. J., Graham, E. K., Mroczek, D. K. & Condon, D. M. Who are the Scrooges? Personality predictors of holiday spending. Soc. Psychol. Personal. Sci. 10 , 775–782, 10.1177/1948550618792883 (2019). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 57. Cui, J. Y., Dong, R. C., Li, W. Q. & Wang, W. J. Research on Personality of Netease Cloud Music User: Based on Internet Behavior and Lyrics Data. Psychol. Sci. 44 , 1403–1410, 10.16719/j.cnki.1671-6981.20210617 (2021). [ Google Scholar ] 58. Brunborg, G. S., Hanss, D., Mentzoni, R. A., Molde, H. & Pallesen, S. Problem gambling and the five-factor model of personality: a large population-based study. Addiction 111 , 1428–1435, 10.1111/add.13388 (2016). [ DOI ] [ PubMed ] [ Google Scholar ] 59. Kern, M. L. et al . The online social self: an open vocabulary approach to personality. Assessment 21 , 158–169, 10.1177/1073191113514104 (2014). [ DOI ] [ PubMed ] [ Google Scholar ] 60. Eichstaedt, J. C. et al . Closed- and open-vocabulary approaches to text analysis: a review, quantitative comparison, and recommendations. Psychol. Methods 26 , 677–698, 10.1037/met0000349 (2021). [ DOI ] [ PubMed ] [ Google Scholar ] 61. Johnson, J. A. Measuring thirty facets of the five factor model with a 120-item public domain inventory: Development of the IPIP-NEO-120. J. Res. Pers. 51 , 78–89, 10.1016/j.jrp.2014.05.003 (2014). [ Google Scholar ] 62. Roger, D. & Dai, X. Preliminary development of the Chinese Adjective Big Five Personality Scale I: Theoretical framework and test reliability. Chin. J. Clin. Psychol. 23 , 381–385, 10.16128/j.cnki.1005-3611.2015.03.001 (2015). [ Google Scholar ] 63. Ge, P. P. Revision of the Big Five Personality Scale IPIP-NEO-120. Doctoral dissertation, Yangzhou University (2017). 64. Tao, J., Jingwen, J., Li, Y. & Wei, Z. A standardized personality lexicon for enhancing natural language processing and personalized human-machine interaction. Figshare 10.6084/m9.figshare.29596547.v3 (2025). 65. DeVellis, R. F. Scale development: Theory and applications (4th ed.). Sage (2017). 66. Marina, O. & Guy, L. A systematic analysis of performance measures for classification tasks. Inf. Process. Manage. 45 , 427–437 (2009). [ Google Scholar ] 67. Wan, M., Wan, N. S. & Zulkarnain, N. Z. Comparative evaluation of lexicons in performing sentiment analysis. J. Adv. Comput. Technol. Appl. 2 , 14–20 (2020). [ Google Scholar ] 68. Khan, Z. A. et al . Developing lexicons for enhanced sentiment analysis in software engineering: An innovative multilingual approach for social media reviews. Comput. Contin. 79 (5), 2771–2793 (2024). [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Data Citations Tao, J., Jingwen, J., Li, Y. & Wei, Z. A standardized personality lexicon for enhancing natural language processing and personalized human-machine interaction. Figshare 10.6084/m9.figshare.29596547.v3 (2025). Supplementary Materials 41597_2026_6783_MOESM1_ESM.xlsx (20.3KB, xlsx) Table S1 Definition of the Big Five personality dimensions and facets Data Availability Statement The data supporting the findings of this study are openly available in figshare at 10.6084/m9.figshare.29596547.v3 64 . No code was generated or used during the study. Articles from Scientific Data are provided here courtesy of Nature Publishing Group ACTIONS View on publisher site PDF (1.4 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 674 · SHA-256 8ec5b90f05b7dfe5
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.