BPG is committed to discovery and dissemination of knowledge
Observational Study Open Access
Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Psychiatry. Sep 19, 2026; 16(9): 119308
Published online Sep 19, 2026. doi: 10.5498/wjp.119308
Gamified video-based affective computing framework for adolescent depression screening
Qi-Zhou Li, Jing-Yun Li, Xin-Xi Chen, Zi-Xu Wang, Xi-Wang Fan, Clinical Research Center for Mental Disorders, Shanghai Pudong New Area Mental Health Center, School of Medicine, Tongji University, Shanghai 200124, China
Li-Yuan-Ke Wang, Department of Psychosomatic Medicine, Suining Central Hospital, Suining 629000, Sichuan Province, China
Ming-Hao Wang, Wei Fang, Beijing Situ Wellbeing Technology Co., Ltd, Beijing 100080, China
Pei-Dong Chen, Department of Rehabilitation Therapy, Qilu Institute of Technology, Jinan 250200, Shandong Province, China
Qing-Shu Bai, Hua Shi, Department of Mental Health and Health Promotion, Chinese PLA Center for Disease Control and Prevention, Beijing 100071, China
Pei Sun, Faculty of Health and Wellness, City University of Macau, Macau 999078, China
Pei Sun, Tsinghua Laboratory of Brain and Intelligence and Department of Psychological and Cognitive Science, Tsinghua University, Beijing 100084, China
Rui Zhong, Mental Health Assessment Center, Shandong Mental Health Center, Jinan 250014, Shandong Province, China
ORCID number: Xi-Wang Fan (0000-0003-4180-0496); Hua Shi (0000-0002-8634-7333).
Co-first authors: Qi-Zhou Li and Li-Yuan-Ke Wang.
Co-corresponding authors: Rui Zhong and Hua Shi.
Author contributions: Li QZ led the methodological design, formal analysis, and writing of the original draft, reviewing and editing of the manuscript; Wang LYK and Wang MH contributed to conceptualization, data collection, and writing of the original draft, reviewing and editing; Li JY and Fang W supported the formal analysis and validation; Chen XX, Wang ZX, Chen PD, and Bai QS contributed to data collection; Sun P, Zhong R, and Fan XW supervised the project and provided conceptual guidance; Shi H provided project administration and supervision, and secured funding. Li QZ and Wang LYK contributed equally to this work as co-first authors. Zhong R and Shi H are designated as co-corresponding authors because they made substantial, distinct, and complementary contributions to the conception, supervision, coordination, and academic integrity of this multidisciplinary study. Zhong R provided important conceptual guidance and project supervision, particularly in relation to mental health assessment, clinical interpretation, and the relevance of the proposed screening framework for adolescent depression. His expertise helped ensure that the study design and interpretation of behavioral indicators were clinically meaningful. Shi H was responsible for overall project administration and supervision and played a key role in securing funding support, coordinating institutional collaboration, and ensuring the smooth implementation of the study. Given that the manuscript integrates clinical psychiatry, adolescent mental health screening, affective computing, and machine learning, effective correspondence requires expertise in both clinical-scientific interpretation and project-level coordination. Therefore, the designation of Zhong R and Shi H as co-corresponding authors accurately reflects their shared senior leadership, complementary responsibilities, and accountability for the work.
AI contribution statement: We declare that AI tools were used solely for language polishing at the final stage of manuscript preparation. Specifically, ChatGPT was used for minor linguistic refinement, after which the manuscript underwent additional professional language editing. No AI tools were used in the conception or design of the study, data collection or analysis, interpretation of results, or generation of any substantive manuscript content. All figures were produced using Python and MATLAB.
Institutional review board statement: This study was approved by Ethics Committee of Shanghai Pudong New District Mental Health Center, No. JW-IIT-2022-011GZ1.
Informed consent statement: Informed consent was obtained from all individual participants included in the study. Participants were informed about the study’s purpose, procedures, risks, and benefits, and their participation was voluntary. They were assured of confidentiality and the right to withdraw at any time without consequences.
Conflict-of-interest statement: All authors have completed the Unified Competing Interest form (available on request from the corresponding author) and declare: No support from any organization for the submitted work; no financial relationships with any organizations that might have an interest in the submitted work in the previous 3 years, no other relationships or activities that could appear to have influenced the submitted work.
STROBE statement: The authors have read the STROBE Statement-checklist of items, and the manuscript was prepared and revised according to the STROBE Statement-checklist of items.
Data sharing statement: The data used and analyzed in this study are not publicly available because they are the property of the Clinical Research Center for Mental Disorders, Shanghai Pudong New Area Mental Health Center, School of Medicine, Tongji University but may be available from the corresponding author on reasonable request.
Corresponding author: Hua Shi, Department of Mental Health and Health Promotion, Chinese PLA Center for Disease Control and Prevention, No. 20 Dongjie Street, Fengtai District, Beijing 100071, China. placdc@139.com
Received: January 26, 2026
Revised: March 5, 2026
Accepted: June 29, 2026
Published online: September 19, 2026
Processing time: 212 Days and 17.8 Hours

Abstract
BACKGROUND

Early detection of adolescent depression remains challenging, as conventional assessments rely on self-reporting and clinician-administered interviews, which constrain scalability and ecological validity. Recent advances in affective computing and computer vision enable the extraction of objective, low-cost behavioral markers from naturalistic visual data.

AIM

To develop and evaluate a video-based machine learning framework for adolescent depression screening using visual features captured during a gamified affective task designed to elicit spontaneous emotional and attentional responses.

METHODS

In this cross-sectional screening-model development and validation study, 383 adolescents aged 10-19 years (188 with clinically diagnosed major depressive disorder, 195 healthy controls) were recruited from community and clinical settings in Shanghai, China (March 2024 to August 2025). Participants performed a gamified Whac-A-Mole task while being recorded by a 720 p webcam (16 fps). Forty visual behavioral features, encompassing facial action units, gaze metrics, head pose, and valence-arousal dynamics, were extracted and used to train and evaluate a random forest classifier using stratified 10-fold cross-validation and an internal train-test split.

RESULTS

The model achieved strong performance, with an area under the curve (AUC) of 0.813 on the internal test set and a cross-validation AUC of 0.886 in the full sample. Accuracy, precision, and F1-score in the cross-validation analysis were 0.812, 0.841, and 0.799, respectively. Key discriminative indicators included AU6 (cheek raiser), AU12 (lip corner puller), gaze yaw variability, and valence fluctuation. Probability calibration demonstrated well-balanced confidence estimates and stable performance across validation folds.

CONCLUSION

This study establishes the feasibility of using affective computing-derived visual features from a gamified setting to identify adolescent depression. By leveraging spontaneous, nonverbal behavior in an engaging and stigma-free context, the framework provides a validated, interpretable, and scalable complementary screening tool to support early risk identification and referral, marking an important step toward ethical and accessible digital screening for adolescent mental health.

Key Words: Depressive disorder; Facial expression; Machine learning; Affective computing

Core Tip: This study introduces a gamified, video-based affective computing approach for adolescent depression screening that captures spontaneous behavioral cues during an interactive game task. By extracting facial action units, gaze behavior, head pose, and valence-arousal features from short video recordings recorded within the game, we trained a machine learning classifier that reliably distinguished depressed adolescents from healthy controls with strong external generalization. The findings demonstrate that visual behavioral signals alone can support accurate, low-cost, and non-invasive depression screening, highlighting the potential of gamified paradigms for scalable mental health screening in addition to traditional clinical interviews.



INTRODUCTION

Mental health problems have emerged as a major global public health concern in recent years. According to the latest World Health Organization report, over one billion people worldwide are affected by mental health disorders. Among them, depression stands out as one of the most prevalent and disabling conditions, characterized by persistent low mood, cognitive impairment, and an elevated risk of suicide[1]. It currently affects approximately 332 million individuals globally. Alarmingly, the prevalence of depression among adolescents has more than doubled in recent years, with one in four adolescents worldwide exhibiting depressive symptoms, a rate exceeding that observed in adults[2,3]. Cross-national analyses further indicate that adolescent depression and related mental health problems have escalated over the past decade, particularly in high-income countries[4].

Early identification of depression is critical for timely intervention; however, existing diagnostic approaches face substantial limitations. Standard assessments, such as the Patient Health Questionnaire (PHQ-9) and Hamilton Depression Rating Scale (HDRS), rely heavily on self-reported symptoms and clinician-administered interviews[5]. Although widely adopted in clinical practice, these methods are prone to reporting bias, recall inaccuracies, and inter-rater variability. Consequently, a considerable proportion of adolescents with depression remain undetected or misclassified, delaying access to appropriate treatment[6].

Recent advances in computational psychiatry have introduced neuroimaging- and machine-learning-based approaches for depression detection, which have demonstrated promising diagnostic performance. Nevertheless, evidence from meta-analyses highlights persistent barriers, including high financial cost, technical complexity, and limited accessibility, that hinder large-scale implementation and clinical translation[7-9]. This has led to growing interest in developing low-cost, non-invasive, and behaviorally grounded alternatives that can operate effectively outside traditional clinical settings. Video-based behavioral analysis represents a particularly promising direction in this regard. Observable cues such as facial action units (FAUs), eye gaze, head movement, and emotional expression provide rich and objective information reflecting depression-associated affective and psychomotor alterations. Empirical studies have confirmed the feasibility of visual behavior as a reliable marker of depression; features including FAUs, gaze trajectories, and head-motion dynamics have yielded robust classification accuracy in distinguishing depressive from non-depressive individuals[5,10-14]. These findings underscore the scalability and ecological validity of camera-based mental health assessment. However, most existing behavioral datasets are derived from clinical interview recordings, such as the Distress Analysis Interview Corpus, or from highly controlled laboratory tasks designed to elicit specific emotions[15-17]. Such contexts may prompt scripted or self-monitored responses, reducing the authenticity of the captured behavioral signals[18]. Despite recent advances, most multimodal and video-based behavioral analysis frameworks have been trained and validated on adult datasets, with only limited efforts explicitly addressing adolescent populations. As adolescence constitutes a pivotal developmental window for the emergence and early recognition of depressive symptoms, the creation of models specifically tailored to this age group remains an urgent priority.

To address these gaps, the present study developed a gamified, video-based experimental paradigm designed to elicit and record spontaneous behavioral and affective responses among adolescents. Using these recordings, we aimed to train a machine learning model to classify depressive vs non-depressive participants. Prior research has shown that gamified tasks enhance engagement and evoke authentic affective and attentional reactions, making them particularly suitable for emotion-based assessments[19,20]. Therefore, we designed an interactive Whac-A-Mole task, and participants completed the task while their facial videos were recorded through an online platform integrated with affective computing algorithms. From these recordings, a range of visual and emotional features were automatically extracted, including FAUs, gaze direction, head pose, and affective dimensions, such as valence, arousal, and basic emotions. Importantly, the purpose of the Whac-A-Mole task was not to diagnose depression through game performance itself, but to provide a standardized and adolescent-friendly context in which depression-relevant nonverbal behaviors could be elicited and recorded. Depression is associated with alterations in affective reactivity, psychomotor behavior, attentional engagement, and spontaneous facial expressivity; these processes can manifest in how individuals orient to stimuli, modulate facial movement, and sustain visually guided responses, even during simple interactive tasks. From this perspective, a brief game-based paradigm serves as a behavioral elicitation setting, enabling objective capture of visual-behavioral signatures that may reflect underlying depressive symptomatology. Accordingly, if behavioral differences between depressed and non-depressed adolescents can be reliably identified in such a low-pressure and engaging context, this paradigm may offer a clinically useful complementary screening approach that is both adolescent-friendly and practically scalable. Taken together, this study aimed to investigate the feasibility of a non-invasive, low-cost, and scalable screening framework for adolescent depression, designed to detect subtle behavioral differences between clinically diagnosed and healthy adolescents across both community and clinical contexts.

The contributions of this paper are as follows: (1) Methodological innovation: We propose a gamified, video-based experimental paradigm specifically designed for adolescents. This paradigm not only elicits spontaneous affective and behavioral responses in an ecologically valid setting but also enables systematic evaluation of whether such naturalistic data possess sufficient quality and informational depth to support machine-learning-based analysis; (2) Train and validation: Leveraging data collected through this paradigm, we trained and evaluated machine-learning classifiers to distinguish adolescents with depression from healthy controls. The resulting models achieved strong discriminative performance, confirming that visual-behavioral cues captured in the gamified setting can serve as reliable and interpretable markers of depressive status; and (3) Practical and clinical relevance: This study provides preliminary evidence for the feasibility of a non-invasive, low-cost, and scalable video-based screening framework for adolescent depression. It also illustrates a potential translational pathway through which gamified behavioral assessments could complement traditional symptom-based evaluations, thereby facilitating early detection across both community and clinical contexts.

MATERIALS AND METHODS
Participants

A total of 437 adolescents aged 10-19 years were recruited from both clinical and community settings in Shanghai between March 2024 and August 2025. Participants with depressive symptoms (n = 243) were referred from the Shanghai Pudong Mental Health Center and met the DSM-5 diagnostic criteria for major depressive disorder (MDD) without comorbid psychiatric conditions, as confirmed by clinical evaluation. All depressive participants were unmedicated at the time of testing. Healthy controls (n = 195) were recruited from local schools and screened to exclude individuals with major emotional or psychiatric disorders.

All participants provided written informed consent. For those under 18 years of age, consent was additionally obtained from a parent or legal guardian. Because the study involved video recording of adolescents in a mental-health screening context, additional safeguards were implemented to protect privacy and autonomy. All participants and, for those under 18 years of age, their parent or legal guardian, were informed that facial videos would be recorded during task performance and used solely for research on behavioral screening rather than as a stand-alone diagnostic decision. Participation was entirely voluntary, and withdrawal was permitted at any time without penalty. All analyses were conducted on extracted visual-behavioral features rather than on identifiable personal information in the reporting stage, and access to the raw recordings was restricted to authorized research personnel under institutional oversight. After enrollment, participants completed two sets of gamified psychological tasks during which facial data were recorded for analysis. Due to low task adherence or poor data quality, 55 participants from the depression group were excluded. The final dataset comprised 383 participants: 188 with MDD and 195 healthy controls. The study protocol was approved by the Ethics Committee of Shanghai Pudong New District Mental Health Center (PDJW-IIT-2022-011GZ1). Participant demographics are summarized in Table 1.

Table 1 Sample characteristics, n (%).
Characteristic
Depression (n = 188)
Health (n = 195)
P value
Age (years), mean (SD)17.09 (3.89)14.06 (0.76)< 0.001
Sex< 0.001
    Male50 (26.9)113 (58.2)
    Female136 (73.1)81 (41.8)
Education level< 0.001
    Middle school year 10 (0.0)185 (95.4)
    Secondary school year 175 (40.3)9 (4.6)
    Secondary school year 2111 (59.7)0 (0.0)
Framework

To elicit spontaneous affective and attentional responses, participants completed a gamified “Whac-A-Mole” task implemented on an online platform integrated with affective computing modules developed by our team. The task was designed to create a structured, engaging, and low-pressure environment that encourages natural behavioral expression, particularly suitable for adolescents. During each session, cartoon moles appeared randomly at various screen positions. Participants were instructed to tap the targets as quickly as possible once a mole appeared. Each session consisted of 50 consecutive rounds, with a 2-second duration per round. If a participant responded within the time limit, the next round began immediately; otherwise, the trial automatically terminated after 2 seconds. The brief response window introduced moderate time pressure, promoting subtle facial reactions and attentional shifts relevant for affective-behavioral analysis. Throughout the entire session, a 720 p high-definition camera operating at 16 frames per second (fps) continuously recorded facial expressions and head movements. These recordings constituted the raw data for subsequent affective computing analysis, including extraction of FAUs, gaze direction, head pose, and emotion-related dimensions (valence and arousal).

Affective computing algorithm

Video data recorded by the 720 p high-definition camera were subsequently processed using an integrated affective computing pipeline designed to capture multiple dimensions of participants’ behavioral and emotional states. This pipeline jointly estimates basic emotion categories, continuous valence-arousal dimensions, head and gaze orientation, and FAUs. Such multimodal integration enables comprehensive modeling of visual affective cues that are highly relevant for depression screening. Each computational module and its underlying architecture are described in detail below.

Pre-processing

Each video was first split into individual image frames. A face detector was applied to obtain face bounding boxes and 3D facial landmarks in each frame. Detected faces were cropped and aligned based on five facial keypoints, then resized to 112 pixels × 112 pixels. For frames where no valid face was detected (e.g., due to occlusion or head turning), the nearest temporally valid frame was used as a substitute to ensure continuity of the feature sequence.

Basic emotion recognition

This module classifies each frame into one of seven basic emotion categories: Neutral, happy, sad, angry, surprise, fear, and disgust. Two complementary feature extractors were employed: (1) A DenseNet-based model[21,22] pre-trained on FER+[23] and AffectNet[24] for facial expression classification, producing 342-dimensional visual features; and (2) An IResNet100-based model[25] pre-trained on FER+[23], RAF-DB[26], and AffectNet[24], yielding 512-dimensional visual features. Additionally, a Masked Autoencoder[27] was pre-trained in a self-supervised manner on a collection of approximately 1.2 million face images from multiple datasets including DFEW[28], EmotionNet[29], and FERV39k[30], producing 768-dimensional features. The extracted visual features were concatenated with audio features (described below) to form multimodal representations, which were then fed into temporal encoders, either a 4-layer Transformer encoder (with 4 attention heads and a feed-forward dimension of 1024) or a bidirectional LSTM, to capture temporal context across video frames. Fully connected classification layers with hidden sizes of {512, 256} and dropout rate of 0.1 produced the final emotion predictions. Models were trained for 30 epochs using the Adam optimizer[31]. An ensemble of multiple models with different feature combinations and temporal encoders was used to improve robustness and generalization following previously described methodology[22,32].

Valence-arousal estimation

Grounded in the Circumplex Model of Affect, this module estimates continuous valence (pleasantness) and arousal (activation) values for each frame. Multimodal feature extraction combined visual features from DenseNet (342-d), IResNet100 expression features (512-d), IResNet100-based FAU features (512-d, pre-trained on an authorized commercial FAU dataset), and MobileNet-based features (512-d, pre-trained on AffectNet for valence-arousal estimation). Audio features included low-level descriptors-eGeMAPS (23-d) and ComParE 2016 (130-d) extracted via openSMILE, as well as deep features from a pre-trained wav2vec 2.0 model (768-d, pre-trained and fine-tuned on 960 hours of Librispeech[33] and a VGGish model (128-d, pre-trained on AudioSet Visual[34]. Audio features were concatenated and projected through a fully connected layer to form a unified multimodal representation. Four types of temporal encoders were explored: LSTM, GRU, a 4-layer Transformer encoder, and a hybrid Transformer-LSTM architecture. The Transformer-LSTM combination first applies the Transformer encoder to capture local temporal patterns, then feeds the resulting representations into an LSTM to model global temporal dependencies[35,36]. Regression heads with hidden sizes of {512, 256} produced frame-level valence and arousal predictions, trained with concordance correlation coefficient loss. Post-processing smoothing (window sizes of 5-10 frames) was applied to refine predictions. This approach achieved first place in the Affective Behavior Analysis in the wild (ABAW) valence-arousal estimation challenge[35,36].

Head pose and gaze estimation

This module estimates head rotation (yaw, pitch, roll) and gaze direction using monocular 3D facial landmark detection. A convolutional neural network (CNN)-based multi-model fusion approach was employed[37]: One branch extracts gaze features from aligned facial red, green, blue images using a multi-input CNN classifier, while a parallel branch uses a 2D CNN-based PointNet to estimate head pose from 3D facial landmarks. Three face alignment strategies were applied to normalize facial information and reduce ambiguity between head pose and eye gaze. The fusion of head poses and gaze features provides reliable indicators of attentional focus and cognitive engagement, achieving 81.5% accuracy on the DGW cross-subject test dataset[37].

FAU detection

Based on the facial action coding system, this module detects micro-expressions and subtle facial muscle activations, including AU6 (cheek raiser), AU12 (lip corner puller), AU15 (lip corner depressor), among others. The architecture employs IResNet100 as the backbone with feature pyramid networks and single stage headless modules to enlarge the receptive field and extract multi-scale facial texture features[38]. To address the inherent class imbalance in AU datasets, a novel face-masking data augmentation strategy was used: Upper or lower face regions were occluded using facial landmark detection, and the AU labels corresponding to covered regions were set to zero, and detection was selectively applied to under-represented AUs[39]. Multiple temporal models were explored for sequence modeling, including Transformer, TCN, GRU, BiGRU, LSTM, BiLSTM, and hybrid architectures combining these modules[39]. Three pre-trained backbones were used as initialization: An AU detection model, a facial expression model, and a face recognition model. Model-level ensembling further boosted performance, achieving F1 scores of 49.82% (2nd place, 3rd ABAW) and 54.22% (5th ABAW)[38,39].

Feature extraction

In total, data from 383 adolescents were analyzed, with each participant represented by 40 features derived from the affective computing outputs. Initially, 20 frame-level variables were extracted from the video recordings. To obtain temporally stable and noise-reduced representations, a sliding-window approach was applied to the continuous time series: Each video was segmented into 20000 millisecond windows with a 10000-millisecond overlap. Within each window, the mean and standard deviation (SD) of every feature were computed to capture both the average expression tendencies and short-term fluctuations in behavioral dynamics. This process effectively doubled the original feature set, yielding 40 aggregated visual descriptors for each participant. The mean and SD values of the 20 visual features for the depressive and healthy control groups are summarized in Table 2. No data points were excluded or relabeled during preprocessing.

Table 2 Visual features (mean and SD).
Visual features
Depression (n = 188)
Health (n = 195)
Facial emotion
    Valence-0.13 (0.10)-0.14 (0.14)
    Arousal0.03 (0.14)0.08 (0.16)
    Neutral0.75 (0.26)0.65 (0.30)
    Happy0.02 (0.05)0.04 (0.07)
    Sad0.08 (0.12)0.11 (0.17)
    Angry0.03 (0.05)0.03 (0.04)
    Surprise0.03 (0.05)0.06 (0.09)
    Fear0.01 (0.01)0.01 (0.02)
    Disgust0.09 (0.16)0.11 (0.16)
Gaze
    Pitch7.92 (4.93)8.66 (5.38)
    Yaw-0.90 (9.28)1.96 (9.01)
Head pose
    Roll-0.04 (2.55)-0.05 (2.63)
    FAUs
    Fau_10.18 (0.14)0.18 (0.16)
    Fau_20.32 (0.19)0.24 (0.19)
    Fau_40.09 (0.09)0.11 (0.12)
    Fau_60.002 (0.01)0.01 (0.02)
    Fau_120.01 (0.02)0.03 (0.07)
    Fau_150.02 (0.04)0.06 (0.08)
    Fau_200.02 (0.04)0.04 (0.06)
    Fau_250.14 (0.13)0.15 (0.17)
Model building methods

After data cleaning, binary classification models were trained to detect depression. The random forest classifier (RFC) was chosen as the primary model owing to its robustness against multicollinearity, capacity to model nonlinear relationships, and relative insensitivity to feature scaling. RFC combines bootstrap aggregation and random feature subspace sampling to mitigate overfitting while maintaining strong predictive performance. The dataset (n = 383; 40 visual-affective features including FAUs, gaze direction, head pose, valence, arousal, and discrete emotions) was divided into a training set (70%, n = 268) and a test set (30%, n = 115) using stratified sampling to preserve class balance. Any demographic variables such as age, sex, education level, and clinical scale scores, were not included as model inputs.

Before model training, correlation diagnostics and statistical screening were performed. Because our goal was to identify the best-performing RFC, feature selection was technically necessary prior to training. Features that exhibited significant differences between the depressive and healthy control groups were prioritized for model construction. However, given that tree-based algorithms like RFC are generally tolerant to redundant inputs, two versions of the model were trained for comparison: (1) RFC-feature, using only statistically significant features identified during feature selection; and (2) RFC-40, using all 40 extracted visual features. To support the feature selection process, both Pearson’s correlation coefficient (r) and Spearman’s rank correlation coefficient (ρ) were computed across all feature pairs to assess linear and monotonic dependencies. Correlation matrices were visualized by heatmaps to identify block-wise relationships within and across feature families. These analyses provided diagnostic insight into feature redundancy and complementarity, serving as a reference rather than strict exclusion criteria. Subsequently, a three-channel feature selection framework was applied to derive a compact and discriminative subset of variables: (1) Statistical filtering: Two-sample t-tests (with false discovery rate correction) were performed to compare depressive and healthy groups for each feature, supplemented by effect size estimation; (2) Information-theoretic evaluation: Mutual information and F-score metrics were computed to quantify each feature’s discriminative power relative to the class labels; and (3) Model-based selection: A preliminary random forest (RF) “probe model” was trained on the full dataset to estimate feature importance scores, with stability verified across resampled subsets.

Hyperparameter optimization for the RFC was conducted using a grid search with five-fold cross-validation to balance bias and variance. The best configuration, selected based on cross-validation accuracy, was as follows: (1) n_estimators = 200 (ensures stable ensemble performance); (2) Max_depth = 8 (controls tree complexity, prevents overfitting); (3) Max_features = “log2” (encourages feature diversity across trees). min_samples_split = 15; (4) Min_samples_leaf = 5 (regularization constraints); and (5) Bootstrap = True (sampling with replacement).

The trained RFC models were serialized using joblib to ensure reproducibility. Model performance was evaluated on the internal test set using Accuracy, Precision, Recall, F1-score, and receiver operating characteristic (ROC)-area under the curve (AUC). Additionally, probability calibration curves were analyzed to assess confidence distribution and threshold sensitivity. For benchmark comparison, support vector machine (SVM) classifiers were also developed. Unlike RFC, SVMs require standardized inputs and benefit from dimensionality reduction to mitigate overfitting in high-dimensional feature spaces. Therefore, principal component analysis was applied prior to SVM training. Both linear and radial basis function kernels were evaluated, and hyperparameters (C, γ) were optimized through grid search with five-fold cross-validation. Although SVMs served as a valuable comparative baseline, RFC consistently outperformed them across all major evaluation metrics. Consequently, RFC was implemented as the primary model for subsequent analyses and reporting.

RESULTS
Feature correlation and selection for RFC

Correlation analysis revealed strong internal dependencies among several FAU and emotion features (|r| > 0.70), particularly between the mean and SD pairs of the same variable (e.g., fau12_mean ↔ fau12_sd, happy_mean ↔ happy_sd). In contrast, gaze and head-pose features exhibited low intercorrelations, indicating that these behavioral modalities provide largely complementary information.

Based on a composite ranking that integrated statistical significance, information-theoretic relevance, and model-based importance, the top 22 most discriminative features were selected for the RFC-feature and SVM models (Figure 1). A cutoff of 22 was empirically determined by identifying the point of diminishing marginal gain in five-fold cross-validation AUC as features were added in rank order; adding further features beyond this point produced less than 0.5% improvement in AUC.

Figure 1
Figure 1 Top 22 feature composite scores.
Model performance and comparison

On the internal test set (n = 115), the RF model outperformed the SVM across all major evaluation metrics, achieving the following results: Accuracy = 0.730, precision = 0.755, recall = 0.661, F1-score = 0.705, and AUC = 0.813. Compared with SVM, the RF model achieved notable gains in accuracy (+0.07), precision (+0.11), and AUC (+0.06) (Figure 2A and B). Although the SVM exhibited slightly higher sensitivity (i.e., recall), the RF achieved a more balanced trade-off between precision and recall, resulting in superior overall performance.

Figure 2
Figure 2 Performance evaluation of the random forest classifier. A: Random forest (RF) improvements of +0.0695 accuracy, +0.1110 precision, +0.0439 F1, and +0.0557 area under the curve (AUC) relative to support vector machine; B: Receiver operating characteristic (ROC) for the RF on the test set (AUC = 0.813); the grey diagonal indicates chance level; C: Confusion matrix of the RF model on the test set (threshold = 0.5). True negatives = 47, false positives = 12, false negatives = 19, true positives = 37. The RF achieved an accuracy of 0.730, precision of 0.755, recall of 0.661, and F1-score of 0.705; D: Prediction confidence distribution of the RF model (test set). Confidence values range from 0.50 to 0.93, with a mean of 0.688 (indicated by the orange dashed line); E: ROC curve of the RF model on the independent hold-out set; F: Predicted-probability distribution (RF) for the Depressed class on the test set (non-depressed in blue, depressed in red). The red dashed line marks the 05-decision threshold; the overlap (approximately 0.35-0.60) explains most errors. ROC: Receiver operating characteristic; AUC: Area under the curve; RF: Random forest.

When comparing the two RF models trained on different feature sets, RFC-40 (all features) consistently outperformed RFC-feature (selected features), underscoring the robustness of the RF in handling high-dimensional data. Consequently, RFC-40 was adopted as the final depression classification model for subsequent analyses.

The confusion matrix (Figure 2C) indicated 12 false positives and 19 false negatives at a 0.5 decision threshold, reflecting a conservative classification tendency that prioritizes specificity over sensitivity, a desirable property for clinical screening applications. The model’s predicted confidence scores (mean = 0.688) were well-calibrated (Figure 2D), with most predictions displaying moderate to high certainty.

Cross-validation performance in the full sample

To further assess model stability, the RFC was additionally evaluated using cross-validation across the full analytic sample (n = 383). The model showed stable and slightly improved performance, with accuracy = 0.812, precision = 0.841, recall = 0.761, F1-score = 0.799, and AUC = 0.886. The confusion matrix derived from the cross-validation predictions (true negative = 168, false positive = 27, false negative = 45, true positive = 143) demonstrated strong diagonal dominance, reflecting consistent discrimination and a low false-positive rate (specificity = 0.862).

Notably, the cross-validation AUC exceeded that of the internal test set, which may reflect greater stability when performance is estimated across repeated folds rather than from a single held-out split. The ROC curve (Figure 2E) further confirmed robust class separation across varying decision thresholds, supporting the reliability of the proposed video-based depression screening framework within the present sample. Specifically, several key discriminative visual features also demonstrated stable contributions across both the internal test set and the cross-validation analysis. For instance, the depressive group exhibited significantly reduced activation in fau6 [mean and SD = 0.002 (0.01) vs 0.01 (0.02)] and fau12 [mean and SD = 0.01 (0.02) vs 0.03 (0.07)], illustrating diminished positive affect, while gaze yaw variability was consistently higher in the depressive group [mean and SD = -0.90 (9.28) vs 1.96 (9.01)], reflecting attentional instability. This consistency suggests that the model captures behavioral alterations associated with adolescent depression.

DISCUSSION

This study demonstrates that visual behavioral features captured from short video recordings can effectively differentiate adolescents with depression from healthy controls. Leveraging visual cues extracted from a gamified paradigm, the RFC achieved robust performance across both the internal test set and cross-validation analyses (AUC = 0.81-0.89). These findings provide strong empirical support for the feasibility of integrating affective computing techniques into low-cost, scalable, and non-invasive screening tools for adolescent depression screening.

Comparison with previous research

Although our experimental paradigm differs substantially from most previous studies, the proposed model achieved higher classification performance than the majority of existing video-only approaches while remaining consistent with established trends in the literature. For instance, a study using in-the-wild smartphone selfies reported modest performance (balanced accuracy approximately 0.60; Matthews correlation coefficient approximately 0.14)[40], reflecting the difficulty of learning from unconstrained and heterogeneous visual input. While the metrics used in that study (balanced accuracy) and ours (AUC) are not directly comparable, as the former is threshold-dependent and reflects performance at a single cutoff, whereas the latter is threshold-independent and quantifies global class separability, their relative magnitudes remain informative. A balanced accuracy of approximately 0.60 represents only marginal improvement over random chance (0.50), whereas an AUC around 0.80 indicates reliable class discrimination across thresholds. From this perspective, our model exhibits comparatively stronger discriminative capacity within its evaluation framework.

In contrast, a recent study[41] reported very high AUCs (> 0.90) using RF and SVC models trained on mobile-derived AUs and head-pose features. However, their small sample size (n = 25) raises concerns regarding overfitting and limited generalizability. When interpreted alongside large-scale evidence, a meta-analysis of 80 studies estimated a pooled AUC of 0.86 (95% confidence interval: 0.82-0.88) for video-based depression detection models[42]. Against this benchmark, our test AUC of 0.88 positions the proposed framework near the upper boundary of reported results. These findings collectively suggest that the gamified, video-based approach proposed in this study is both promising and competitive. Nonetheless, results should be interpreted cautiously, as variations in sample composition, recording conditions, and validation strategies may substantially influence reported model performance.

In addition, our model exhibits potential advantages compared with the video-only baselines reported in multimodal depression detection studies. In binary classification of depressed vs non-depressed individuals, our model achieved an F1 score approximately 12% higher on the test set than both the video-only and audiovisual ensemble models from the AVEC 2016 challenge[43]. Although the datasets and preprocessing pipelines differ, this relative improvement supports the notion that our gamified paradigm elicits more robust and diagnostically informative behavioral cues from facial expressions alone. Similarly, another study[44] observed that visual-only models generalize poorly to unseen data (F1 for the depressed class = 0.18), whereas our framework maintained stable performance in cross-validation analyses within the present sample (AUC = 0.886). Pampouchidou et al[44] further reported that only one visual descriptor (LMHI-FaceHOG) meaningfully contributed to classification, while most others provided limited discriminative value. In contrast, our analyses revealed that training on all 40 extracted visual features (RFC-40), rather than the 22 features retained after statistical screening and composite ranking (RFC-feature), consistently produced higher classification performance. This suggests that in the context of our gamified task, integrating low-level visual cues with affective-computing-derived representations captured from the same visual stream produces a richer and more informative feature space.

Taken together, these results support the view that our gamified paradigm effectively enhances the informational content of visual behavior, thereby improving the classification performance of video-based affective computing models. Nevertheless, further research using larger and more diverse datasets is warranted to verify the generalizability of this advantage across different populations and task designs.

Practical implications

In practical clinical applications, the choice of operating threshold should be tailored toward specific screening objectives. Under community-based mass screening, the threshold could be lowered to prioritize sensitivity, ensuring that most at-risk adolescents are identified for further clinical evaluation. Conversely, for specialized clinical settings where the aim is to confirm a clinical diagnosis or monitor treatment response, a higher threshold might be prioritized to maximize clinical specificity and avoid unnecessary interventions. Accordingly, this flexible structure contributes framework adaptability toward different resource-constrained environments in practical applications.

Limitations and future directions

Although the proposed video-based model demonstrated strong discriminative capability and stable performance across internal evaluation and cross-validation analyses, several limitations should be acknowledged. First, the study cohort was deliberately confined to adolescents to enhance clinical specificity; however, within this age range, the groups were not entirely balanced in age and sex distribution. Although these imbalances may introduce confounding, three lines of evidence argue against a primary demographic explanation. First, the model was intentionally restricted to behavioral features, excluding demographic variables and clinical scores as predictors. Second, supplementary analyses showed that none of the top 10 behavioral features was significantly associated with age in the healthy control group (all P > 0.05), and gazeYaw_mean was positively associated with symptom severity in the depression group (r = 0.16, P = 0.024), linking the feature set to clinical rather than demographic variation. Third, consistent performance across the internal evaluation and cross-validation analyses supports the interpretation that the identified features reflect pathological rather than purely demographic variation. Nevertheless, residual confounding from developmental stage or sex-related expression norms cannot be fully excluded and warrants targeted fairness analyses in future work. Additionally, model evaluation was conducted within a single study sample under closely aligned recording conditions, including similar recording platforms, camera specifications, and frame rates, which limited the extent to which distributional shift and external generalizability could be assessed. To confirm generalizability, it is important to involve additional multi-site and multi-device validation. Future studies should include varying the capture hardware, ambient lighting, and recording contexts to ensure the framework remains robust against technical and environmental distributional shifts. Another important observation concerns the overlap in predicted probability distributions for depressed and non-depressed participants (approximately 0.35-0.60; Figure 2F). This finding is consistent with real-world adolescent manifestations, where subthreshold or mixed depressive features frequently occur, and single-time-point clinical labels inevitably introduce noise. However, such overlap indicates that model predictions near the decision boundary are sensitive to threshold selection and base-rate prevalence, potentially constraining screening precision despite a high overall AUC. It is also important to note that the affective computing system described here is designed as a screening support tool rather than a standalone diagnostic instrument. In clinical deployment, the system outputs (emotion distributions, valence-arousal trajectories, AU activation patterns, and gaze metrics) serve as quantitative behavioral indicators that complement, but do not replace professional clinical assessment. The system is intended to be integrated into a two-stage clinical workflow: Automated screening to flag adolescents exhibiting affective patterns associated with elevated depression risk, followed by comprehensive clinical evaluation by qualified mental health professionals for flagged cases. We acknowledge that false-negative cases (i.e., at-risk individuals not flagged by the system) represent a critical concern in mental health screening. To mitigate this, we adopted several strategies: (1) Setting conservative detection thresholds to favor sensitivity over specificity during the initial screening stage; (2) Employing multi-dimensional behavioral indicators (combining emotion, AU, gaze, and valence-arousal features) rather than relying on any single modality, thereby reducing the chance of missing subtle affective cues; and (3) Recommending periodic re-screening rather than one-time assessment, as affective patterns may vary across sessions. We report AUC as the primary metric because it captures model performance across all possible operating thresholds; however, we recognize that in clinical practice, the selection of specific thresholds must balance false-negative and false-positive rates according to the requirements of the target population and clinical setting.

To address these issues, several future directions are suggested. To minimize demographic confounding, subsequent studies could adopt stratified recruitment or matching by age and sex, implement propensity-weighted or stratified cross-validation schemes, and systematically report fairness metrics such as differences in true/false positive rates (ΔTPR/ΔFPR) to evaluate equity in model performance. To enhance external validity, a multi-site, multi-device validation protocol that deliberately varies capture hardware, frame rate, and recording context should be implemented. Cross-site transfer learning (e.g., training on site A and testing on sites B/C) combined with distributionally robust learning objectives and uncertainty quantification methods (e.g., conformal calibration) under pre-registered evaluation criteria would provide a more rigorous assessment of generalizability. Beyond these demographic and technical aspects, future work should move beyond binary diagnosis. Employing a multi-task learning architecture that jointly models categorical diagnosis and continuous symptom severity (e.g., PHQ-9, HDRS) may improve sample efficiency, probability calibration, and resilience to label noise. Expanding the label space toward multi-label or hierarchical frameworks incorporating common comorbidities (e.g., anxiety, attentional deficit hyperactivity disorder, bipolar spectrum disorders) could help disentangle shared vs disorder-specific behavioral markers, aligning model outputs more closely with real-world adolescent psychopathology. Lastly, to explicitly address the boundary overlap, future studies should perform calibration and decision-curve analyses across various operating thresholds, reporting predictive values (PPV/NPV) under realistic prevalence conditions. Longitudinal or repeated assessments could stabilize diagnostic labels among borderline cases, while targeted multimodal augmentation, such as integrating speech, physiological, or short self-report signals for low-confidence samples, may enhance predictive reliability without compromising scalability or intrusiveness. Beyond these technical directions, future work should also address how the framework could be adapted and validated in more diverse real-world deployment contexts. For example, school-based or community health settings present different environmental constraints (variable lighting, consumer-grade devices, non-clinical administrators) compared with the current laboratory-aligned protocol. Pilot deployments in such settings, guided by implementation science frameworks, would yield practical insights into the robustness and acceptability of the system under routine operational conditions. Similarly, cross-cultural validation is warranted, as normative patterns of facial expressivity and gaze behavior may vary across ethnic and socio-cultural groups, potentially influencing model performance in ways that controlled single-site studies cannot anticipate. Engaging diverse adolescent populations in iterative co-design would not only strengthen generalizability but also ensure that the tool remains equitable and contextually appropriate across different healthcare systems.

CONCLUSION

In conclusion, this study introduces and validates a depression classification model using video-based features extracted through affective computing for adolescent depression screening, demonstrating robust screening accuracy and stable performance across internal testing and cross-validation analyses. By integrating gamified data collection with interpretable machine learning, the proposed system bridges the gap between conventional clinician-administered assessments and scalable digital screening solutions. The gamification strategy effectively enhances user engagement, minimizes self-report bias, and reduces stigma, enabling broader participation among adolescents who might otherwise be reluctant to engage with traditional symptom questionnaires.

Beyond methodological innovation, this work highlights the feasibility of deriving clinically meaningful affective representations from short, spontaneous video interactions, offering a low-cost, non-invasive, and privacy-preserving alternative for large-scale mental health monitoring. Although further validation across diverse sites, devices, and cultural settings remains necessary, the current results establish a strong empirical foundation for future multimodal and longitudinal extensions. Ultimately, this study represents an important step toward building accessible, ethical, and evidence-based digital tools capable of supporting early detection and equitable mental health care for adolescents worldwide.

ACKNOWLEDGEMENTS

We would like to express our sincere gratitude to all participants who contributed to this study, as well as their families and guardians for their support. We extend thanks to the many members of the research team who helped execute the study.

References
1.  Wang C, Tong Y, Tang T, Wang X, Fang L, Wen X, Su P, Wang J, Wang G. Association between adolescent depression and adult suicidal behavior: A systematic review and meta-analysis. Asian J Psychiatr. 2024;100:104185.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 10]  [Reference Citation Analysis (0)]
2.  Miller L, Campo JV. Depression in Adolescents. N Engl J Med. 2021;385:445-449.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 30]  [Cited by in RCA: 140]  [Article Influence: 28.0]  [Reference Citation Analysis (1)]
3.  Piccin J, Buchweitz C, Manfro PH, Pereira RB, Rohrsetzer F, Souza L, Viduani A, Caye A, Kohrt BA, Mondelli V, Swartz JR, Fisher HL, Kieling C. Predicting the incidence of depression in adolescence using a sociodemographic risk score: prospective follow-up of the IDEA-RiSCo study. BMJ Ment Health. 2025;28:e301207.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 5]  [Cited by in RCA: 3]  [Article Influence: 3.0]  [Reference Citation Analysis (0)]
4.  Zhang X, Mori Y, Abio A, Khorasani ZK, Gilbert S, Grimland M, Bezborodovs N, Wan Mohd Yunus WMA, Silwal S, Westerlund M, Praharaj SK, Ndetei DM, Kaneko H, Heinonen E, Sourander A. Cross-national research on adolescent mental health: a systematic review comparing research in low, middle and high-income countries. BMJ Glob Health. 2025;10:e019267.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 10]  [Reference Citation Analysis (0)]
5.  Ray A, Kumar S, Reddy R, Mukherjee P, Garg R.   Multi-level Attention Network using Text, Audio and Video for Depression Prediction. In: Proceedings of the 9th International on Audio/Visual Emotion Challenge and Workshop. New York: ACM, 2019: 81-88.  [PubMed]  [DOI]  [Full Text]
6.  Sampson M  Diagnosis, misdiagnosis, and becoming better: An investigation into epistemic injustice and mental health. Lampeter: University of Wales Trinity Saint David, 2021. Available from: https://repository.uwtsd.ac.uk/id/eprint/1849/.  [PubMed]  [DOI]
7.  Chen Z, Liu X, Yang Q, Wang YJ, Miao K, Gong Z, Yu Y, Leonov A, Liu C, Feng Z, Chuan-Peng H. Evaluation of Risk of Bias in Neuroimaging-Based Artificial Intelligence Models for Psychiatric Diagnosis: A Systematic Review. JAMA Netw Open. 2023;6:e231671.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 23]  [Reference Citation Analysis (0)]
8.  Flint C, Cearns M, Opel N, Redlich R, Mehler DMA, Emden D, Winter NR, Leenings R, Eickhoff SB, Kircher T, Krug A, Nenadic I, Arolt V, Clark S, Baune BT, Jiang X, Dannlowski U, Hahn T. Systematic misestimation of machine learning performance in neuroimaging studies of depression. Neuropsychopharmacology. 2021;46:1510-1517.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 30]  [Cited by in RCA: 75]  [Article Influence: 15.0]  [Reference Citation Analysis (0)]
9.  Cohen SE, Zantvoord JB, Wezenberg BN, Bockting CLH, van Wingen GA. Magnetic resonance imaging for individual prediction of treatment response in major depressive disorder: a systematic review and meta-analysis. Transl Psychiatry. 2021;11:168.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 16]  [Cited by in RCA: 54]  [Article Influence: 10.8]  [Reference Citation Analysis (0)]
10.  Shen G, Jia J, Nie L, Feng F, Zhang C, Hu T, Chua T, Zhu W.   Depression Detection via Harvesting Social Media: A Multimodal Dictionary Learning Solution. In: Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence. Melbourne: IJCAI, 2017: 3838-3844.  [PubMed]  [DOI]  [Full Text]
11.  Scherer S, Stratou G, Lucas G, Mahmoud M, Boberg J, Gratch J, Rizzo A, Morency L. Automatic audiovisual behavior descriptors for psychological disorder analysis. Image Vis Comput. 2014;32:648-658.  [PubMed]  [DOI]  [Full Text]
12.  Gahalawat M, Fernandez Rojas R, Guha T, Subramanian R, Goecke R.   Explainable Depression Detection via Head Motion Patterns. In: Proceedings of the 25th International Conference on Multimodal Interaction. Paris: ACM, 2023: 261-270.  [PubMed]  [DOI]  [Full Text]
13.  Islam R, Bae SW. FacePsy: An Open-Source Affective Mobile Sensing System - Analyzing Facial Behavior and Head Gesture for Depression Detection in Naturalistic Settings. Hum Comput Interact. 2024;8:1-32.  [PubMed]  [DOI]  [Full Text]
14.  Zhang Z, Zhang S, Ni D, Wei Z, Yang K, Jin S, Huang G, Liang Z, Zhang L, Li L, Ding H, Zhang Z, Wang J. Multimodal Sensing for Depression Risk Detection: Integrating Audio, Video, and Text Data. Sensors (Basel). 2024;24:3714.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 2]  [Cited by in RCA: 14]  [Article Influence: 7.0]  [Reference Citation Analysis (0)]
15.  Sun Y, Zhang H, Men J. Research on Identification and Classification of Depression in College Students through Feature Analysis. IEIE Trans Smart Process Comput. 2023;12:511-517.  [PubMed]  [DOI]  [Full Text]
16.  Hu B, Tao Y, Yang M. Detecting depression based on facial cues elicited by emotional stimuli in video. Comput Biol Med. 2023;165:107457.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 7]  [Reference Citation Analysis (0)]
17.  Zhang W, Mao K, Chen J. A Multimodal Approach for Detection and Assessment of Depression Using Text, Audio and Video. Phenomics. 2024;4:234-249.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 12]  [Reference Citation Analysis (0)]
18.  Lumsden J, Edwards EA, Lawrence NS, Coyle D, Munafò MR. Gamification of Cognitive Assessment and Cognitive Training: A Systematic Review of Applications and Efficacy. JMIR Serious Games. 2016;4:e11.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 213]  [Cited by in RCA: 233]  [Article Influence: 23.3]  [Reference Citation Analysis (0)]
19.  Hamari J, Koivisto J, Sarsa H.   Does Gamification Work? -- A Literature Review of Empirical Studies on Gamification. In: Proceedings of the 47th Hawaii International Conference on System Sciences. Hawaii: IEEE, 2014: 3025-3034.  [PubMed]  [DOI]  [Full Text]
20.  Mekler ED, Brühlmann F, Tuch AN, Opwis K. Towards understanding the effects of individual gamification elements on intrinsic motivation and performance. Comput Hum Behav. 2017;71:525-534.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 345]  [Cited by in RCA: 122]  [Article Influence: 13.6]  [Reference Citation Analysis (0)]
21.  Liu C, Tang T, Lv K, Wang M.   Multi-Feature Based Emotion Recognition for Video Clips. In: Proceedings of the 20th ACM International Conference on Multimodal Interaction. New York: ACM, 2018: 630-634.  [PubMed]  [DOI]  [Full Text]
22.  Liu C, Zhang X, Liu X, Zhang T, Meng L, Liu Y, Deng Y, Jiang W.   Facial Expression Recognition Based on Multi-modal Features for Videos in the Wild. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Vancouver: IEEE, 2023: 5872-5879.  [PubMed]  [DOI]  [Full Text]
23.  Barsoum E, Zhang C, Ferrer CC, Zhang Z.   Training deep networks for facial expression recognition with crowd-sourced label distribution. In: Proceedings of the 18th ACM International Conference on Multimodal Interaction. New York: ACM, 2016: 279-283.  [PubMed]  [DOI]  [Full Text]
24.  Mollahosseini A, Hasani B, Mahoor MH. AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild. IEEE Trans Affective Comput. 2019;10:18-31.  [PubMed]  [DOI]  [Full Text]
25.  Deng J, Guo J, Yang J, Xue N, Kotsia I, Zafeiriou S. ArcFace: Additive Angular Margin Loss for Deep Face Recognition. IEEE Trans Pattern Anal Mach Intell. 2022;44:5962-5979.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 28]  [Cited by in RCA: 42]  [Article Influence: 10.5]  [Reference Citation Analysis (0)]
26.  Li S, Deng W, Du J.   Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu: IEEE, 2017: 2852-2861.  [PubMed]  [DOI]  [Full Text]
27.  He K, Chen X, Xie S, Li Y, Dollár P, Girshick R.   Masked Autoencoders Are Scalable Vision Learners. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022: 16000-16009.  [PubMed]  [DOI]  [Full Text]
28.  Jiang X, Zong Y, Zheng W, Tang C, Xia W, Lu C, Liu J.   DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the Wild. In: Proceedings of the 28th ACM International Conference on Multimedia. New York: ACM, 2020: 2881-2889.  [PubMed]  [DOI]  [Full Text]
29.  Benitez-Quiroz CF, Srinivasan R, Martinez AM.   EmotioNet: An Accurate, Real-Time Algorithm for the Automatic Annotation of a Million Facial Expressions in the Wild. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE, 2016: 5562-5570.  [PubMed]  [DOI]  [Full Text]
30.  Wang Y, Sun Y, Huang Y, Liu Z, Gao S, Zhang W, Ge W, Zhang W.   FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE, 2022: 20922-20931.  [PubMed]  [DOI]  [Full Text]
31.  Kingma DP, Ba JL.   Adam: A Method for Stochastic Optimization. In: 3rd International Conference on Learning Representations (ICLR 2015). San Diego: ICLR, 2015.  [PubMed]  [DOI]  [Full Text]
32.  Zhang T, Liu C, Liu X, Liu Y, Meng L, Sun L, Jiang W, Zhang F, Zhao J, Jin Q.   Multi-Task Learning Framework for Emotion Recognition In-the-Wild. In: Karlinsky L, Michaeli T, Nishino K. Computer Vision - ECCV 2022 Workshops. Springer, Cham, 2023: 143-156.  [PubMed]  [DOI]  [Full Text]
33.  Panayotov V, Chen G, Povey D, Khudanpur S.   Librispeech: An ASR corpus based on public domain audio books. In: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Brisbane: IEEE, 2015: 5206-5210.  [PubMed]  [DOI]  [Full Text]
34.  Gemmeke JF, Ellis DPW, Freedman D, Jansen A, Lawrence W, Moore RC, Plakal M, Ritter M.   Audio Set: An ontology and human-labeled dataset for audio events. In: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). New Orleans: IEEE, 2017: 776-780.  [PubMed]  [DOI]  [Full Text]
35.  Liu X, Sun L, Jiang W, Zhang F, Deng Y, Huang Z, Meng L, Liu Y, Liu C.   EVAEF: Ensemble Valence-Arousal Estimation Framework in the Wild. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Vancouver: IEEE, 2023: 5863-5871.  [PubMed]  [DOI]  [Full Text]
36.  Meng L, Liu Y, Liu X, Huang Z, Jiang W, Zhang T, Liu C, Jin Q.   Valence and Arousal Estimation based on Multimodal Temporal-Aware Features for Videos in the Wild. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). New Orleans: IEEE, 2022: 2345-2352.  [PubMed]  [DOI]  [Full Text]
37.  Lyu K, Wang M, Meng L.   Extract the Gaze Multi-dimensional Information Analysis Driver Behavior. In: Proceedings of the 2020 International Conference on Multimodal Interaction. New York: ACM, 2020.  [PubMed]  [DOI]  [Full Text]
38.  Jiang W, Wu Y, Qiao F, Meng L, Deng Y, Liu C.   Model Level Ensemble for Facial Action Unit Recognition at the 3rd ABAW Challenge. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). New Orleans: IEEE, 2022: 2337-2344.  [PubMed]  [DOI]  [Full Text]
39.  Duan G, Deng H, Fu H, Wang L, Yang H. Comparisons of Electrolyte Balance Efficacy of Two Gelatin-Balanced Crystalloid for Surgery Patients Under General Anesthesia: A Multi-Center, Prospective, Randomized, Single-Blind, Controlled Study. Int J Gen Med. 2023;16:5855-5868.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Reference Citation Analysis (0)]
40.  Nepal S, Pillai A, Wang W, Griffin T, Collins AC, Heinz M, Lekkas D, Mirjafari S, Nemesure M, Price G, Jacobson NC, Campbell AT. MoodCapture: Depression Detection Using In-the-Wild Smartphone Images. Proc SIGCHI Conf Hum Factor Comput Syst. 2024;2024:996.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 3]  [Reference Citation Analysis (0)]
41.  Hardiansyah B, Hermawati FA, Saputra DA, Caesa DD. Depression Detection via Facial Expressions and Movement Analysis Using Machine Learning with Optimized Feature Selection. Int J Act Behav Comput. 2025;2025:1-25.  [PubMed]  [DOI]  [Full Text]
42.  Wang L, Wang C, Li C, Murai T, Bai Y, Song Z, Zhang S, Zhang Q, Huang Y, Bi X, Jiang J. AI-assisted multi-modal information for the screening of depression: a systematic review and meta-analysis. NPJ Digit Med. 2025;8:523.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1]  [Cited by in RCA: 12]  [Article Influence: 12.0]  [Reference Citation Analysis (0)]
43.  Valstar M, Gratch J, Schuller B, Ringeval F, Lalanne D, Torres Torres M, Scherer S, Stratou G, Cowie R, Pantic M.   AVEC 2016: Depression, Mood, and Emotion Recognition Workshop and Challenge. In: Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge. New York: ACM, 2016: 3-10.  [PubMed]  [DOI]  [Full Text]
44.  Pampouchidou A, Simantiraki O, Fazlollahi A, Pediaditis M, Manousos D, Roniotis A, Giannakakis G, Meriaudeau F, Simos P, Marias K, Yang F, Tsiknakis M.   Depression Assessment by Fusing High and Low Level Features from Audio, Video, and Text. In: Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge. New York: ACM, 2016: 27-34.  [PubMed]  [DOI]  [Full Text]
Footnotes

Peer review: Externally peer reviewed.

Peer-review model: Single blind

Specialty type: Psychiatry

Country of origin: China

Peer-review report’s classification

Scientific quality: Grade A, Grade B, Grade B, Grade B, Grade D

Novelty: Grade A, Grade B, Grade C, Grade C, Grade D

Creativity or innovation: Grade A, Grade C, Grade C, Grade C, Grade D

Scientific significance: Grade B, Grade B, Grade B, Grade B

P-Reviewer: Chen YX, Academic Fellow, PhD, Postdoctoral Fellow, China; Ng DKWJ, PhD, Researcher, Malaysia; Zhu HC, PhD, Professor, China S-Editor: Qu XL L-Editor: Filipodia P-Editor: Yang YQ

Write to the Help Desk