Introduction: Modern lifestyles have led to various mental health disorders and significant psychological issues—such as anxiety, stress, and depression—in many individuals. Psychologists encounter challenges during the interview process that stem from the process itself, the interviewee, and the interviewer; these errors can be attributed to the intricate and highly complex network of interactions occurring during face-to-face exchanges between the interviewer and the interviewee. Given that human beings and human interactions are inherently prone to error, and that the diagnosis of disorders is often influenced by the emotions and feelings of both the client and the therapist, there is a critical need for a tool capable of predicting depression with high accuracy and without human intervention.
Method: In this study, real data was used to predict depression and demonstrate the proposed method. The dataset comprised 600 individuals, of whom 297 were diagnosed as not having depression and 303 as having depression. The C5.0 decision tree was employed to extract rules, while a genetic algorithm was used to select the optimal rules yielding the highest prediction accuracy. The proposed method was evaluated based on the weighted count of correctly predicted samples.
Results: The results demonstrated that the proposed method offers significantly higher accuracy than many existing machine learning techniques. Furthermore, among all the extracted rules, those exhibiting the greatest impact—along with minimal overlap and conflict—regarding the detection of the presence or absence of depression were selected. Additionally, the algorithm identified the most influential features associated with depression. Based on the analyzed data, suicidal ideation, hypochondriasis, and childhood trauma were identified, in that order, as three significant and influential factors in depression.
Conclusion: The use of data mining techniques and machine learning tools is highly efficient and effective for predicting depression. Extracting the most significant features influencing the output variable can substantially reduce time and costs for clients while enhancing the accuracy of the therapist's diagnosis.
Type of Study:
Original Article |
Subject:
Data Mining Received: 2025/12/31 | Accepted: 2026/04/6