0Viewes
Evaluasi Pendekatan Hybrid Filtering dengan Machine Learning sebagai Decision Support pada Email Borderline di Mail Server Zimbra 8.8.15
Repository Analytics
Statistic Details
0Downloaded
0Accessed per month
0Countries
Statistic not available yet or restricted.
Loading...
Date
Authors
AR, Muhammad
Journal Title
Journal ISSN
Volume Title
Publisher
Politeknik Negeri Batam
Abstract
Rule-based email filtering systems such as SpamAssassin determine message status using accumulated rule scores and predefined operational thresholds. However, messages with scores close to the threshold may contain mixed spam and ham indicators, making their classification less certain. This study evaluates a hybrid filtering approach that retains rule-based filtering as the primary mechanism and applies machine learning as a parallel decision-support layer for borderline emails on a Zimbra 8.8.15 mail server. The dataset consisted of 25,220 operational email records, including 1,210 messages assigned integer SpamAssassin scores of 3 or 4 by the operational mail gateway. These two score categories represented the transition zone beginning at operational level 3 and ending before level 5 and were defined as the borderline subset in this study. The evaluation included the original stratified split, a Message-ID group-aware split, and a temporal holdout. In the main evaluation, Random Forest achieved a macro F1-score of 0.9995, a false positive rate of 0.0002, and a ROC-AUC of 0.9999. Its global performance remained stable in the group-aware evaluation with a macro F1-score of 0.9980 and in the temporal holdout with a macro F1-score of 0.9945. Performance on the temporal borderline subset was lower, indicating greater sensitivity to class distribution changes and the limited number of ham samples. The prototype generated confidence scores, review priorities, and administrator validation records without replacing the existing rule-based decision. Integration testing was conducted using 15 controlled test emails covering high-score, borderline, and normal categories. The rule-based threshold treated borderline and normal emails identically, while the model assigned spam probabilities of 0.4489-0.4930 and 0.0497 respectively, separating messages that required administrator review from those that did not. The results show that Random Forest provided the most stable performance as a parallel evaluator within the evaluated environment.
Description
Citation
APA
