An evaluation of Naive Bayes variants in content-based learning for spam filtering

13 years 9 months ago

Download alex.seewald.at

We describe an in-depth analysis of spam-ﬁltering performance of a simple Naive Bayes learner and two extended variants. A set of seven mailboxes comprising about 65,000 mails from seven diﬀerent users, as well as a representative snapshot of 25,000 mails which were received over 18 weeks by a single user, were used for evaluation. Our main motivation was to test whether two extended variants of Naive Bayes learning, SA-Train and CRM114, were superior to simple Naive Bayes learning, represented by SpamBayes. Surprisingly, we found that the performance of these systems was remarkably similar and that the extended systems have signiﬁcant weaknesses which are not apparent for the simpler Naive Bayes learner. The simpler Naive Bayes learner, SpamBayes, also oﬀers the most stable performance in that it deteriorates least over time. Overall, SpamBayes should be preferred over the more complex variants.

Alexander K. Seewald

Real-time Traffic

IDA 2007 | Information Technology | Naive Bayes | Simple Naive Bayes | Simpler Naive Bayes |

claim paper

Added	14 Dec 2010
Updated	14 Dec 2010
Type	Journal
Year	2007
Where	IDA
Authors	Alexander K. Seewald

Sciweavers

An evaluation of Naive Bayes variants in content-based learning for spam filtering

IDA 2007 | Information Technology | Naive Bayes | Simple Naive Bayes | Simpler Naive Bayes |

Explore & Download

Productivity Tools

Document Tools

Image Tools

Sciweavers