URL-based phishing detection using XGBoost with engineered features

Indonesian Journal of Electrical Engineering and Computer Science

URL-based phishing detection using XGBoost with engineered features

Abstract

URL-based phishing involves fake uniform resource locators (URLs) created by attackers to trick users into believing they are visiting a legitimate website and thereby steal their confidential information. While several powerful machine learning (ML) and deep learning (DL) studies exist to detect phishing, they still face limitations. Many studies rely on third-party intervention to extract features, which introduces delays that make them unsuitable for fast detection. Another limitation is that existing studies often use small datasets, and traditional features hinder models' ability to learn new phishing techniques, resulting in poor generalization. Therefore, developing new features is crucial to ensure that anti-phishing tools can keep pace with evolving phishing tactics. In addition, the existing studies do not report detection time, which is important for fast detection, and reduces methodological clarity. This paper aims to address these limitations by applying a neural network model and traditional ML classification algorithms to support browser-based phishing detection that balances high accuracy with fast detection. Our XGBoost model achieved 98% accuracy on the test set, utilizing 40 third-party-independent features. Additionally, we achieved an average response time of 0.026225 seconds and an average computation time of 1.6977×10⁻6 seconds per URL, which demonstrates competitive speed. We provided a table of features from recent studies, together with their documented sources, to support future research and analyze key URL-based features.

Discover Our Library

Embark on a journey through our expansive collection of articles and let curiosity lead your path to innovation.

Explore Now
Library 3D Ilustration