Authors S IliyazDepartment of Computer Science Engineering (Cyber Security), GATES Institute of Technology, Gooty, Andhra Pradesh, IndiaS KaveriDepartment of Computer Science Engineering (Cyber Security), GATES Institute of Technology, Gooty, Andhra Pradesh, IndiaG Vishnu MedhasDepartment of Computer Science Engineering (Cyber Security), GATES Institute of Technology, Gooty, Andhra Pradesh, IndiaV AnjaliDepartment of Computer Science Engineering (Cyber Security), GATES Institute of Technology, Gooty, Andhra Pradesh, IndiaG Pavan Sai ReddyDepartment of Computer Science Engineering (Cyber Security), GATES Institute of Technology, Gooty, Andhra Pradesh, IndiaA NavadeepDepartment of Computer Science Engineering (Cyber Security), GATES Institute of Technology, Gooty, Andhra Pradesh, IndiaD SharfuddinDepartment of Computer Science Engineering (Cyber Security), GATES Institute of Technology, Gooty, Andhra Pradesh, India Abstract In detecting malicious websites, a common approach is the use of blacklists which are not exhaustive in them-selves and are unable to generalize to new malicious sites. Detecting newly encountered malicious websites automatically will help reduce the vulnerability to this form of attack. In this study, we explored the use of ten machine learning models to classify malicious websites based on lexical features and understand how they generalize across datasets. Specifically, we trained, validated, and tested these models on different sets of datasets and then carried out a cross-datasets analysis. From our analysis, we found that K-Nearest Neighbour is the only model that performs consistently high across data. Other models such as Random Forest, Decision Trees, Logistic Regression, and Support Vector Machines also consistently outperform a baseline model of predicting every link as malicious across all metrics and datasets. Also, we found no evidence that any subset of lexical features generalizes across models or datasets. This research should be relevant to cybersecurity professionals and academic researchers as it could form the basis for real-life detection systems or further research work. Keywords Lexical features machine learning malicious URLs Citation of this Article . Licence Copyright (c) 2026 Current Journal of Engineering and Science Research. This work is licensed under a Creative Commons Attribution Non Commercial 4.0 International Licence. References .