Predicting international tourist arrivals in north sumatra using machine learning and google trends keyword selection
DOI:
https://doi.org/10.52465/joscex.v7i3.107Keywords:
Tourism forecasting , International tourist arrivals , Google trends, Random forest , Keyword selectionAbstract
This study examined the prediction of monthly international tourist arrivals in North Sumatra by combining historical visitation data with Google Trends keyword selection in a machine learning framework. Monthly data from January 2010 to December 2025 were analyzed. The proposed approach constructed temporal features from past arrivals and search-interest features from tourism-related keywords grouped into destination, travel-intent, and attraction-specific segments. A seasonal naive model was used as the baseline, while Random Forest and Gradient Boosting were applied as the main prediction models. The results showed that the historical-only Random Forest model achieved the best performance among the main forecasting scenarios and clearly outperformed the seasonal baseline. The inclusion of all Google Trends keywords did not improve prediction accuracy consistently. However, the destination segment provided more useful predictive information than the other keyword groups. Further keyword selection revealed that Bukit Lawang was the most robust single keyword, while a compact subset consisting of Danau Toba, Medan, Sumatra Utara, and Bukit Lawang produced the best test performance. These findings indicated that Google Trends improved forecasting accuracy only when relevant keywords were selected carefully.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Soft Computing Exploration

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
