The Effect of Search Intent Data on Predicting Post-Crisis Daily Tourist Arrivals in Sri Lanka
Dept. of Computer Science & Engineering, University of Moratuwa, Sri Lanka
A multi-model forecasting framework for Sri Lanka's post-crisis daily tourist arrivals (2023–2025), integrating official arrivals data with Yandex search intent, localized Google Trends, exchange rates, and weather records across seven model architectures.
Read abstract
Accurately predicting daily international tourist arrivals is challenging for tourism-dependent economies like Sri Lanka, especially during its post-crisis recovery (2023–2025). Existing methods often ignore non-Google search engines, missing critical signals from key source markets like Russia, where Yandex dominates. This paper introduces a multi-model forecasting framework that integrates official daily arrivals with Yandex data, localized Google Trends, exchange rates, and weather records. We evaluate seven architectures: Pure SARIMA, SARIMAX, XGBoost, Random Forest, LSTM, Support Vector Regression (SVR), and a Sequential Hybrid, over a 550-day rolling test period. To prevent data leakage, feature selection utilizes Granger causality and cross-correlation analysis strictly on training data. The results demonstrate that SVR with an RBF kernel performs best, achieving a rolling MAPE of 8.94% and an RMSE of 736 arrivals/day. Crucially, Yandex queries for tours and flights emerge as the strongest exogenous predictor across all machine learning models, completely outperforming Russian Google Trends data, which proved uninformative. These findings confirm that aligning predictive features with a tourist market's preferred search engine significantly enhances forecasting accuracy, providing highly actionable insights for modern tourism intelligence systems.