Model Machine Learning untuk Rekomendasi Hotel Berdasarkan Profil dan Preferensi Wisatawan

Authors

Keywords:

klasifikasi, machine learning, preferensi wisatawan, random forest,, rekomendasi hotel, classification, hotel recommendation, tourist preferences

Abstract

Perkembangan platform pariwisata digital mendorong kebutuhan terhadap sistem rekomendasi hotel yang lebih personal dan berbasis data. Namun, sistem rekomendasi yang bergantung pada popularitas, rating, atau riwayat interaksi pengguna kurang optimal ketika data historis belum tersedia. Penelitian ini mengusulkan pendekatan classification-based recommendation untuk memprediksi kategori pilihan hotel wisatawan berdasarkan profil demografis, pola perjalanan, faktor ekonomi, dan preferensi fasilitas. Dataset diperoleh dari 290 responden dengan variabel meliputi usia, pendapatan, status pernikahan, pekerjaan, rekan perjalanan, lama liburan, besaran biaya liburan, jarak, fasilitas penginapan, dan layanan internet. Tiga algoritma klasifikasi digunakan, yaitu Decision Tree, Gaussian Naive Bayes, dan Random Forest. Evaluasi model dilakukan menggunakan metrik accuracy, precision, recall, F1-score, dan confusion matrix. Hasil penelitian menunjukkan bahwa Random Forest memperoleh performa terbaik dengan accuracy sebesar 68,97% dan F1-score sebesar 67,87%. Analisis feature importance menunjukkan bahwa lama liburan, pekerjaan, besaran biaya liburan, jarak, dan pendapatan menjadi faktor dominan dalam pemilihan hotel. Temuan ini menunjukkan bahwa pendekatan klasifikasi dapat digunakan sebagai dasar awal rekomendasi hotel yang lebih personal, terutama pada kondisi ketika data rating atau riwayat interaksi pengguna belum tersedia.

 

The development of digital tourism platforms has increased the need for more personalized and data-driven hotel recommendation systems. However, recommendation systems that rely on popularity, ratings, or user interaction history are less optimal when historical data are not yet available. This study proposes a classification-based recommendation approach to predict tourists’ hotel choice categories based on demographic profiles, travel patterns, economic factors, and facility preferences. The dataset was obtained from 290 respondents and included variables such as age, income, marital status, occupation, travel companions, length of vacation, vacation budget, distance, accommodation facilities, and internet services. Three classification algorithms were applied: Decision Tree, Gaussian Naive Bayes, and Random Forest. The models were evaluated using accuracy, precision, recall, F1-score, and confusion matrix. The results show that Random Forest achieved the best performance, with an accuracy of 68.97% and an F1-score of 67.87%. Feature importance analysis indicates that length of vacation, occupation, vacation budget, distance, and income are the dominant factors influencing hotel selection. These findings suggest that the classification-based approach can serve as an initial basis for more personalized hotel recommendations, particularly when rating data or user interaction history are not yet available.

Downloads

Published

2026-07-06

Issue

Section

Articles