Model Machine Learning untuk Rekomendasi Hotel Berdasarkan Profil dan Preferensi Wisatawan
Keywords:
klasifikasi, machine learning, preferensi wisatawan, random forest,, rekomendasi hotel, classification, hotel recommendation, tourist preferencesAbstract
Perkembangan platform pariwisata digital mendorong kebutuhan terhadap sistem rekomendasi hotel yang lebih personal dan berbasis data. Namun, sistem rekomendasi yang bergantung pada popularitas, rating, atau riwayat interaksi pengguna kurang optimal ketika data historis belum tersedia. Penelitian ini mengusulkan pendekatan classification-based recommendation untuk memprediksi kategori pilihan hotel wisatawan berdasarkan profil demografis, pola perjalanan, faktor ekonomi, dan preferensi fasilitas. Dataset diperoleh dari 290 responden dengan variabel meliputi usia, pendapatan, status pernikahan, pekerjaan, rekan perjalanan, lama liburan, besaran biaya liburan, jarak, fasilitas penginapan, dan layanan internet. Tiga algoritma klasifikasi digunakan, yaitu Decision Tree, Gaussian Naive Bayes, dan Random Forest. Evaluasi model dilakukan menggunakan metrik accuracy, precision, recall, F1-score, dan confusion matrix. Hasil penelitian menunjukkan bahwa Random Forest memperoleh performa terbaik dengan accuracy sebesar 68,97% dan F1-score sebesar 67,87%. Analisis feature importance menunjukkan bahwa lama liburan, pekerjaan, besaran biaya liburan, jarak, dan pendapatan menjadi faktor dominan dalam pemilihan hotel. Temuan ini menunjukkan bahwa pendekatan klasifikasi dapat digunakan sebagai dasar awal rekomendasi hotel yang lebih personal, terutama pada kondisi ketika data rating atau riwayat interaksi pengguna belum tersedia.
The development of digital tourism platforms has increased the need for more personalized and data-driven hotel recommendation systems. However, recommendation systems that rely on popularity, ratings, or user interaction history are less optimal when historical data are not yet available. This study proposes a classification-based recommendation approach to predict tourists’ hotel choice categories based on demographic profiles, travel patterns, economic factors, and facility preferences. The dataset was obtained from 290 respondents and included variables such as age, income, marital status, occupation, travel companions, length of vacation, vacation budget, distance, accommodation facilities, and internet services. Three classification algorithms were applied: Decision Tree, Gaussian Naive Bayes, and Random Forest. The models were evaluated using accuracy, precision, recall, F1-score, and confusion matrix. The results show that Random Forest achieved the best performance, with an accuracy of 68.97% and an F1-score of 67.87%. Feature importance analysis indicates that length of vacation, occupation, vacation budget, distance, and income are the dominant factors influencing hotel selection. These findings suggest that the classification-based approach can serve as an initial basis for more personalized hotel recommendations, particularly when rating data or user interaction history are not yet available.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Candra Agustina, Angga Ardiansyah, Raja Sabarudin, Lisnawanty Lisnawanty, Eka Rahmawati

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish manuscripts through IJCIT agree to the following provisions:
1. The copyright holder of the article is the author.
2. The author grants the right to publish the scientific article to IJCIT as the first publisher. At the same time, the author grants permission/license regarding the Creative Commons Attribution License to other parties to distribute the article.
3. Non-exclusivity matters in the distribution of the Journal, namely the publication of the author's scientific article can be agreed separately (for example, a request to be included in the institutional library or published as a book) then adjust the author as one of the parties and IJCIT as the first publisher.
4. Authors can publish articles online before and during the manuscript submission process (for example, in the Repository or on the website of the organization/institution), as this can encourage the creation of estimates and exchange of citations.
5. Manuscripts and related materials published through this Journal are distributed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA).



