ORIGINAL ARTICLE
Evaluating machine learning algorithms: The role of sample size and quality
More details
Hide details
1
Faculty of Geodesy and Cartography, Warsaw University of Technology, Pl. Politechniki 1, 00-661 Warsaw, Poland
These authors had equal contribution to this work
A - Research concept and design; B - Collection and/or assembly of data; C - Data analysis and interpretation; D - Writing the article; E - Critical revision of the article; F - Final approval of article
Submission date: 2026-02-26
Final revision date: 2026-06-12
Acceptance date: 2026-07-07
Publication date: 2026-07-27
Corresponding author
Przemysław Kupidura
Faculty of Geodesy and Cartography, Warsaw University of Technology, Pl. Politechniki 1, 00-661 Warsaw, Poland
Reports on Geodesy and Geoinformatics 2026;122:1-13
KEYWORDS
TOPICS
ABSTRACT
This study evaluates the performance of selected machine learning methods, Maximum Likelihood (MLC), Random Forest (RF), Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), and Artificial Neural Networks (ANN), for land use/land cover (LULC) classification using Sentinel-2 satellite imagery. Each algorithm was tested across multiple classification scenarios that systematically varied both training sample size and training sample quality. To emulate realistic imperfections in reference data, controlled levels of label noise were introduced into the training set, and the resulting changes in classification performance were analysed. In addition, each model’s susceptibility to overfitting was assessed by comparing performance on training data and independent test data. To reduce the influence of a single random draw of training and testing pixels, all sampling-based experiments were repeated 10 times using different random seed values, and the reported results were summarized using repeated-run statistics. This design enabled a comparable assessment of how classification accuracy, stability, and generalization depend on dataset size and quality. The results indicate that SVM was the most consistent and reliable method across the tested scenarios, achieving high classification accuracy over a wide range of training sample sizes and maintaining strong robustness under noisy labels. Compared with the other methods, SVM showed lower sensitivity to degraded training data and a smaller tendency to overfit, which makes it a strong baseline choice when reference data are limited or imperfect.
REFERENCES (38)
1.
Belgiu M., Drăguţ L. (2016). Random forest in remote sensing: A review of applications and future directions. ISPRS Journal of Photogrammetry and Remote Sensing. 114: 24–31-24–31. doi:10.1016/j.isprsjprs.2016.01.011.
2.
Bigdeli A., Maghsoudi A., Ghezelbash R. (2023). A comparative study of the XGBoost ensemble learning and multilayer perceptron in mineral prospectivity modeling: a case study of the Torud-Chahshirin belt, NE Iran. Earth Science Informatics. 17 (1): 483–499-483–499. doi:10.1007/s12145-023-01184-4.
3.
Boser B. E., Guyon I. M., Vapnik V. N. (1992). A training algorithm for optimal margin classifiers. Proceedings of the fifth annual workshop on Computational learning theory. 144–152-144–152. doi:10.1145/130385.130401.
4.
Budach L., Feuerpfeil M., Ihde N., Nathansen A., Noack N., Patzlaff H., Naumann F., Harmouch H. (2022). The effects of data quality on machine learning performance. arXiv preprint arXiv:2207.14529.
5.
Burkholder A., Warner T. A., Culp M., L, enberger R. (2011). Seasonal trends in separability of leaf reflectance spectra for Ailanthus altissima and four other tree species. Photogrammetric Engineering and Remote Sensing. 77 (8): 793–804-793–804. doi:10.14358/PERS.77.8.793.
6.
Chen T., Guestrin C. (2016). XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 785–794-785–794. doi:10.1145/2939672.2939785.
7.
Cohen J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement. 20 (1): 37–46-37–46. doi:10.1177/001316446002000104.
8.
Cracknell M. J., Reading A. M. (2014). Geological mapping using remote sensing data: A comparison of five machine learning algorithms, their response to variations in the spatial distribution of training data and the use of explicit spatial information. Computers and Geosciences. 63: 22–33-22–33. doi:10.1016/j.cageo.2013.10.008.
9.
Duro D. C., Franklin S. E., Dubé M. G. (2012). A comparison of pixel-based and object-based image analysis with selected machine learning algorithms for the classification of agricultural landscapes using SPOT-5 HRG imagery. Remote sensing of environment. 118: 259–272-259–272. doi:10.1016/j.rse.2011.11.020.
10.
Figueroa R. L., Zeng-Treitler Q., K, ula S., Ngo L. H. (2012). Predicting sample size required for classification performance. BMC Medical Informatics and Decision Making. 12 (1). doi:10.1186/1472-6947-12-8.
11.
Fu Y., Shen R., Song C., Dong J., Han W., Ye T., Yuan W. (2023). Exploring the effects of training samples on the accuracy of crop mapping with machine learning algorithm. Science of Remote Sensing. 7: 100081-100081. doi:10.1016/j.srs.2023.100081.
12.
Ghayour L., Neshat A., Paryani S., Shahabi H., Shirzadi A., Chen W., Al-Ansari N., Geertsema M., Pourmehdi Amiri M., Gholamnia M., Dou J., Ahmad A. (2021). Performance Evaluation of Sentinel-2 and Landsat 8 OLI Data for Land Cover/Use Classification Using a Comparison between Machine Learning Algorithms. Remote Sensing. 13 (7): 1349-1349. doi:10.3390/rs13071349.
13.
Halevy A., Norvig P., Pereira F. (2009). The Unreasonable Effectiveness of Data. IEEE Intelligent Systems. 24 (2): 8–12-8–12. doi:10.1109/mis.2009.36.
14.
H, D. J., Christen P., Kirielle N. (2021). F*: an interpretable transformation of the F-measure. Machine Learning. 110 (3): 451–456-451–456. doi:10.1007/s10994-021-05964-1.
15.
Koppaka R., Moh T. S. (2020). Machine Learning in Indian Crop Classification of Temporal Multi-Spectral Satellite Image. 2020 14th International Conference on Ubiquitous Information Management and Communication (IMCOM). 1–8-1–8. doi:10.1109/imcom48794.2020.9001718.
16.
Kupidura P., Kępa A., Krawczyk P. (2024). Comparative analysis of the performance of selected machine learning algorithms depending on the size of the training sample. Reports on Geodesy and Geoinformatics. 118 (1). doi:10.2478/rgg-2024-0015.
17.
Kupidura P., Niemyski S. (2024). Analysis of the effectiveness of selected machine learning algorithms in the classification of satellite image content depending on the size of the training sample. Teledetekcja Środowiska. 64.
18.
Labatut V., Cherifi H. (2012). Accuracy measures for the comparison of classifiers. arXiv preprint arXiv:1207.3790.
19.
Li X., Chen W., Cheng X., Wang L. (2016). A Comparison of Machine Learning Algorithms for Mapping of Complex Surface-Mined and Agricultural Landscapes Using ZiYuan-3 Stereo Satellite Imagery. Remote Sensing. 8 (6): 514-514. doi:10.3390/rs8060514.
20.
Liu J., Zuo Y., Wang N., Yuan F., Zhu X., Zhang L., Zhang J., Sun Y., Guo Z., Guo Y., Song X., Song C., Xu X. (2021). Comparative Analysis of Two Machine Learning Algorithms in Predicting Site-Level Net Ecosystem Exchange in Major Biomes. Remote Sensing. 13 (12): 2242-2242. doi:10.3390/rs13122242.
21.
Mansor N. S., Awang H., Malami S. T. S., Zolkafli A., Taiye M. A., Maulana H. (2024). Support Vector Machine for Satellite Images Classification Using Radial Basis Function Kernel Method. In: Computing and Informatics. 301–312-301–312. Springer Nature Singapore. doi:10.1007/978-981-99-9589-9_23.
22.
Maxwell A. E., Warner T. A. (2015). Differentiating mine-reclaimed grasslands from spectrally similar land cover using terrain variables and object-based machine learning classification. International Journal of Remote Sensing. 36 (17): 4384–4410-4384–4410. doi:10.1080/01431161.2015.1083632.
23.
Maxwell A. E., Warner T. A., Strager M. P., Pal M. (2014). Combining RapidEye Satellite Imagery and Lidar for Mapping of Mining and Mine Reclamation. Photogrammetric Engineering and Remote Sensing. 80 (2): 179–189-179–189. doi:10.14358/pers.80.2.179-189.
24.
Maxwell A., Strager M., Warner T., Zégre N., Yuill C. (2014). Comparison of NAIP orthophotography and RapidEye satellite imagery for mapping of mining and mine reclamation. GIScience and Remote Sensing. 51 (3): 301–320-301–320. doi:10.1080/15481603.2014.912874.
25.
Maxwell A., Warner T., Strager M., Conley J., Sharp A. (2015). Assessing machine-learning algorithms and image-and lidar-derived variables for GEOBIA classification of mining and mine reclamation. International Journal of Remote Sensing. 36 (4): 954–978-954–978. doi:10.1080/01431161.2014.1001086.
26.
Mousavinezhad M., Feizi A., Aalipour M. (2023). Performance Evaluation of Machine Learning Algorithms in Change Detection and Change Prediction of a Watershed’s Land Use and Land Cover. International Journal of Environmental Research. 17 (2). doi:10.1007/s41742-023-00518-w.
27.
Powers D. (2007). Evaluation: From precision, recall and F-factor to ROC, informedness, markedness and correlation.
28.
Ramezan C. A., Warner T. A., Maxwell A. E., Price B. S. (2021). Effects of Training Set Size on Supervised Machine-Learning Land-Cover Classification of Large-Area High-Resolution Remotely Sensed Data. Remote Sensing. 13 (3): 368-368. doi:10.3390/rs13030368.
29.
Raudys S., Jain A. (1991). Small sample size effects in statistical pattern recognition: recommendations for practitioners. IEEE Transactions on Pattern Analysis and Machine Intelligence. 13 (3): 252–264-252–264. doi:10.1109/34.75512.
30.
Richards J. A., Jia X. (2006). Remote Sensing Digital Image Analysis: An Introduction. Springer Berlin Heidelberg. doi:10.1007/3-540-29711-1.
31.
Seydi S. T., Kanani-Sadat Y., Hasanlou M., Sahraei R., Chanussot J., Amani M. (2022). Comparison of Machine Learning Algorithms for Flood Susceptibility Mapping. Remote Sensing. 15 (1): 192-192. doi:10.3390/rs15010192.
32.
Shang M., Wang S. X., Zhou Y., Du C. (2018). Effects of Training Samples and Classifiers on Classification of Landsat-8 Imagery. Journal of the Indian Society of Remote Sensing. 46 (9): 1333–1340-1333–1340. doi:10.1007/s12524-018-0777-z.
33.
Shih H. c., Stow D. A., Tsai Y. H. (2018). Guidance on and comparison of machine learning classifiers for Landsat-based land cover and land use mapping. International Journal of Remote Sensing. 40 (4): 1248–1274-1248–1274. doi:10.1080/01431161.2018.1524179.
34.
Sim J., Wright C. C. (2005). The Kappa Statistic in Reliability Studies: Use, Interpretation, and Sample Size Requirements. Physical Therapy. 85 (3): 257–268-257–268. doi:10.1093/ptj/85.3.257.
35.
Sobieraj J., Fernández M., Metelski D. (2022). A Comparison of Different Machine Learning Algorithms in the Classification of Impervious Surfaces: Case Study of the Housing Estate Fort Bema in Warsaw (Poland). Buildings. 12 (12): 2115-2115. doi:10.3390/buildings12122115.
36.
Volke M. I., Abarca-Del-Rio R. (2020). Comparison of machine learning classification algorithms for land cover change in a coastal area affected by the 2010 Earthquake and Tsunami in Chile. Natural Hazards and Earth System Sciences Discussions. 2020: 1–14-1–14. doi:10.5194/nhess-2020-41.
37.
Zhao Z., Islam F., Waseem L. A., Tariq A., Nawaz M., Islam I. U., Bibi T., Rehman N. U., Ahmad W., Aslam R. W., Raza D., Hatamleh W. A. (2024). Comparison of Three Machine Learning Algorithms Using Google Earth Engine for Land Use Land Cover Classification. Rangeland Ecology and Management. 92: 129–137-129–137. doi:10.1016/j.rama.2023.10.007.
38.
Zheng W., Jin M. (2020). The Effects of Class Imbalance and Training Data Size on Classifier Learning: An Empirical Study. SN Computer Science. 1 (2). doi:10.1007/s42979-020-0074-0.