Document Type : Research
Author
PhD in Soil Science; Soil and Water Research Expert; Tehran Agricultural and Natural Resources Research and Education Center; Agricultural Research, Education and Extension Organization (AREEO); Varamin; Iran.
Abstract
Objectives
Soil pH is a key chemical property that governs nutrient availability, microbial activity, and overall soil fertility, thereby exerting a significant influence on plant growth and productivity. Accurate spatial prediction of soil pH is essential for site-specific soil management and precision agriculture. The present study aims: 1) to prepare a digital map of the spatial distribution of soil pH in the 0–25 cm soil layer at Badr Watershed, southern Qorveh County, using environmental covariates derived from terrain attributes, remote sensing data, and geopedological information; 2) to evaluate the predictive performance of machine learning and statistical models, including Random Forest (RF), Artificial Neural Network (ANN), Decision Tree (DT), K-Nearest Neighbor (KNN), and Multiple Linear Regression (MLR) used for soil pH estimation; 3) to identify the most influential environmental variables controlling the spatial variability of soil pH; and 4) to determine the modeling approach capable of producing the most accurate and reliable digital soil pH maps to support precision agriculture and sustainable soil management.
Material and Methods
This study was conducted in Badr Watershed located in the southern part of Qorveh County, Kurdistan Province, Iran, covering an area of approximately 6,700 ha. The objective was to develop a digital soil pH map of the 0–25 cm soil layer by integrating environmental covariates with machine learning and statistical modeling approaches.
A geopedological map was first generated in a Geographic Information System (GIS) environment based on the Zink geopedological approach using geological and topographic information. Soil sampling locations were subsequently selected using the Latin Hypercube Sampling (LHS) technique to ensure representative coverage of the environmental variability across the study area. A total of 125 surface soil samples (0–25 cm) were collected and soil pH was determined in saturated soil paste using a calibrated pH meter following standard laboratory procedures.
A comprehensive set of environmental covariates was prepared to represent the soil-forming factors, including terrain attributes derived from a Digital Elevation Model (DEM), remote sensing indices extracted from Landsat 8 Imagery, and geopedological variables. These covariates were used as predictor variables for digital soil mapping.
Soil pH was modeled using five predictive approaches, including Random Forest (RF), Artificial Neural Network (ANN), Decision Tree (DT), K-Nearest Neighbor (KNN), and Multiple Linear Regression (MLR), implemented in the R statistical software environment. Model performance was evaluated using both 10-fold cross-validation and 5-fold random validation procedures. Predictive accuracy was assessed based on the coefficient of determination (R²), root mean square error (RMSE), and other relevant statistical performance indices. The resulting prediction maps were subsequently generated to characterize the spatial distribution of soil pH across the study area.
Results
The environmental covariates exhibited varying degrees of influence on the spatial prediction of soil pH. Variable importance analysis identified geomorphology as the most influential predictor, followed by watershed network base level, carbonate index, slope aspect, relative slope position, slope length factor (LS), and curvature index.
The the above-mentioned five statistical and machine learning models (namely, MLR, RF, ANN, DT, and KNN) were evaluated in terms of their predictive performance using both 10-fold cross-validation and 5-fold random validation procedures. The results revealed differences in prediction accuracy among the models evaluated.
Based on the 10-fold cross-validation results, the Multiple Linear Regression (MLR) model achieved the highest predictive performance, with a coefficient of determination (R²) of 0.698 and a root mean square error (RMSE) of 0.190, indicating its strong capability of soil pH estimation across the study area. In contrast, the 5-fold random validation results identified the K-Nearest Neighbor (KNN) model as the most accurate and precise approach for soil pH prediction.
Comparison of the validation approaches further indicated that ensemble or combined prediction strategies would generally outperform individual models. In this regard, integrating the prediction outputs from multiple models improved the overall reliability and spatial accuracy of the digital soil pH maps generated, resulting in a more robust characterization of soil pH variability across the study watershed.
Conclusion
This study demonstrated the effectiveness of integrating environmental covariates with statistical and machine learning approaches for digital mapping of soil pH in Badr Watershed. The results confirmed that terrain- and geopedology-related variables played fundamental roles in explaining the spatial variability of soil pH, highlighting the importance of incorporating multiple environmental factors into digital soil mapping frameworks.
Model evaluation revealed that predictive performance varied depending on the validation strategy adopted. This is evidenced by the fact that Multiple Linear Regression (MLR) provided the highest predictive accuracy under the 10-fold cross-validation scheme, whereas the K-Nearest Neighbor (KNN) model performed best under the 5-fold random validation approach. These findings indicate that no single model is universally superior and that model performance depends on both the characteristics of the validation procedure and the spatial structure of the dataset.
Furthermore, the superior performance of combined prediction approaches suggests that, compared to individual models, ensemble modeling can improve the reliability and spatial accuracy of digital soil pH maps. The soil pH maps thus generated provide valuable spatial information that can support site-specific soil management, precision agriculture, and sustainable land-use planning in the study area.
Overall, the integration of geopedological information, terrain derivatives, remote sensing data, and advanced predictive models yields a robust framework for digital soil pH mapping. Future studies may be recommended to investigate hybrid and ensemble machine learning techniques, incorporate additional environmental covariates and higher-resolution remote sensing data, and evaluate the transferability of the modeling framework proposed herein to other regions with different soil and environmental conditions.
Keywords