Vol. 336 No. 9 (2025)
DOI https://doi.org/10.18799/24131830/2025/9/4754
Efficiency and validity of cluster analysis of trace elements content in snow cover dust
Relevance. Clustering, as a method of data analysis, has found wide application in various fields of knowledge where classification of research objects is required. The search for algorithms that facilitate the most efficient use of the method is obvious. The success of forming a classification tree of hierarchical cluster analysis depends on the data standardization methods used. Aim. To conduct a comparative analysis of methods for standardizing the composition of chemical elements of snow dust for assessing the environmental hazard and validity of the results of hierarchical cluster analysis. Objects and methods. As an example, we used the microelement composition of the solid phase of snow in the city of Tyumen and background points more than 10 km away from the city. The content of pollutants in the snow cover reflects atmospheric air pollution. Using the example of analyzing the content of chemical elements in the solid phase of snow cover in the city of Tyumen, the simplest methods of preliminary data processing are substantiated in order to standardize them for subsequent statistical analysis. The paper considers four methods of data standardization in comparison with the original data. The effectiveness of clustering was assessed using the integral indicator of environmental pollution, and its validity – using the Kalinski–Harabash index. To confirm the main conclusions, the results are compared with data for the Tomsk region. Results. The paper shows a graphical display of geochemical spectra using different methods of data standardization, as well as an analysis of the differences in clustering results. To compare them, the data on the microelement composition of the snow cover in the Tomsk region were used. Conclusions. The “Weight” method of weights (%) turned out to be the most effective in graphically displaying the geochemical spectrum, allowing us to identify differences in the relative content of trace elements in the city and in background conditions. It was believed that the higher their values, the more effective the clustering; the control was the same indicators for snow cover in the Tomsk region, which turned out to be consistent with the indicators for Tyumen. It was established that standardization with a median and quantiles of 0.25 and 0.75 “Median” is most effective.
Keywords:
trace elements, snow dust, data standardization, cluster analysis, validity, air pollution


