In our study we presented an effective method for clustering of Web pages. From flat HTML files we extracted keywords, formed feature vectors as representation of Web pages and applied them to a clustering method. We took advantage of the Fuzzy C-Means clustering algorithm (FCM). We demonstrated an organized and schematic manner of data collection. Various categories of Web pages were retrieved from ODP (Open Directory Project) in order to create our datasets. The results of clustering proved that the method performs well for all datasets. Finally, we presented a comprehensive experimental study examining: the behavior of the algorithm for different input parameters, internal structure of datasets and classification experiments.
Machine Learning has assumed a prominent position in the plethora of design and analysis of intelligent systems. Learning is the holy grail of Machine Learning, and with rapidly growing complexity and the size of the constructed networks (the trend which is profoundly visible in deep learning architectures), the overwhelming computing is staggering. The return on investment clearly diminishes: even a very limited improvement in performance (commonly expressed as a classification rate or prediction error) does call for intensive computing because of learning a large number of parameters. The recent developments in green Artificial Intelligence (or better to say, green Machine Learning) has identified and emphasized a genuine need for a holistic multicriteria assessment of the design practices of Machine Learning architectures by involving computing overhead, interpretability, robustness, and identifying sound trade-offs present in these problems. We discuss a realization of green Machine Learning and advocate how Granular Computing contributes to the augmentation of the existing technology. In particular, some paradigms that exhibit a sound potential to support the sustainability of Machine Learning such as federated learning and transfer learning, are identified, critically evaluated, and cast into some general perspective.
Knowledge-based clustering algorithms can improve traditional clustering models by introducing domain knowledge to identify the underlying data structure. While there have been several approaches to clustering with the guidance of knowledge tidbits, most of them mainly focus on numeric knowledge without considering the uncertain nature of information. To capture the uncertainty of information, pure numeric knowledge tidbits are expanded to knowledge granules in this article. Then, two questions arise: how to obtain granular knowledge and how to use those knowledge granules in clustering. To the end, a novel knowledge extraction and granulation (KEG) method and a granular knowledge-based fuzzy clustering model are proposed in this study. First, inspired by the concept of natural neighbors, an automatic KEG is developed. In KEG, high-density points are filtered from the dataset and then merged with their natural neighbors to form several dense areas, i.e., granular knowledge. Furthermore, the granular knowledge expressed by interval or triangular numbers is leveraged into the clustering algorithm, which is the framework of fuzzy clustering with granular knowledge. To concretize this model into clustering algorithms, the classical fuzzy C-Means clustering algorithm has been selected to incorporate the granular knowledge produced by KEG. Then, the corresponding fuzzy C-Means clustering with interval knowledge granules (IKG-FCM) and triangular knowledge granules (TKG-FCM) are proposed. Experiments on synthetic and real-world datasets demonstrate that IKG-FCM and TKG-FCM always achieve better clustering performance with less time cost, especially on imbalanced data, compared with state-of-the-art algorithms.
The task of identifying native and foreign elements and rejecting foreign ones in the pattern recognition problem is discussed in this paper. Such the task is a nonstandard aspect of pattern recognition, which is rarely present in research. In this paper, ensembles of support vector machines solving two-classes and one-class problems are employed as classification tools and as basic tools for rejecting of foreign elements. Evaluation of quality of classification and rejection methods are proposed in the paper and finally some experiments are performed in order to illustrate acquainted terms and methods.
Effectiveness and clarity of software objects, their adherence to coding standards and programming habits of programmers are important features of overall quality of software systems. This paper proposes an approach towards a quantitative software quality assessment with respect to extensibility, reusability, clarity and efficiency. It exploits techniques of Computational Intelligence (CI) that are treated as a consortium of granular computing, neural networks and evolutionary techniques. In particular, we take advantage of self-organizing maps to gain a better insight into the data, and study genetic decision trees-a novel algorithmic framework to carry out classification of software objects with respect to their quality. Genetic classifiers serve as a "quality filter" for software objects. Using these classifiers, a system manager can predict quality of software objects and identify low quality objects for review and possible revision. The approach is applied to an object-oriented visualization-based software system for biomedical data analysis.
In this paper, we develop a general framework of a granular representation of ECG signals. The crux of the approach lies in the development and ongoing processing realized in the setting of information granules-fuzzy sets. They serve as basic conceptual and semantically meaningful entities using which we describe signals and build their models (such as various predictive schemes or classifiers). A comprehensive two-phase scheme of the design of the information granules is proposed and described. At the first phase, we discuss the temporal granulation through a series of temporal windows (granular windows) and an aggregation of the values of signal by means of fuzzy sets. To address this issue, offered is a detailed method of building a fuzzy set based on numeric data and a certain optimization criterion that strikes a balance between the highest experimental relevance of the fuzzy set supported by numeric data and its substantial specificity. At the next phase of the granular design, a collection of information granules is further summarized with the use of fuzzy clustering (Fuzzy C-Means). The resulting prototypes (centroids) formed by this grouping process serve as elements of the granular vocabulary. We discuss ways of using these vocabularies in the knowledge-based representation, modeling, and classification of ECG beats.