A novel fuzzy adaptive knowledge-based inference neural network (FAKINN) is proposed in this study. Conventional fuzzy cluster-based neural networks (FCBNNs) suffer from the challenge of a direct extraction of fuzzy rules that can capture and represent the interclass heterogeneity and intraclass homogeneity when the data possess complex structures. Moreover, the capability of the cluster-based rule generator in FCBNNs may decrease with the increase of data dimensionality. These drawbacks impede the generation of desired fuzzy rules, and affect the inference results depending on the fuzzy rules, thereby limiting their generalization ability. To address these drawbacks, an adaptive knowledge generator (AKG), consisting of the observation paradigm (OP) and clustering strategy (CS), is effectively designed to improve the generalization ability in FAKINN. The OP distills the characteristic information (CI) from data to highlight the homogeneity and heterogeneity of objects, and the CS, viz., the weighted condition-driven fuzzy clustering method (WCFCM), is proposed to summarize the CI to construct fuzzy rules. Moreover, the feedback between the OP and CS can control the dimensionality of CI, which endows FAKINN with the potential to tackle high-dimensional data. The main originality of the study focuses on the AKG and WCFCM that are proposed to develop the structural design methodology of FNNs. The performance of FAKINN is evaluated on various benchmarks with 27 comparative methods, and two real-world problems are adopted to validate its effectiveness. Experimental results show that FAKINN outperforms the comparison methods.
Querying and reporting from large volumes of structured, semistructured, and unstructured data often requires some flexibility. This flexibility provided by fuzzy sets allows for categorization of the surrounding world in a flexible, human-mind-like manner. Apache Hive is a data warehousing framework working on top of the Hadoop platform for big data processing. Hive allows executing queries and aggregating and analyzing data stored in Hadoop distributed file system and other repositories. Hive responds to the current needs for efficient big data warehousing, which is impossible with traditional data warehouses due to their rigid nature. This article presents the FuzzyHive library that extends the Hive framework with fuzzy sets based techniques for querying, analyzing, and reporting on big data warehouses. We formalize the fuzzy techniques used while operating on Hive-based data warehouses (including fuzzy filtering on dimensional attributes, projection with fuzzy transformation, fuzzy grouping, and joining). We also show how we embedded these operations in Hive query language, which was not studied so far. Such extensions make big data warehousing more flexible and contribute to the portfolio of tools used by the community of people working with fuzzy sets and data analysis. The FuzzyHive library complements the spectrum of available solutions for fuzzy data processing and querying in large datasets. We investigate Hive fuzzy querying performance, effectiveness, and scalability for various data storage formats (text, Avro, and Parquet). Our experiments demonstrate that the proposed extensions introduce more elasticity and are also efficient for big data warehousing, which is the first such kind of solution for this environment.
In this study, we propose a cluster-oriented development of fuzzy models. An overall design process is focused on an efficient usage of fuzzy clustering, Fuzzy C-Means (FCM), in particular, to form information granules-clusters that are used in the construction of the fuzzy model. Fuzzy models are regarded as mappings from information granules expressed in the input and output spaces. This position motivates us to look at the development of the models through the perspective of the construction and efficient usage of information granules. This study directly associates fuzzy clustering with fuzzy modeling both in terms of conceptual and algorithmic linkages. The augmented FCM method is formed predominantly for modeling purposes so that a balance between the structural content present in the input and output spaces is achieved and this way the performance of the resulting fuzzy model is optimized. It is shown that the cluster-oriented modeling gives rise to the Mamdani-like fuzzy rules and a zero-order Takagi-Sugeno model (under a certain decoding scheme). We identify an interesting and direct linkage between the developed fuzzy models and a fundamental idea of encoding-decoding (or granulation-degranulation) encountered in processing fuzzy sets and Granular Computing, in general. Furthermore, refinements of zero-order fuzzy models are investigated leading to first-order fuzzy models with linear functions standing in the conclusions of the rules. A series of experiments is reported where we used synthetic and real-world data using which an issue of generalization capabilities is elaborated in detail.
Different Earth observation resources (EORs) [e.g., satellites, airships, and unmanned aerial vehicles (UAVs)] are usually managed by different organization sub-planners, which lack interactions and cooperation among one another. Such independent resource operations are no longer efficient to meet diverse and vast observation requests, especially in emergency situations, such as earthquakes, flooding, and forest fire disasters. This paper addresses the issue of coordinated planning of heterogeneous EORs, including satellites, airships, and UAVs. A hierarchical coordinated planning architecture is proposed to integrate heterogeneous EORs for the construction of a distributed and loosely coupled Earth observation system. The architecture comprises four component categories, namely, observation resource, sub-planner, coordination, and information management. Moreover, we propose two task assignment algorithms to coordinate and allocate observation tasks to sub-planners. The first algorithm is a highest-weight-first-allocated algorithm, and the second is a tabu-list-based simulated annealing (SA-TL) algorithm. Experiments and comparative studies demonstrate the efficiency of the coordinated planning architecture and SA-TL algorithm. We also show that the system responds dynamically to unexpected situations through effective disturbance-handling mechanisms.
Mining through vast arrays of heterogeneous data (no matter where they come from) is a challenging and rewarding pursuit. To make the findings meaningful, the methods of data mining need to be presented to the end-user in a highly intelligible manner. The role of information granules is to cast data mining in the setting of some meaningful nonnumeric entities exhibiting a well-defined semantics. We propose a concept of a linguistic selector and view it as a focal element of data mining. The linguistic selectors are weighted and-combinations of the individual linguistic terms: fuzzy sets or fuzzy relations defined over features (variables) existing in the problem. When operating on a database, they are matched with the individual records and return a certain level of compatibility (matching). The fundamental aspect of such linguistic selectors deals with their relevance in terms of their semantics and statistical relevance. We quantify these two features and propose an optimization problem leading to the design of the meaningful linguistic selectors.