The learning abilities and high transparency are the two important and highly desirable features of any model of software quality. The transparency and user-centricity of quantitative models of software engineering are of paramount relevancy as they help us gain a better and more comprehensive insight into the revealed relationships characteristic to software quality and software processes. In this study, we are concerned with logic-driven architectures of logic models based on fuzzy multiplexers (fMUXs). Those constructs exhibit a clear and modular topology whose interpretation gives rise to a collection of straightforward logic expressions. The design of the logic models is based on the genetic optimization and genetic algorithms, in particular. Through the prudent usage of this optimization framework, we address the issues of structural and parametric optimization of the logic models. Experimental studies exploit software data that relates software metrics (measures) to the number of modifications made to software modules.Request access from your librarian to read this chapter's full text.
Hesitant fuzzy preference relation (HFPR) is a valid tool to describe the hesitation, ambiguity and uncertainty of decision makers. As a crucial criterion to ensure the rationality of preferences and final decision results, the consistency of preference relations is a valuable research topic. In this study, the additive consistency of HFPRs is investigated. Two kinds of consistency, completely additive and weakly additive, for HFPRs are introduced. Some 0–1 mixed programming models and simple algebraic operations are developed to detect the additive consistency type for HFPRs. Further, the priority weights of a consistent HFPR can be derived by constructing and solving two linear programming models. If an HFPR is identified to be inconsistent, we present a straightforward method to rectify inconsistency. Then, an integrated algorithm is proposed to ascertain the additive consistency type, improve consistency and rank alternatives. Finally, the applicability and validity of this proposal are verified through a case study, discussion and comparative analysis with the existing methods.
Granular Computing has emerged as a unified and coherent framework of designing, processing, and interpretation of information granules. Information granules are formalized within various frameworks such as sets (interval mathematics), fuzzy sets, rough sets, shadowed sets, probabilities (probability density functions), to name several the most visible approaches. In spite of the apparent diversity of the existing formalisms, there are some underlying commonalities articulated in terms of the fundamentals, algorithmic developments and ensuing application domains. In this study, we introduce two pivotal concepts: a principle of justifiable granularity and a method of an optimal information allocation where information granularity is regarded as an important design asset. We show that these two concepts are relevant to various formal setups of information granularity and offer constructs supporting the design of information granules and their processing. A suite of applied studies is focused on knowledge management in which case we identify several key categories of schemes present there.
As an emerging computing paradigm of information processing, Granular Computing exhibits great potential in human-centric decision problems such as feature selection and feature extraction, pattern recognition and knowledge discovery. Optimization plays an important role in these areas. The optimization problems arising in Granular Computing area are called granular optimization problems in which information granules are treated as information processing units and therefore granules denote the related solutions. Particle swarm optimization (PSO) has been demonstrated to be a very competitive algorithm in solving global optimization problems. In this paper, we develop a novel PSO variant called granular PSO to solve problems of granular optimization. Each granule in this study is expressed as a multi-dimension hyper-box with each coordinate being described by an interval. In the proposed granular PSO, the velocity and position of a particle is represented by intervals rather than single numerical values. The velocity and position update strategy is modified accordingly. In granular PSO, the solution space search behavior of a particle is realized in granule-to-granule manner rather than point-to-point format. We provide experimental simulations to demonstrate the effectiveness of the proposed granular PSO algorithm.
We have been witnessing a flurry of unprecedented developments in the technologies of data mining and their diverse applications. One can point to flagship projects of high visibility such as smart cities in which a variety of IT technologies, smart sensors, and human-centric interfaces play a pivotal role. In all of these applications, floods of data need to be stored, processed, and interpreted. The Internet of Things has shaped the current landscape of IT and influenced relationships between the technology and users. The advantages of all these projects exhibit far reaching and indisputable benefits in numerous areas of management, marketing, engineering, and manufacturing, including new manufacturing paradigms such as the Industry 4.0 initiative. Facial recognition is a new and booming technology with tangible benefits in biometrics, advertising, medical areas, and ways of identifying missing people. Innovative facial recognition by Facebook has opened new directions and unleashed new opportunities. The volume of surveillance cameras alone (in the United States itself amounting to 30 million; Vlahos, 2009) clearly points at the enormous volume of data. This technology has been designed mostly using white, Caucasian subjects; there is also a need to expand so that the wide range of global diversity is included. Data Mining and Knowledge Discovery serve as umbrella terms that host a truly remarkable spectrum of applications in all areas of human endeavors. By the same token, which is not surprising at all, the technology of data mining and knowledge discovery acts as a double-edged sword. The controversies around the use of the data and highly visible cases of violating data privacy have attracted public attention; one can refer to Cambridge Analytica or Epic Games. No doubt, data handling ethics scenarios are a legal and political minefield that calls for delicate balancing mechanisms between achieving benefits of data mining and preventing unethical practices. The lack of data mining ethics in various organizations at different levels has become a highly contentious issue. There are collections of different legal frameworks existing in various countries. What makes the situation worse, sometimes there is a lack of full understanding of possible implications of the usage of the rapidly progressing technologies. The academic community has been cognizant of these issues from the outset of emergence of the concept of knowledge discovery and technologies of data mining by proactively studying a variety of countermeasures including such approaches as adversarial learning or a suite of diverse privacy mechanisms, among others. One can point here to interesting studies and review materials on privacy preserving in data mining (Cuzzocrea, 2017) (Wang, Luo, Zhao, & Le, 2009), security (Xu, Jiang, Wang, Yuan, & Ren, 2014), legal aspects (Carmichael, Stalla-Bourdillon, & Staab, 2016), and general considerations on ethics and technology (Munoz, 2004). WIREs Data Mining and Knowledge Discovery was designed to be a comprehensive interdisciplinary resource on all subjects within the umbrella of Data Mining and Knowledge Discovery, including new types of technologies like facial recognition and smart cities. As the Editor-in-Chief, I am fully aware of the drawbacks of these technologies and take all necessary measures to diligently cope with this subject and bring ethical issues into the picture. We regard comprehensive discussions on new, potentially controversial technologies and their ethical implications as a mission equally important as the dissemination of high-quality technical knowledge about mining data. As a matter of fact, the journal has a topic dedicated to ethics in data mining with subtopics on fairness, privacy, and legal issues. We actively solicit articles and encourage article proposals on these important and timely topics. We are confident that with our rigorous editorial and peer review practices, we can provide a truly comprehensive resource on the technologies of Data Mining and Knowledge Discovery, the benefits, and the drawbacks.
In recent years, image processing in a Euclidean domain has been well studied. Practical problems in computer vision and geometric modeling involve image data defined in irregular domains, which can be modeled by huge graphs. In this paper, a wavelet frame-based fuzzy C -means (FCM) algorithm for segmenting images on graphs is presented. To enhance its robustness, images on graphs are first filtered by using spatial information. Since a real image usually exhibits sparse approximation under a tight wavelet frame system, feature spaces of images on graphs can be obtained. Combining the original and filtered feature sets, this paper uses the FCM algorithm for segmentation of images on graphs contaminated by noise of different intensities. Finally, some supporting numerical experiments and comparison with other FCM-related algorithms are provided. Experimental results reported for synthetic and real images on graphs demonstrate that the proposed algorithm is effective and efficient, and has a better ability for segmentation of images on graphs than other improved FCM algorithms existing in the literature. The approach can effectively remove noise and retain feature details of images on graphs. It offers a new avenue for segmenting images in irregular domains.
The job-shop scheduling problem (JSP) is NP hard, which has very important practical significance. Because of many uncontrollable factors, such as machine delay or human factors, it is difficult to use a single real-number to express the processing and completion time of the jobs. JSP with fuzzy processing time and completion time (FJSP) can model the scheduling more comprehensively, which benefits from the developments of fuzzy sets. Fuzzy relative entropy leads to a method that can evaluate the quality of a feasible solution following the comparison between the actual value and the ideal value (the due date). Therefore, the multiobjective FJSP can be transformed into a single-objective optimization problem and solved by a hybrid adaptive differential evolution (HADE) algorithm. The maximum completion time, the total delay time, and the total energy consumption of jobs will be considered. HADE adopts a mutation strategy based on DE-current-to-best. Its parameters (CR and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">F</i> ) are all made adaptive and normally distributed. The new individuals are selected according to the fitness value (FRE) obtained from a population consisting of <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">N</i> parents and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">N</i> children in HADE. The algorithm is analyzed from different viewpoints. As the experimental results demonstrate, the performance of the HADE algorithm is better than those of some other state-of-the-art algorithms (namely, ant colony optimization, artificial bee colony, and particle swarm optimization).
A novel approach based on supervised hierarchical clustering is developed with the purpose of discovering structure in data where labels are provided. Labels can come in the form of discrete-valued class labels or continuous-valued output variables to aid hierarchical clustering in discovering the structure and the number of clusters, in particular. In the proposed method, Clusters are linked together if their discrete-valued labels are the same or in the case of continuous output variables if their outputs are similar. Similarity within a cluster in the continuous case is expressed by a measure of internal cluster dispersion. Several experiments on synthetic data with discrete-valued class labels are conducted to demonstrate the algorithm's ability to discover class or data structure.