526 publications from this institution
The widespread use of machine learning algorithms and the high level of expertise required to utilize them have fuelled the demand for solutions that can be used by non-experts. One of the main challenges non-experts face in applying machine learning to new problems is algorithm selection - the identification of the algorithm(s) that will deliver top performance for a given dataset, task, and evaluation measure. We present AutoGRD, a novel meta-learning approach for algorithm recommendation. AutoGRD first represents datasets as graphs and then extracts their latent representation that is used to train a ranking meta-model capable of accurately recommending top-performing algorithms for previously unseen datasets. We evaluate our approach on 250 datasets and demonstrate its effectiveness both for classification and regression tasks. AutoGRD outperforms state-of-the-art meta-learning and Bayesian methods.
No abstract is provided for this article.
Recommender systems provide consumers with ratings of items. These ratings are based on a set of ratings that were obtained from a wide scope of users. Predicting the ratings can be formulated as a regression problem. Ensemble regression methods are effective tools that improve the results of simple regression algorithms by iteratively applying the simple algorithm to a diverse set of inputs. The present paper describes a simple and effective ensemble regressor for the prediction of missing ratings in recommender systems. The ensemble method is an adaptation of the AdaBoost regression algorithm for recommendation tasks. In all iterations, interpolation weights for all nearest neighbors are simultaneously derived by minimizing the root mean squared error. From iteration to iteration instances that are hard to predict are reinforced by manipulating their weights in the goal function that needs to be minimized. The experimental evaluation demonstrates that the ensemble methodology significantly improves the predictive performance of single neighborhood-based collaborative filtering.
In privacy-preserving data mining (PPDM), a widely used method for achieving data mining goals while preserving privacy is based on k-anonymity. This method, which protects subject-specific sensitive data by anonymizing it before it is released for data mining, demands that every tuple in the released table should be indistinguishable from no fewer than k subjects. The most common approach for achieving compliance with k-anonymity is to replace certain values with less specific but semantically consistent values. In this paper we propose a different approach for achieving k-anonymity by partitioning the original dataset into several projections such that each one of them adheres to k-anonymity. Moreover, any attempt to rejoin the projections, results in a table that still complies with k-anonymity. A classifier is trained on each projection and subsequently, an unlabelled instance is classified by combining the classifications of all classifiers. Guided by classification accuracy and k-anonymity constraints, the proposed data mining privacy by decomposition (DMPD) algorithm uses a genetic algorithm to search for optimal feature set partitioning. Ten separate datasets were evaluated with DMPD in order to compare its classification performance with other k-anonymity-based methods. The results suggest that DMPD performs better than existing k-anonymity-based algorithms and there is no necessity for applying domain dependent knowledge. Using multiobjective optimization methods, we also examine the tradeoff between the two conflicting objectives in PPDM: privacy and predictive performance.
No abstract is provided for this article.
Many real life problems are characterized by the structure of data derived from multiple sensors. The sensors may be independent, yet their information considers the same entities. Thus, there is a need to efficiently use the information rendered by numerous datasets emanating from different sensors. A novel methodology to deal with such problems is suggested in this work. Measures for evaluating probabilistic classification are used in a new efficient voting approach called "selective voting", which is designed to combine the classification of the models (sensor fusion). Using "selective voting", the number of sensors is decreased significantly while the performance of the integrated model's classification is increased. This method is compared to other methods designed for combining multiple models as well as demonstrated on a real-life problem from the field of human resources.
No abstract is provided for this article.
No abstract is provided for this article.
Substantial electronically stored textual data such as clinical narratives reports often need to be retrieved to find relevant information for clinical and research purposes. The context of negation, a negative finding, is of special importance, since many of the most frequently described findings are such. Hence, when searching free-text narratives for patients with a certain medical condition, if negation is not taken into account, many of the documents retrieved were irrelevant. We present a new cascaded pattern learning method for automatic identification of negative context in clinical narratives reports. Studying the training corpuses, the classification errors and patterns selected by the classifier, we noticed that it is possible to create a more powerful ensemble structure than the structure obtained from general-purpose ensemble method (such as Adaboost). We compare the new algorithm to previous methods proposed for the same task of similar medical narratives, and show its advantages: accuracy improvement compared to other machine learning methods, and much faster than manual knowledge engineering techniques with matching accuracy
Decision forests, such as random forest (RF) are widely used for tabular data, mainly due to their predictive performance and ease of usage. However, given that the forest’s trees may produce contradictory predictions for a certain sample, the usage of decision forests in applications that involve decision-making necessitates further reliability assessment of the predictions for generating a trustworthy combined prediction. Model distillation using a born-again tree is a common approach for converting a decision forest into a single decision tree (DT) while preserving the forest’s predictive performance and supporting decision-makers. In this paper, we introduce PnT (Path in Tree), a novel approach that learns a path-based encoding from a decision forest. PnT applies an iterative algorithm, where in each iteration, a batch of trees is trained and then used to identify informative paths. These paths are then encoded and utilized for PnT-DT, an approach for producing a contradiction-free born-again DT. We also show that PnT can be leveraged for PnT-RF, a born-again forest approach, capable of improving the predictive performance of a plain decision forest. We evaluate PnT-DT and PnT-RF on 40 classification datasets and demonstrate that both PnT-DT and PnT-RF significantly outperform existing state-of-the-art (SOTA) born-again DT and decision forest methods in terms of predictive performance.
No abstract is provided for this article.
The formation of new malwares every day poses a significant challenge to anti-virus vendors since antivirus tools, using manually crafted signatures, are only capable of identifying known malware instances and their relatively similar variants. To identify new and unknown malwares for updating their anti-virus signature repository, anti-virus vendors must daily collect new, suspicious files that need to be analyzed manually by information security experts who then label them as malware or benign. Analyzing suspected files is a time-consuming task and it is impossible to manually analyze all of them. Consequently, anti-virus vendors use machine learning algorithms and heuristics in order to reduce the number of suspect files that must be inspected manually. These techniques, however, lack an essential element – they cannot be daily updated. In this work we introduce a solution for this updatability gap. We present an active learning (AL) framework and introduce two new AL methods that will assist anti-virus vendors to focus their analytical efforts by acquiring those files that are most probably malicious. Those new AL methods are designed and oriented towards new malware acquisition. To test the capability of our methods for acquiring new malwares from a stream of unknown files, we conducted a series of experiments over a ten-day period. A comparison of our methods to existing high performance AL methods and to random selection, which is the naïve method, indicates that the AL methods outperformed random selection for all performance measures. Our AL methods outperformed existing AL method in two respects, both related to the number of new malwares acquired daily, the core measure in this study. First, our best performing AL method, termed “Exploitation”, acquired on the 9th day of the experiment about 2.6 times more malwares than the existing AL method and 7.8 more times than the random selection. Secondly, while the existing AL method showed a decrease in the number of new malwares acquired over 10days, our AL methods showed an increase and a daily improvement in the number of new malwares acquired. Both results point towards increased efficiency that can possibly assist anti-virus vendors.
No abstract is provided for this article.
Online cancer communities help members support one another, provide new perspectives about living with cancer, normalize experiences, and reduce isolation. The American Cancer Society's 166000-member Cancer Survivors Network (CSN) is the largest online peer support community for cancer patients, survivors, and caregivers. Sentiment analysis and topic modeling were applied to CSN breast and colorectal cancer discussion posts from 2005 to 2010 to examine how sentiment change of thread initiators, a measure of social support, varies by discussion topic. The support provided in CSN is highest for medical, lifestyle, and treatment issues. Threads related to 1) treatments and side effects, surgery, mastectomy and reconstruction, and decision making for breast cancer, 2) lung scans, and 3) treatment drugs in colon cancer initiate with high negative sentiment and produce high average sentiment change. Using text mining tools to assess sentiment, sentiment change, and thread topics provides new insights that community managers can use to facilitate member interactions and enhance support outcomes.
Most recommender systems, such as collaborative filtering, cannot provide personalized recommendations until a user profile has been created. This is known as the new user cold-start problem. Several systems try to learn the new users' profiles as part of the sign up process by asking them to provide feedback regarding several items. We present a new, anytime preferences elicitation method that uses the idea of pairwise comparison between items. Our method uses a lazy decision tree, with pairwise comparisons at the decision nodes. Based on the user's response to a certain comparison, we select on-the-fly what pairwise comparison should next be asked. A comparative field study has been conducted to examine the suitability of the proposed method for eliciting the user's initial profile. The results indicate that the proposed pairwise approach provides more accurate recommendations than existing methods and requires less effort when signing up newcomers.
Many e-commerce sites use recommender systems, which suggest products that consumers may want to purchase in order to increase site revenue. Though recommender systems have achieved great success, they have not reached their full potential. Most current systems share a common weakness: they fail to take into account dynamic properties of the offering which could dramatically improve the effectiveness of a recommendation; these characteristics include the product price, promotion indication, and seller's reputation. Particularly, in a multi-seller platform (e.g., eBay, Amazon), where competing firms sell products differentiated mainly by the seller's reputation and product price, modeling consumer's sensitivity to these dynamic properties and incorporating it into a recommender system will optimize sellers’ revenue and market penetration. In this research, we introduce a novel approach for a personal price aware multi-seller recommender system (PMSRS) which implicitly models a consumer's willingness to pay (WTP) for a specific product, taking into account discount indication and seller reputation, and incorporating it within a context-aware recommendation model to improve its effectiveness. We use six months of transactional data from eBay.com to test the proposed approach and prove its validity and effectiveness. Our results show that the proposed approach provides a good estimation of the consumer's WTP, and that incorporating the consumer's WTP and seller's reputation into a recommender system significantly improves its prediction accuracy (F-score improvements of 84% compared to a matrix factorization recommendation model which doesn't take into account the seller's reputation or consumer's WTP).