526 publications from this institution
Proactive recommender systems push recommendations to users without their explicit request whenever a recommendation that suits a user is available. These systems strive to optimize the match between recommended items and users' preferences. We assume that recommendations might be reflected with low accuracy not only due to the recommended items' suitability to the user, but also because of the recommendations' timings. We therefore claim that it is possible to learn a model of good and bad contexts for recommendations that can later be integrated in a recommender system. Using mobile data collected during a three week user study, we suggest a two-phase model that is able to classify whether a certain context is at all suitable for any recommendation, regardless of its content. Results reveal that a hybrid model that first decides whether it should use a personal or a non-personal timing model, and then classifies accordingly whether the timing is proper for recommendations, is superior to both the personal or non-personal timing models.
No abstract is provided for this article.
This updated compendium provides a methodical introduction with a coherent and unified repository of ensemble methods, theories, trends, challenges, and applications. More than a third of this edition comprised of new materials, highlighting descriptions of the classic methods, and extensions and novel approaches that have recently been introduced.Along with algorithmic descriptions of each method, the settings in which each method is applicable and the consequences and tradeoffs incurred by using the method is succinctly featured. R code for implementation of the algorithm is also emphasized.The unique volume provides researchers, students and practitioners in industry with a comprehensive, concise and convenient resource on ensemble learning methods.
Sequential data is everywhere, and it can serve as a basis for research that will lead to improved processes. For example, road infrastructure can be improved by identifying bottlenecks in GPS data, or early diagnosis can be improved by analyzing patterns of disease progression in medical data. The main obstacle is that access and use of such data is usually limited or not permitted at all due to concerns about violating user privacy, and rightly so. Anonymizing sequence data is not a simple task, since a user creates an almost unique signature over time. Existing anonymization methods reduce the quality of information in order to maintain the level of anonymity required. Damage to quality may disrupt patterns that appear in the original data and impair the preservation of various characteristics. Since in many cases the researcher does not need the data as is and instead is only interested in the patterns that exist in the data, we propose PrivGen, an innovative method for generating data that maintains patterns and characteristics of the source data. We demonstrate that the data generation mechanism significantly limits the risk of privacy infringement. Evaluating our method with real-world datasets shows that its generated data preserves many characteristics of the data, including the sequential model, as trained based on the source data. This suggests that the data generated by our method could be used in place of actual data for various types of analysis, maintaining user privacy and the data's integrity at the same time.
XML transactions are used in many information systems to store data and interact with other systems. Abnormal transactions, the result of either an on-going cyber attack or the actions of a benign user, can potentially harm the interacting systems and therefore they are regarded as a threat. In this paper we address the problem of anomaly detection and localization in XML transactions using machine learning techniques. We present a new XML anomaly detection framework, XML-AD. Within this framework, an automatic method for extracting features from XML transactions was developed as well as a practical method for transforming XML features into vectors of fixed dimensionality. With these two methods in place, the XML-AD framework makes it possible to utilize general learning algorithms for anomaly detection. Central to the functioning of the framework is a novel multi-univariate anomaly detection algorithm, ADIFA. The framework was evaluated on four XML transactions datasets, captured from real information systems, in which it achieved over 89% true positive detection rate with less than a 0.2% false positive rate.
In the past decade, the usage of mobile devices has gone far beyond simple activities like calling and texting. Today, smartphones contain multiple embedded sensors and are able to collect useful sensing data about the user and infer the user's context. The more frequent the sensing, the more accurate the context. However, continuous sensing results in huge energy consumption, decreasing the battery's lifetime. We propose a novel approach for cost-aware sensing when performing continuous latent context detection. The suggested method dynamically determines user's sensors sampling policy based on three factors: (1) User's last known context; (2) Predicted information loss using KL-Divergence; and (3) Sensors' sampling costs. The objective function aims at minimizing both sampling cost and information loss. The method is based on various machine learning techniques including autoencoder neural networks for latent context detection, linear regression for information loss prediction, and convex optimization for determining the optimal sampling policy. To evaluate the suggested method, we performed a series of tests on real-world data recorded at a high-frequency rate; the data was collected from six mobile phone sensors of twenty users over the course of a week. Results show that by applying a dynamic sampling policy, our method naturally balances information loss and energy consumption and outperforms the static approach.% We compared the performance of our method with another state of the art dynamic sampling method and demonstrate its consistent superiority in various measures. %Our methods outperformed, and were able to improve we achieved better results in either sampling cost or information loss, and in some cases we improved both.
Feature selection is an essential process for machine learning tasks since it improves generalization capabilities, and reduces run-time and a model’s complexity. In many applications, the cost of collecting the features must be taken into account. To cope with the cost problem, we developed a new cost-sensitive fitness function based on histogram comparison. This function is integrated with a genetic search method to form a new feature selection algorithm termed CASH (cost-sensitive attribute selection algorithm using histograms). The CASH algorithm takes into account feature collection costs as well as feature grouping and misclassification costs. Our experiments in various domains demonstrated the superiority of CASH over several other cost-sensitive genetic algorithms.
Background: Patients with chronic lymphocytic leukemia (CLL) are known to have a suboptimal immune response of both humoral and cellular arms. Recently, a BNT162b2 mRNA COVID-19 vaccine was introduced with a high efficacy of 95% in immunocompetent individuals. Approximately half of the patients with CLL fail to mount a humoral response to the vaccine, as detected by anti-spike antibodies. Currently, there is no data available regarding T-cell immune responses following the vaccine of these patients. Aim of the study: To investigate T-cell response determined by interferon gamma (IFNγ) secretion in patients with CLL following BNT162b mRNA Covid-19 vaccine, in comparison with serologic response. Methods: CLL patients from 3 medical centers in Israel were included in the study. All patients received two 30-μg doses of BNT162b2 vaccine (Pfizer), administered intramuscularly 3 weeks apart. For evaluation of SARS-CoV-2 Spike-specific T-cell responses, blood samples were stimulated ex-vivo with Spike protein and secreted IFNγ was quantified (ELISA DuoSet, R&D Systems, Minneapolis, Minnesota, USA). T-cell immune response was considered to be positive for values above 25 pg/ml of Spike-specific response. T-cell subpopulations were characterized by flow cytometry (CD3, CD4, CD8). Anti-spike antibody tests were performed using the Architect AdviseDx SARS-CoV-2 IgG II (Abbot, Lake Forest, Illinois, USA). Statistical analysis was performed using Mann-Whitney test for continuous variables while the Wald Chi-square test was used for comparing categorical variables. Results: 83 patients with CLL were tested for T-cell response. Blood samples were collected after a median time of 139 days post administration of the second dose of vaccine (IQ range 134-152). Out of 83 patients, 68 were eligible for the analysis (with positive internal control). Median age of the cohort was 68 years (56-72); and 44 (65%) were males. 19 (28%) patients were treatment-naïve, most of whom were Binet stage A or B. 31 (46%) patients were on therapy: 17 with a BTK-inhibitor, and 13 with a venetoclax-based regimen. 29 (42%) patients were previously treated with anti-CD20, 13 of whom in the 12 months period prior to vaccination. T cell immune response to the vaccine was evident in 22 (32%) patients. CIRS Score>6 and specifically hypertension were statistically significantly associated with a lower T-cell response (univariate analysis, p-value<0.05). Variables that were associated with the development of T-cell response were presence of del(13q), IgM ≥ 40 mg/dL, and IgA ≥ 80 mg/dL (p-value 0.05-0.1). There was no significant difference with regards to age, gender, other CLL-specific prognostic markers, treatment, and T-cell subpopulation distribution according to flow cytometry (Table 1). The presence of T-cell response highly correlated with both the detection of anti-spike IgG antibodies following the second dose (p=0.0239) and at the time of T-cell testing (n=66, p=0.048, Table 2). While 50% of patients who tested positive for anti-spike IgG antibodies also developed positive T-cell response, only 17% of patients who did not develop T-cell response, tested positive for anti-spike antibodies. Importantly, 24% of the patients who tested negative for anti-spike IgG antibodies, developed positive T cell response. Moreover, the level of the T-cell response (log transformed) correlated linearly with (log transformed) anti-spike IgG titer (adjusted r=0.26 and p =0.026 according to Pearson correlation, Figure 1). Conclusion: We show that cellular immune response to the BNT162b2 mRNA COVID-19 vaccine, is blunted in most CLL patients and that there is a correlation between T-cell response and serologic response to the vaccine. These results need to be validated in a larger cohort. Figure 1 Disclosures Itchaki: AbbVie: Consultancy, Honoraria, Research Funding; Janssen: Consultancy, Honoraria, Research Funding. Benjamini: Janssen: Consultancy, Honoraria, Research Funding; AbbVie: Consultancy, Honoraria, Research Funding. Tadmor: AbbVie: Consultancy, Honoraria, Research Funding; Janssen: Consultancy, Honoraria, Research Funding.
No abstract is provided for this article.
No abstract is provided for this article.
No abstract is provided for this article.
Context-aware systems enable the sensing and analysis of user context in order to provide personalized services. Our study is part of growing research efforts examining how high-dimensional data collected from mobile devices can be utilized to infer users' dynamic preferences. We present a novel method for inferring contextual user preferences by using an unsupervised deep learning technique applied to mobile sensor data. We train an auto-encoder for each user preference with contextual data that based on past user interaction with the system. Given new contextual sensor data from a user, the patterns discovered from each auto-encoder are used to predict the most likely preference in the given context. This can greatly enhance a variety of services, such as mobile online advertising and context-aware recommender systems. We demonstrate our contribution with a point of interest (POI) recommender system in which we label contextual preferences based on the interaction of users with categories of items. Empirical results utilizing a real world dataset of mobile users show a significant improvement (16% to 73% improvement) in classification accuracy compared with state of the art classification methods.
No abstract is provided for this article.
In this paper, we present a black-box attack against API call based machine learning malware classifiers, focusing on generating adversarial sequences combining API calls and static features (e.g., printable strings) that will be misclassified by the classifier without affecting the malware functionality. We show that this attack is effective against many classifiers due to the transferability principle between RNN variants, feed forward DNNs, and traditional machine learning classifiers such as SVM. We also implement GADGET, a software framework to convert any malware binary to a binary undetected by malware classifiers, using the proposed attack, without access to the malware source code.
A model-based co-clustering divides the data based on two main axes and simultaneously trains a supervised model for each co-cluster using all other input features. For example, in the rating prediction task of recommender system, the main two axes are items and users. In each co-cluster, we train a regression model for predicting the rating based on other features such as user's characteristics (e.g., gender), item's characteristics (e.g., genre), contextual features (e.g., location), and so on. In reality, users and items do not necessarily belong to a single co-cluster, but rather can be associated with several co-clusters. We extend the model-based co-clustering to support fuzzy co-clustering. In this setting, each item–user pair is associated to every co-cluster with some membership grade. This grade indicates the level of relevance of the item–user pair to the co-cluster. Furthermore, we propose a distributed algorithm, based on a map-reduce approach, to handle big datasets. Evaluating the fuzzy co-clustering algorithm on three datasets shows a significant improvement comparing with a regular co-clustering algorithm. In addition, a map-reduce version of the fuzzy co-clustering algorithm significantly reduces the runtime.