In this paper, we are the first to construct a software programming taxonomy from Stackoverflow. More precisely, we propose a machine learning based method with novel features to capture the hierarchical semantic structure of tags in Stackoverflow. A graph pruning algorithm is applied to eliminate the conflicts by constructing a Directed Acyclic Graph (DAG). As a result, our dataset, named Software.zhishi.schema, contains 38,205 concepts together with 36,249 subsumption relations. In order to further test the usability of our published data, we adopt a similarity computing task of words from software programming which is one of the most fundamental tasks in the software repository mining area. The results show that our dataset can outperform other knowledge bases due to its high coverage with finergrained domain concepts.
Jens H. Kuhn, Artem Babaian, Laura M. Bergner, Paul Dény, Dieter Glebe, Masayuki Horie, Eugene V Koonin, Mart Krupovìč, Sofia Paraskevopoulou, Marcos de la Peña, Teemu Smura, Jussi Hepojoki
Discussion(0)
No comments yet. Be the first to comment.