TR Dizin İndeksli Yayınlar Koleksiyonu / TR Dizin Indexed Publications Collection
Permanent URI for this collectionhttps://hdl.handle.net/20.500.14365/4
Browse
2 results
Search Results
Article Stop Word Detection as a Binary Classification Problem(2017) Karaoğlan, Bahar; Metin, Senem KumovaIn a wide group of languages, the stop words, which have only grammatical roles and not contributing to information content, may be simply exposed by their relatively higher occurrence frequencies. But, in agglutinative or inflectional languages, a stop word may be observed in several different surface forms due to the inflection producing noise. In this study, some of the well-known binary classification methods are employed to overcome the inflectional noise problem in stop word detection. The experiments are conducted on corpora of an agglutinative language, Turkish, in which the amount of inflection is high and a non-agglutinative language, English, in which the inflection is lower for stop words. The evaluations demonstrated that in Turkish corpus, the classification methods improve stop word detection with respect to frequency-based method. On the other hand, the classification methods applied on English corpora showed no improvement in the performance of stop word detection.Article Certainty Factor Model in Paraphrase Detection(Pamukkale Univ, 2021) Metin, Senem Kumova; Karaoglan, Bahar; Kisla, Tarik; Soleymanzadeh, KatiraIn this paper, we address the problem of uncertainty management in identification of paraphrase sentence pairs. Paraphrase sentences are simply sets/pairs of sentences that express the same facts and/or opinions using different words or order of words. We propose the use of certainty factor (CF) model in paraphrase detection. A set of succeeding paraphrase detection features (generic and distance based features) is built by filtering and this set is used as evidences in CF model. The CF model is evaluated by F1 and accuracy measures on Microsoft Research Paraphrase corpus. The results are compared to the well-known Bayesian reasoning. The experimental results showed that CF model is an alternating paraphrase detection method to Bayes model.
