To see the other types of publications on this topic, follow the link: Clustering of text information.

Journal articles on the topic 'Clustering of text information'

Create a spot-on reference in APA, MLA, Chicago, Harvard, and other styles

Select a source type:

Consult the top 50 journal articles for your research on the topic 'Clustering of text information.'

Next to every source in the list of references, there is an 'Add to bibliography' button. Press on it, and we will generate automatically the bibliographic reference to the chosen work in the citation style you need: APA, MLA, Harvard, Chicago, Vancouver, etc.

You can also download the full text of the academic publication as pdf and read online its abstract whenever available in the metadata.

Browse journal articles on a wide variety of disciplines and organise your bibliography correctly.

1

Vedmiediev, Daniil, and Nataliia Shapoval. "Text Message Clustering." Electronics and Control Systems 4, no. 78 (2023): 16–20. http://dx.doi.org/10.18372/1990-5548.78.18255.

Full text
Abstract:
The division into groups of text messages is considered, which can be useful when building a personalized approach in different systems. Тo solve this problem, the Embedded Word2Vec was proposed. To enhance the division into groups, the suggestion of employing mini-batch k-means is presented, offering a method with lower computational demands. This recommendation aligns with the practical need for efficient and scalable clustering methods, especially when dealing with large datasets. Furthermore, the proposed metric based on the greatest common sequence is highlighted as a valuable tool for ev
APA, Harvard, Vancouver, ISO, and other styles
2

Strouse, DJ, and David J. Schwab. "The Information Bottleneck and Geometric Clustering." Neural Computation 31, no. 3 (2019): 596–612. http://dx.doi.org/10.1162/neco_a_01136.

Full text
Abstract:
The information bottleneck (IB) approach to clustering takes a joint distribution [Formula: see text] and maps the data [Formula: see text] to cluster labels [Formula: see text], which retain maximal information about [Formula: see text] (Tishby, Pereira, & Bialek, 1999 ). This objective results in an algorithm that clusters data points based on the similarity of their conditional distributions [Formula: see text]. This is in contrast to classic geometric clustering algorithms such as [Formula: see text]-means and gaussian mixture models (GMMs), which take a set of observed data points [Fo
APA, Harvard, Vancouver, ISO, and other styles
3

Zhang, Wen, Taketoshi Yoshida, Xijin Tang, and Qing Wang. "Text clustering using frequent itemsets." Knowledge-Based Systems 23, no. 5 (2010): 379–88. http://dx.doi.org/10.1016/j.knosys.2010.01.011.

Full text
APA, Harvard, Vancouver, ISO, and other styles
4

Nasir, Jamal A., Iraklis Varlamis, Asim Karim, and George Tsatsaronis. "Semantic smoothing for text clustering." Knowledge-Based Systems 54 (December 2013): 216–29. http://dx.doi.org/10.1016/j.knosys.2013.09.012.

Full text
APA, Harvard, Vancouver, ISO, and other styles
5

Fan, Wentao. "Application and analysis of text similarity in text clustering in the Chinese context." Applied and Computational Engineering 21, no. 1 (2023): 71–77. http://dx.doi.org/10.54254/2755-2721/21/20231120.

Full text
Abstract:
With the development of the Internet, information sharing is higher, and the amount of information that each user is exposed to is increasing. How to find the information peoples want from so much information is a very important question. The vast majority of these resources are related to textual information. The most intuitive manifestation of these problems is that when people usually use search engines, enter a piece of text, and search out the relevant website, if the algorithm is not good, the search results will be very unsatisfactory. Therefore, this paper studies the application of te
APA, Harvard, Vancouver, ISO, and other styles
6

Papapetrou, Odysseas, Wolf Siberski, and Norbert Fuhr. "Decentralized Probabilistic Text Clustering." IEEE Transactions on Knowledge and Data Engineering 24, no. 10 (2012): 1848–61. http://dx.doi.org/10.1109/tkde.2011.120.

Full text
APA, Harvard, Vancouver, ISO, and other styles
7

Bewoor, Mrunal S., and Suhas H. Patil. "Empirical Analysis of Single and Multi Document Summarization using Clustering Algorithms." Engineering, Technology & Applied Science Research 8, no. 1 (2018): 2562–67. https://doi.org/10.5281/zenodo.1207394.

Full text
Abstract:
<em>Abstract</em>&mdash;The availability of various digital sources has created a demand for text mining mechanisms. Effective summary generation mechanisms are needed in order to utilize relevant information from often overwhelming digital data sources. In this view, this paper conducts a survey of various single as well as multi-document text summarization techniques. It also provides analysis of treating a query sentence as a common one, segmented from documents for text summarization. Experimental results show the degree of effectiveness in text summarization over different clustering algo
APA, Harvard, Vancouver, ISO, and other styles
8

Kogan, J., C. Nicholas, and V. Volkovich. "Data Mining - Text mining with information-theoretic clustering." Computing in Science & Engineering 5, no. 6 (2003): 52–59. http://dx.doi.org/10.1109/mcise.2003.1238704.

Full text
APA, Harvard, Vancouver, ISO, and other styles
9

Liu, Yuan-chao, Chong Wu, and Ming Liu. "Research of fast SOM clustering for text information." Expert Systems with Applications 38, no. 8 (2011): 9325–33. http://dx.doi.org/10.1016/j.eswa.2011.01.126.

Full text
APA, Harvard, Vancouver, ISO, and other styles
10

Onan, Aytug, Hasan Bulut, and Serdar Korukoglu. "An improved ant algorithm with LDA-based representation for text document clustering." Journal of Information Science 43, no. 2 (2016): 275–92. http://dx.doi.org/10.1177/0165551516638784.

Full text
Abstract:
Document clustering can be applied in document organisation and browsing, document summarisation and classification. The identification of an appropriate representation for textual documents is extremely important for the performance of clustering or classification algorithms. Textual documents suffer from the high dimensionality and irrelevancy of text features. Besides, conventional clustering algorithms suffer from several shortcomings, such as slow convergence and sensitivity to the initial value. To tackle the problems of conventional clustering algorithms, metaheuristic algorithms are fr
APA, Harvard, Vancouver, ISO, and other styles
11

Kiran, V. Gaidhane* Prof. L. H. Patil Prof. C. U. Chouhan. "AN EFFICIENT APPROACH FOR TEXT MINING USING SIDE INFORMATION." INTERNATIONAL JOURNAL OF ENGINEERING SCIENCES & RESEARCH TECHNOLOGY 5, no. 7 (2016): 1137–41. https://doi.org/10.5281/zenodo.58632.

Full text
Abstract:
Nowadays, in many text mining applications, information is present in the form of text documents. Text document contains various types of information such as side information or metadata. The different types of information such as document provenance information, title of the document, links in the document, user-access behavior from web logs, or other non-textual attributes treated as side information contained into the text document. Such attributes contains a large amount of information for clustering purposes. It is difficult to estimate the importance of this side-information when text do
APA, Harvard, Vancouver, ISO, and other styles
12

Mohammed, Noorullah R., and Moulana Mohammed. "Assessment of Twitter Data Clusters with Cosine-Based Validation Metrics Using Hybrid Topic Models." Ingénierie des systèmes d information 25, no. 6 (2020): 755–69. http://dx.doi.org/10.18280/isi.250606.

Full text
Abstract:
Text data clustering is performed for organizing the set of text documents into the desired number of coherent and meaningful sub-clusters. Modeling the text documents in terms of topics derivations is a vital task in text data clustering. Each tweet is considered as a text document, and various topic models perform modeling of tweets. In existing topic models, the clustering tendency of tweets is assessed initially based on Euclidean dissimilarity features. Cosine metric is more suitable for more informative assessment, especially of text clustering. Thus, this paper develops a novel cosine b
APA, Harvard, Vancouver, ISO, and other styles
13

Li, Xiaorong, and Zhinian Shu. "Research on Big Data Text Clustering Algorithm Based on Swarm Intelligence." Wireless Communications and Mobile Computing 2022 (April 15, 2022): 1–10. http://dx.doi.org/10.1155/2022/7551035.

Full text
Abstract:
In order to break through the limitations of current clustering algorithms and avoid the direct impact of disturbance on the clustering effect of abnormal big data texts, a big data text clustering algorithm based on swarm intelligence is proposed. According to the characteristics of swarm intelligence, a differential privacy protection model is constructed; At the same time, the data location information operation is carried out to complete the location information preprocessing. Based on KD tree division, the allocation of the differential privacy budget is completed, and the location inform
APA, Harvard, Vancouver, ISO, and other styles
14

A.Ananda, Shankar, and Kumar Dr.K.R.Ananda. "Data Mining Technique for Opinion Retrieval in Healthcare System." International Journal of Data Mining & Knowledge Management Process (IJDKP) 5, no. 5 (2019): 75–84. https://doi.org/10.5281/zenodo.3463239.

Full text
Abstract:
The aim of this paper is to use Text mining(TM) concepts in the field of Health care System. We no that now days decision making in health care involves number of opinions given by the group of medical experts for specific disease in the form of decisions which will be presented in medical database in the form of text. These decisions are then mined from database with the help of Data Mining techniques. Text document clustering is considered as tool for performing information based operations. For clustering normally K-means clustering technique is used. In this paper we use Bisecting K-means
APA, Harvard, Vancouver, ISO, and other styles
15

T. NAVEEN, KUMAR, and DEVI RAMA. "A HYBRID CLUSTERING WITH SIDE INFORMATION IN TEXT MINING." i-manager's Journal on Computer Science 4, no. 2 (2016): 23. http://dx.doi.org/10.26634/jcom.4.2.8122.

Full text
APA, Harvard, Vancouver, ISO, and other styles
16

Nikhath, A. Kousar, and K. Subrahmanyam. "Feature selection, optimization and clustering strategies of text documents." International Journal of Electrical and Computer Engineering (IJECE) 9, no. 2 (2019): 1313. http://dx.doi.org/10.11591/ijece.v9i2.pp1313-1320.

Full text
Abstract:
Clustering is one of the most researched areas of data mining applications in the contemporary literature. The need for efficient clustering is observed across wide sectors including consumer segmentation, categorization, shared filtering, document management, and indexing. The research of clustering task is to be performed prior to its adaptation in the text environment. Conventional approaches typically emphasized on the quantitative information where the selected features are numbers. Efforts also have been put forward for achieving efficient clustering in the context of categorical informa
APA, Harvard, Vancouver, ISO, and other styles
17

Nikhath, A. Kousar, and K. Subrahmanyam. "Feature selection, optimization and clustering strategies of text documents." International Journal of Electrical and Computer Engineering (IJECE) 9, no. 2 (2019): 1313–20. https://doi.org/10.11591/ijece.v9i2.pp1313-1320.

Full text
Abstract:
Clustering is one of the most researched areas of data mining applications in the contemporary literature. The need for efficient clustering is observed across wide sectors including consumer segmentation, categorization, shared filtering, document management, and indexing. The research of clustering task is to be performed prior to its adaptation in the text environment. Conventional approaches typically emphasized on the quantitative information where the selected features are numbers. Efforts also have been put forward for achieving efficient clustering in the context of categorical informa
APA, Harvard, Vancouver, ISO, and other styles
18

VarshaC., Pande*1 Dr. Harshala B. Pethe2 &. Dr. Abha. S. Khandelwal3. "CLUSTERING AND CLASSIFICATION TECHNIQUES USING TEXT MINING." GLOBAL JOURNAL OF ENGINEERING SCIENCE AND RESEARCHES [NC-Rase 18] (November 12, 2018): 9–15. https://doi.org/10.5281/zenodo.1483957.

Full text
Abstract:
The text is nothing but the combination of characters. Therefore, analyzing and extracting information patterns from such data sets are more complex. Several methods have been proposed for analyzing such texts and extracting information.Data mining, a specific area named text mining is used to classify the huge semi structured or unstructured data needs proper clustering. Maximum text documents involves fast retrieval of information, arrangement of documents, exploring of information from the documents. Declaration of text input data and classification of the documents is a complex process. Te
APA, Harvard, Vancouver, ISO, and other styles
19

Agata, Teru, and Mari Agata. "Determining the possibility of deciphering an unintelligible text by text clustering." Library and Information Science 61 (June 30, 2009): 1–23. http://dx.doi.org/10.46895/lis.61.1.

Full text
APA, Harvard, Vancouver, ISO, and other styles
20

Rahman, Mahmuda. "Information retrieval with text mining for Decision Support System." Bangladesh Journal of Scientific Research 24, no. 2 (2012): 117–26. http://dx.doi.org/10.3329/bjsr.v24i2.10768.

Full text
Abstract:
Key words: Natural language processing; C4.5 classification; DSS, machine learning; KNN clustering; SVMDOI: http://dx.doi.org/10.3329/bjsr.v24i2.10768 Bangladesh J. Sci. Res. 24(2):117-126, 2011 (December)
APA, Harvard, Vancouver, ISO, and other styles
21

Wu, Yu-Chieh. "Chinese Text Categorization via Bottom-Up Weighted Word Clustering." International Journal of Enterprise Information Systems 11, no. 1 (2015): 50–61. http://dx.doi.org/10.4018/ijeis.2015010104.

Full text
Abstract:
Most of the researches on text categorization are focus on using bag of words. Some researches provided other methods for classification such as term phrase, Latent Semantic Indexing, and term clustering. Term clustering is an effective way for classification, and had been proved as a good method for decreasing the dimensions in term vectors. The authors used hierarchical term clustering and aggregating similar terms. In order to enhance the performance, they present a modify indexing with terms in cluster. Their test collection extracted from Chinese NETNEWS, and used the Centroid-Based class
APA, Harvard, Vancouver, ISO, and other styles
22

Shahana Bano, Mrs, B. Divyanjali, A. K M L R V Virajitha, and M. Tejaswi. "Document Summarization Using Clustering and Text Analysis." International Journal of Engineering & Technology 7, no. 2.32 (2018): 456. http://dx.doi.org/10.14419/ijet.v7i2.32.15740.

Full text
Abstract:
Document summarization is a procedure of shortening the content report with a product, so as to make the outline with the significant parts of unique record.Now a days ,users are very much tired about their works and they don’t have much time to spend reading a lot of information .they just want the maximum and accurate information which describes everything and occupies minimum space.This paper discusses an important approach for document summarization by using clustering and text analysis. In this paper, we are performing the clustering and text analytic techniques for reducing the data redu
APA, Harvard, Vancouver, ISO, and other styles
23

Bao, Jun Peng, Jun Yi Shen, Xiao Dong Liu, and Hai Yan Liu. "The heavy frequency vector-based text clustering." International Journal of Business Intelligence and Data Mining 1, no. 1 (2005): 42. http://dx.doi.org/10.1504/ijbidm.2005.007317.

Full text
APA, Harvard, Vancouver, ISO, and other styles
24

Wang, Xue. "Construction of Alumni Information Analysis Model Based on Big Data." Mathematical Problems in Engineering 2022 (August 9, 2022): 1–10. http://dx.doi.org/10.1155/2022/1587793.

Full text
Abstract:
In order to integrate and utilize alumni resources in a better way, big data is utilized to construct alumni information analysis model based on improved hierarchical clustering algorithm, so as to realize mining and retrieval of alumni information. First,the basic principle of hierarchical clustering algorithm is analyzed concretely. Moreover, the improvement is performed on this basis, and a method of calculating the distance between class clusters based on the ant colony optimization is proposed, which uses the shortest distance of the ant colony algorithm to optimally solve the distance be
APA, Harvard, Vancouver, ISO, and other styles
25

Hari, Prasad Bomma. "Data mining techniques and their applicability for data engineers in development and reporting." International Journal of Multidisciplinary Research and Growth Evaluation 03, no. 02 (2022): 625–27. https://doi.org/10.54660/.IJMRGE.2022.3.2-625-627.

Full text
Abstract:
AbstractWhen companies need to build decision making reports on huge amounts of data, datamining is their go-to method. Data mining is like being a detective who finds hiddenclues in a sea of information. Data engineers use clever techniques to turn raw datainto useful insights that businesses can act on. This paper explores various data miningtechniques like association rule learning, clustering, classification, regression analysis,and text mining. Additionally, the paper highlights the importance of new technologieslike machine learning, AI, and big data platforms, and how these advancements
APA, Harvard, Vancouver, ISO, and other styles
26

Zhang, Lei, Hai Qiang Chen, Wei Jie Li, Yan Zhao Liu, and Run Pu Wu. "Short Text Clustering Algorithms for Weibo Topic Detection." Advanced Materials Research 971-973 (June 2014): 1747–51. http://dx.doi.org/10.4028/www.scientific.net/amr.971-973.1747.

Full text
Abstract:
Text clustering is a popular research topic in the field of text mining, and now there are a lot of text clustering methods catering to different application requirements. Currently, Weibo data acquisition is through the API provided by big microblogging platforms. In this essay, we will discuss the algorithm of extracting popular topics posted by Weibo users by text clustering after massive data collection. Due to the fact that traditional text analysis may not be applicable to short texts used in Weibo, text clustering shall be carried out through combining multiple posts into long texts, ba
APA, Harvard, Vancouver, ISO, and other styles
27

Jin, Chun Xia, Hui Zhang, and Qiu Chan Bai. "Text Clustering Algorithm of Co-Occurrence Word Based on Association-Rule Mining." Applied Mechanics and Materials 599-601 (August 2014): 1749–52. http://dx.doi.org/10.4028/www.scientific.net/amm.599-601.1749.

Full text
Abstract:
According to the analysis of text feature, the document with co-occurrence words expresses very stronger and more accurately topic information. So this paper puts forward a text clustering algorithm of word co-occurrence based on association-rule mining. The method uses the association-rule mining to extract those word co-occurrences of expressing the topic information in the document. According to the co-occurrence words to build the modeling and co-occurrence word similarity measure, then this paper uses the hierarchical clustering algorithm based on word co-occurrence to realize text cluste
APA, Harvard, Vancouver, ISO, and other styles
28

Toledo, Assaf, Elad Venezian, and Noam Slonim. "Revisiting Sequential Information Bottleneck: New Implementation and Evaluation." Entropy 24, no. 8 (2022): 1132. http://dx.doi.org/10.3390/e24081132.

Full text
Abstract:
We introduce a modern, optimized, and publicly available implementation of the sequential Information Bottleneck clustering algorithm, which strikes a highly competitive balance between clustering quality and speed. We describe a set of optimizations that make the algorithm computation more efficient, particularly for the common case of sparse data representation. The results are substantiated by an extensive evaluation that compares the algorithm to commonly used alternatives, focusing on the practically important use case of text clustering. The evaluation covers a range of publicly availabl
APA, Harvard, Vancouver, ISO, and other styles
29

Kumar, Sushil, and Komal Kumar Bhatia. "Clustering Based Approach for Novelty Detection in Text Documents." Asian Journal of Computer Science and Technology 8, no. 2 (2019): 116–21. http://dx.doi.org/10.51983/ajcst-2019.8.2.2130.

Full text
Abstract:
As the information is overloaded over the internet accessing of information from the internet according to a given query provides redundant and irrelevant information. It is necessary to retrieve relevant and novel information from a given query by the user. With the result of this the user will require minimum effort to access the information need. In this work we proposed a clustering based approach for novelty detection which will provide the relevant and novel documents for the information need. Based on the user query the incoming stream of documents will be clustered using k-means algori
APA, Harvard, Vancouver, ISO, and other styles
30

Yang, Feng Xia. "High Quality Algorithm for Chinese Short Messages Text Clustering Based on Semantic." Advanced Materials Research 756-759 (September 2013): 3341–45. http://dx.doi.org/10.4028/www.scientific.net/amr.756-759.3341.

Full text
Abstract:
Existing data clustering method lacks considering of latent similar information existing among words,and it leads to unsatisfactory clustering result.Aiming at Chinese short message text clustering,this paper proposes a clustering algorithm based on semantic.It offers Chinese concept,and the measuring methods to calculate the similarity degree about words and Chinese short message text.It completes the clustering of Chinese short messages text through fission downwards and mergence of twos upwards.Experimental results show that this algorithm has better clustering quality than traditional algo
APA, Harvard, Vancouver, ISO, and other styles
31

Abdalgader, Khaled, Atheer A. Matroud, and Khaled Hossin. "Experimental study on short-text clustering using transformer-based semantic similarity measure." PeerJ Computer Science 10 (May 29, 2024): e2078. http://dx.doi.org/10.7717/peerj-cs.2078.

Full text
Abstract:
Sentence clustering plays a central role in various text-processing activities and has received extensive attention for measuring semantic similarity between compared sentences. However, relatively little focus has been placed on evaluating clustering performance using available similarity measures that adopt low-dimensional continuous representations. Such representations are crucial in domains like sentence clustering, where traditional word co-occurrence representations often achieve poor results when clustering semantically similar sentences that share no common words. This article present
APA, Harvard, Vancouver, ISO, and other styles
32

SUNAYAMA, Wataru, Shuhei HAMAOKA, and Kiyoshi OKUDA. "Recursive Clustering of a Text Data Set for Information Collection." Journal of Japan Society for Fuzzy Theory and Intelligent Informatics 24, no. 3 (2012): 697–706. http://dx.doi.org/10.3156/jsoft.24.697.

Full text
APA, Harvard, Vancouver, ISO, and other styles
33

LI, Xiao-Guang, Ge YU, Da-Ling WANG, and Yu-Bin BAO. "Latent Concept Extraction and Text Clustering Based on Information Theory*." Journal of Software 19, no. 9 (2008): 2276–84. http://dx.doi.org/10.3724/sp.j.1001.2008.02276.

Full text
APA, Harvard, Vancouver, ISO, and other styles
34

Tuan Zakaria, Tuan Norhafizah, Mohd Juzaiddin Ab Aziz, Mohd Rosmadi Mokhtar, and Saadiyah Darus. "TEXT CLUSTERING FOR REDUCING SEMANTIC INFORMATION IN MALAY SEMANTIC REPRESENTATION." Asia-Pacific Journal of Information Technology and Multimedia 09, no. 02 (2020): 11–24. http://dx.doi.org/10.17576/apjitm-2020-0902-02.

Full text
APA, Harvard, Vancouver, ISO, and other styles
35

Zhang, Yuan, Yanping Zhang, and Runmei Zhang. "Text information classification method based on secondly fuzzy clustering algorithm." Journal of Intelligent & Fuzzy Systems 38, no. 6 (2020): 7743–54. http://dx.doi.org/10.3233/jifs-179844.

Full text
APA, Harvard, Vancouver, ISO, and other styles
36

Jie Cao, Zhiang Wu, Junjie Wu, and Hui Xiong. "SAIL: Summation-bAsed Incremental Learning for Information-Theoretic Text Clustering." IEEE Transactions on Cybernetics 43, no. 2 (2013): 570–84. http://dx.doi.org/10.1109/tsmcb.2012.2212430.

Full text
APA, Harvard, Vancouver, ISO, and other styles
37

Pullwitt, Daniel. "Integrating contextual information to enhance SOM-based text document clustering." Neural Networks 15, no. 8-9 (2002): 1099–106. http://dx.doi.org/10.1016/s0893-6080(02)00082-5.

Full text
APA, Harvard, Vancouver, ISO, and other styles
38

SONG, SA-KWANG, DONG HYUN JANG, and SUNG HYON MYAENG. "Text Summarization Based on Sentence Clustering with Rhetorical Structure Information." International Journal of Computer Processing of Languages 18, no. 02 (2005): 153–70. http://dx.doi.org/10.1142/s0219427905001250.

Full text
APA, Harvard, Vancouver, ISO, and other styles
39

Granados, Ana, David Camacho, and Francisco Borja Rodríguez. "Is the contextual information relevant in text clustering by compression?" Expert Systems with Applications 39, no. 10 (2012): 8537–46. http://dx.doi.org/10.1016/j.eswa.2012.01.215.

Full text
APA, Harvard, Vancouver, ISO, and other styles
40

Meng, LingYan. "Text Clustering and Economic Analysis of Free Trade Zone Governance Strategies Based on Random Matrix and Subject Analysis." Mathematical Problems in Engineering 2022 (August 30, 2022): 1–9. http://dx.doi.org/10.1155/2022/2731807.

Full text
Abstract:
The steps of generating basic data by the LDA model and calculating text by the weighted algorithm have a good effect on text clustering. In this paper, the LDA topic model is used to effectively improve the accuracy of strategy text clustering. FTZ economics text clustering simulates FTA economics text data and economic data, imports economics and economic figures and word lists, and uses the traditional vector space model for factor representation. After that, the text vectors are independent of each other, ignoring the semantic relationship, which affects the clustering analysis results. A
APA, Harvard, Vancouver, ISO, and other styles
41

Ji, KeKe, ZhengZhong Li, Jian Chen, GuanYan Wang, KeLiang Liu, and Yi Luo. "Freeway accident duration prediction based on social network information." Neural Network World 32, no. 2 (2022): 93–112. http://dx.doi.org/10.14311/nnw.2022.32.006.

Full text
Abstract:
Accident duration prediction is the basis of freeway emergency management, and timely and accurate accident duration prediction can provide a reliable basis for road traffic diversion and rescue agencies. This study proposes a method for predicting the duration of freeway accidents based on social network information by collecting Weibo data of freeway accidents in Sichuan province and using the advantage that human language can convey multi-dimensional information. Firstly, text features are extracted through a TF-IDF model to represent the accident text data quantitatively; secondly, the var
APA, Harvard, Vancouver, ISO, and other styles
42

Calvagna, Andrea, Emiliano Tramontana, and Gabriella Verga. "Automated Social Media Text Clustering Based on Financial Ontologies." Information 15, no. 4 (2024): 210. http://dx.doi.org/10.3390/info15040210.

Full text
Abstract:
Social media networks provide an aggregation of news and content, allowing users to share and discuss topics of greatest interest to them. Users can enrich the news by providing context and opinions that are useful to other users. Understanding topics of interest sheds light on the collective thinking of a group of individuals and offers important insights for exploring a given field. Among the fields of interest on social media networks, finance stands out. Automatically identifying and organizing the main issues that users discuss can be useful for multiple purposes, e.g., identifying the pr
APA, Harvard, Vancouver, ISO, and other styles
43

Roussinov, Dmitri, and J. Leon Zhao. "Text clustering and summary techniques for CRM message management." Journal of Enterprise Information Management 17, no. 6 (2004): 424–29. http://dx.doi.org/10.1108/17410390410566715.

Full text
APA, Harvard, Vancouver, ISO, and other styles
44

Wang, Yao, Fuguo Liu, and Guodong Li. "Clustering Analysis of Hotel Network Reviews Based on Text Mining Method." Industry Science and Engineering 1, no. 4 (2024): 51–59. http://dx.doi.org/10.62381/i245406.

Full text
Abstract:
With the development of information technology, users use online platforms to post real-time online comments to express their preferences and opinions on goods or services. Online review information expresses users' behavioral habits and special preferences. In depth, analysis of hotel online reviews can improve the adaptability of hotel services to user needs. Effective mining of the vast user review data will provide value for the development of the tourism industry. Using text mining methods to process hotel review data, multiple clustering methods were compared and analyzed for positive an
APA, Harvard, Vancouver, ISO, and other styles
45

Zheng, Na, and Jie Yu Wu. "Cluster Analysis for Internet Public Sentiment in Universities by Combining Methods." International Journal of Recent Contributions from Engineering, Science & IT (iJES) 6, no. 3 (2018): 60. http://dx.doi.org/10.3991/ijes.v6i3.9670.

Full text
Abstract:
A clustering method based on the Latent Dirichlet Allocation and the VSM model to compute the text similarity is presented. The Latent Dirichlet Allocation subject models and the VSM vector space model weights strategy are used respectively to calculate the text similarity. The linear combination of the two results is used to get the text similarity. Then the k-means clustering algorithm is chosen for cluster analysis. It can not only solve the deep semantic information leakage problems of traditional text clustering, but also solve the problem of the LDA that could not distinguish the texts b
APA, Harvard, Vancouver, ISO, and other styles
46

Radomirović, Branislav, Vuk Jovanović, Bosko Nikolić, et al. "Text Document Clustering Approach by Improved Sine Cosine Algorithm." Information Technology and Control 52, no. 2 (2023): 541–61. http://dx.doi.org/10.5755/j01.itc.52.2.33536.

Full text
Abstract:
Due to the vast amounts of textual data available in various forms such as online content, social media comments, corporate data, public e-services and media data, text clustering has been experiencing rapid development. Text clustering involves categorizing and grouping similar content. It is a process of identifying significant patterns from unstructured textual data. Algorithms are being developed globally to extract useful and relevant information from large amounts of text data. Measuring the significance of content in documents to partition the collection of text data is one of the most
APA, Harvard, Vancouver, ISO, and other styles
47

Zhang, Jian, Weichao Gao, and Yanhe Jia. "WES-BTM: A Short Text-Based Topic Clustering Model." Symmetry 15, no. 10 (2023): 1889. http://dx.doi.org/10.3390/sym15101889.

Full text
Abstract:
User comments often contain their most practical requirements. Using topic modeling of user comments, it is possible to classify and downscale text data, mine the information in user comments, and understand users’ requirements and preferences. However, user comment texts are usually short and lack rich word frequency and contextual information with sparsity. The traditional topic model cannot model and analyze these short texts well. The biterm topic model (BTM), while solving the sparsity problem, suffers from accuracy and noise problems. In order to eliminate information barriers and furthe
APA, Harvard, Vancouver, ISO, and other styles
48

Sarannya, S., M. Venkatesan, and Prabhavathy Panner. "Double Clustering Based Neural Feedback Method for Unstructured Text Data." Journal of Computational and Theoretical Nanoscience 18, no. 4 (2021): 1306–11. http://dx.doi.org/10.1166/jctn.2021.9385.

Full text
Abstract:
Text clustering has now a days become a very major technique in many fields including data mining, Natural Language Processing etc. It’s also broadly used for information retrieval and assimilation of textual data. Majority of the works which were carried out previously focuses on the clustering algorithms where feature extraction is done without considering the semantic meaning of word based on its context. In the given work, we introduce a double clustering algorithm using K -Means, by using in conjuction, a Bi-directional Long Short-Term Memory and a Convolutional Neural Network for the pur
APA, Harvard, Vancouver, ISO, and other styles
49

Pourvali, Mohsen, and Salvatore Orlando. "Enriching Documents by Linking Salient Entities and Lexical-Semantic Expansion." Journal of Intelligent Systems 29, no. 1 (2018): 1109–21. http://dx.doi.org/10.1515/jisys-2018-0098.

Full text
Abstract:
Abstract This paper explores a multi-strategy technique that aims at enriching text documents for improving clustering quality. We use a combination of entity linking and document summarization in order to determine the identity of the most salient entities mentioned in texts. To effectively enrich documents without introducing noise, we limit ourselves to the text fragments mentioning the salient entities, in turn, belonging to a knowledge base like Wikipedia, while the actual enrichment of text fragments is carried out using WordNet. To feed clustering algorithms, we investigate different do
APA, Harvard, Vancouver, ISO, and other styles
50

Sohrabi, Babak, Iman Raeesi Vanani, and Ehsan Abedin. "Human Resources Management and Information Systems Trend Analysis Using Text Clustering." International Journal of Human Capital and Information Technology Professionals 9, no. 3 (2018): 1–24. http://dx.doi.org/10.4018/ijhcitp.2018070101.

Full text
Abstract:
Human resources management has seen a significant change by the emergence of information systems from a traditional or popularly called personnel management to the modern one. The purpose of this article is to study the trends of information systems in the field of human resources management in combination with information systems through text mining approaches on a broad exploration of internationally published papers. Among text analytics methods for extracting trends, text clustering has been applied to the dataset of highly-ranked information systems journals. The data set was obtained fro
APA, Harvard, Vancouver, ISO, and other styles
We offer discounts on all premium plans for authors whose works are included in thematic literature selections. Contact us to get a unique promo code!