To see the other types of publications on this topic, follow the link: Multimodal Information Retrieval.

Journal articles on the topic 'Multimodal Information Retrieval'

Create a spot-on reference in APA, MLA, Chicago, Harvard, and other styles

Select a source type:

Consult the top 50 journal articles for your research on the topic 'Multimodal Information Retrieval.'

Next to every source in the list of references, there is an 'Add to bibliography' button. Press on it, and we will generate automatically the bibliographic reference to the chosen work in the citation style you need: APA, MLA, Harvard, Chicago, Vancouver, etc.

You can also download the full text of the academic publication as pdf and read online its abstract whenever available in the metadata.

Browse journal articles on a wide variety of disciplines and organise your bibliography correctly.

1

Xu, Hong. "Multimodal bird information retrieval system." Applied and Computational Engineering 53, no. 1 (2024): 96–102. http://dx.doi.org/10.54254/2755-2721/53/20241282.

Full text
Abstract:
Multimodal bird information retrieval system can help people popularize bird knowledge and help bird conservation. In this paper, we use the self-built bird dataset, the ViT-B/32 model in CLIP model as the training model, python as the development language, and PyQT5 to complete the interface development. The system mainly realizes the uploading and displaying of bird pictures, the multimodal retrieval function of bird information, and the introduction of related bird information. The results of the trial run show that the system can accomplish the multimodal retrieval of bird information, ret
APA, Harvard, Vancouver, ISO, and other styles
2

Cui, Chenhao, and Zhoujun Li. "Prompt-Enhanced Generation for Multimodal Open Question Answering." Electronics 13, no. 8 (2024): 1434. http://dx.doi.org/10.3390/electronics13081434.

Full text
Abstract:
Multimodal open question answering involves retrieving relevant information from both images and their corresponding texts given a question and then generating the answer. The quality of the generated answer heavily depends on the quality of the retrieved image–text pairs. Existing methods encode and retrieve images and texts, inputting the retrieved results into a language model to generate answers. These methods overlook the semantic alignment of image–text pairs within the information source, which affects the encoding and retrieval performance. Furthermore, these methods are highly depende
APA, Harvard, Vancouver, ISO, and other styles
3

Li, Peize, Qingyi Si, Peng Fu, Zheng Lin, and Yan Wang. "Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering." Proceedings of the AAAI Conference on Artificial Intelligence 39, no. 5 (2025): 4851–59. https://doi.org/10.1609/aaai.v39i5.32513.

Full text
Abstract:
Retrieval-based multi-image question answering (QA) task involves retrieving multiple question-related images and synthesizing these images to generate an answer. Conventional "retrieve-then-answer" pipelines often suffer from cascading errors because the training objective of QA fails to optimize the retrieval stage. To address this issue, we propose a novel method to effectively introduce and reference retrieved information into the QA. Given the image set to be retrieved, we employ a multimodal large language model (visual perspective) and a large language model (textual perspective) to obt
APA, Harvard, Vancouver, ISO, and other styles
4

Yao, Ziru, and Fuzheng Zhao. "The Optimization of Media Information Retrieval and Adaptive Information Management Based on Deep Reinforcement Learning." Journal of Organizational and End User Computing 37, no. 1 (2025): 1–45. https://doi.org/10.4018/joeuc.378389.

Full text
Abstract:
The rapid growth of large-scale information and the dynamic nature of user behaviors pose significant challenges for modern information retrieval systems, which often struggle to adapt to non-stationary environments and fail to fully utilize multimodal data, leading to suboptimal performance. To address these issues, this study proposes the adaptive deep reinforcement learning (RL) framework for information retrieval and management, which combines RL, multimodal data fusion, and an adaptive update mechanism to dynamically adjust to evolving user preferences and document collections. The adapti
APA, Harvard, Vancouver, ISO, and other styles
5

Mounika, Pilli, K. Venkata Subba Reddy, and N. Ramakrishnaiah. "Comprehensive review and analysis on multi modal image retrieval." Journal of Information and Optimization Sciences 46, no. 2 (2025): 403–13. https://doi.org/10.47974/jios-1923.

Full text
Abstract:
Multimodal image retrieval, which involves retrieving images using various modalities such as text, audio, or other images, has significant research importance due to its wide-ranging applications and potential to enhance user experiences across multiple domains. Traditional image retrieval systems, content-based image retrieval (CBIR) systems rely solely on visual features, which can be limiting. By integrating multiple modalities, multimodal retrieval systems can offer more robust and accurate results. A research problem is considered multimodal when it integrates information from more than
APA, Harvard, Vancouver, ISO, and other styles
6

Kulvinder Singh, Et al. "Enhancing Multimodal Information Retrieval Through Integrating Data Mining and Deep Learning Techniques." International Journal on Recent and Innovation Trends in Computing and Communication 11, no. 9 (2023): 560–69. http://dx.doi.org/10.17762/ijritcc.v11i9.8844.

Full text
Abstract:
Multimodal information retrieval, the task of re trieving relevant information from heterogeneous data sources such as text, images, and videos, has gained significant attention in recent years due to the proliferation of multimedia content on the internet. This paper proposes an approach to enhance multimodal information retrieval by integrating data mining and deep learning techniques. Traditional information retrieval systems often struggle to effectively handle multimodal data due to the inherent complexity and diversity of such data sources. In this study, we leverage data mining techniqu
APA, Harvard, Vancouver, ISO, and other styles
7

Chen, Yali, Bin Hu, and Yajuan Liu. "Optimizing document management and retrieval with multimodal transformers and knowledge graphs." PLOS One 20, no. 6 (2025): e0323966. https://doi.org/10.1371/journal.pone.0323966.

Full text
Abstract:
In the digital age, multimodal archival data is experiencing explosive growth, and how to efficiently and accurately retrieve information from it has become a key challenge. Traditional retrieval methods struggle to effectively handle multi-source heterogeneous multimodal data, leading to poor retrieval accuracy and efficiency. To address this issue, this paper proposes the MDKG-RL model, which organically integrates knowledge graph reasoning, deep reinforcement learning dynamic optimization, and multimodal Transformer architecture to achieve deep semantic understanding of multimodal data and
APA, Harvard, Vancouver, ISO, and other styles
8

Lang, Jian, Zhangtao Cheng, Ting Zhong, and Fan Zhou. "Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal Learning." Proceedings of the AAAI Conference on Artificial Intelligence 39, no. 17 (2025): 18035–43. https://doi.org/10.1609/aaai.v39i17.33984.

Full text
Abstract:
Multimodal learning with incomplete modality is practical and challenging. Recently, researchers have focused on enhancing the robustness of pre-trained MultiModal Transformers (MMTs) under missing modality conditions by applying learnable prompts. However, these prompt-based methods face several limitations: (1) incomplete modalities provide restricted modal cues for task-specific inference, (2) dummy imputation for missing content causes information loss and introduces noise, and (3) static prompts are instance-agnostic, offering limited knowledge for instances with various missing condition
APA, Harvard, Vancouver, ISO, and other styles
9

Pan, Xiaoning. "Application of Multimedia Information Retrieval Technology in Japanese Text Content Information Query Platform." International Journal of e-Collaboration 20, no. 1 (2024): 1–20. http://dx.doi.org/10.4018/ijec.356491.

Full text
Abstract:
The research and development of Japanese operating platform has become a new industry. At the same time, it is necessary to strengthen the research on query and feedback technology in MIR(Multimedia Information Retrieval). Based on this, this paper introduces the research status and progress of query and feedback technology in multimedia retrieval, and introduces the overall framework and technical links of MIR. This paper discusses the application of MIR technology in Japanese text content information query platform. A new multimodal retrieval mechanism is proposed, which combines cross-refer
APA, Harvard, Vancouver, ISO, and other styles
10

UbaidullahBokhari, Mohammad, and Faraz Hasan. "Multimodal Information Retrieval: Challenges and Future Trends." International Journal of Computer Applications 74, no. 14 (2013): 9–12. http://dx.doi.org/10.5120/12951-9967.

Full text
APA, Harvard, Vancouver, ISO, and other styles
11

Calumby, Rodrigo Tripodi. "Diversity-oriented Multimodal and Interactive Information Retrieval." ACM SIGIR Forum 50, no. 1 (2016): 86. http://dx.doi.org/10.1145/2964797.2964811.

Full text
APA, Harvard, Vancouver, ISO, and other styles
12

S. Gomathy, K. P. Deepa, T. Revathi, and L. Maria Michael Visuwasam. "Genre Specific Classification for Information Search and Multimodal Semantic Indexing for Data Retrieval." SIJ Transactions on Computer Science Engineering & its Applications (CSEA) 01, no. 01 (2013): 10–15. http://dx.doi.org/10.9756/sijcsea/v1i1/01010159.

Full text
APA, Harvard, Vancouver, ISO, and other styles
13

ZHANG, Jing. "Video retrieval model based on multimodal information fusion." Journal of Computer Applications 28, no. 1 (2008): 199–201. http://dx.doi.org/10.3724/sp.j.1087.2008.00199.

Full text
APA, Harvard, Vancouver, ISO, and other styles
14

Mourão, André, Flávio Martins, and João Magalhães. "Multimodal medical information retrieval with unsupervised rank fusion." Computerized Medical Imaging and Graphics 39 (January 2015): 35–45. http://dx.doi.org/10.1016/j.compmedimag.2014.05.006.

Full text
APA, Harvard, Vancouver, ISO, and other styles
15

Revuelta-Martínez, Alejandro, Luis Rodríguez, Ismael García-Varea, and Francisco Montero. "Multimodal interaction for information retrieval using natural language." Computer Standards & Interfaces 35, no. 5 (2013): 428–41. http://dx.doi.org/10.1016/j.csi.2012.11.002.

Full text
APA, Harvard, Vancouver, ISO, and other styles
16

Chen, Xu, Alfred O. Hero, III, and Silvio Savarese. "Multimodal Video Indexing and Retrieval Using Directed Information." IEEE Transactions on Multimedia 14, no. 1 (2012): 3–16. http://dx.doi.org/10.1109/tmm.2011.2167223.

Full text
APA, Harvard, Vancouver, ISO, and other styles
17

Hubert, Gilles, and Josiane Mothe. "An adaptable search engine for multimodal information retrieval." Journal of the American Society for Information Science and Technology 60, no. 8 (2009): 1625–34. http://dx.doi.org/10.1002/asi.21091.

Full text
APA, Harvard, Vancouver, ISO, and other styles
18

Lin, Xiaoqin, Chentao Han, Jian Yao, Yue Li, Xujun Wang, and Shufeng Jia. "MKNNet: Knowledge-aligned multimodal transformer for information retrieval." Alexandria Engineering Journal 127 (August 2025): 1029–39. https://doi.org/10.1016/j.aej.2025.06.055.

Full text
APA, Harvard, Vancouver, ISO, and other styles
19

Lau, David, Ganthan Narayana Samy, Dr Fiza Abdul Rahim, et al. "Multimodal RAG Analysis of Product Datasheet." Open International Journal of Informatics 12, no. 2 (2024): 1–12. https://doi.org/10.11113/oiji2024.12n2.309.

Full text
Abstract:
Large language models such as ChatGPT serves as multipurpose chatbot that can provide information across diverse disciplines. However, in order to generate timely and accurate response, retrieval-augmented generation method has been devised to enhance the response of these models. The release of vision models has paved the way for practitioners to perform multimodal retrieval augmented generation on documents that commonly consist of a combination of text, images and tables. Hence, this method is explored to analyze a product datasheet and match it with minimum specification required by potent
APA, Harvard, Vancouver, ISO, and other styles
20

Cao, Yu, Shawn Steffey, Jianbiao He, et al. "Medical Image Retrieval: A Multimodal Approach." Cancer Informatics 13s3 (January 2014): CIN.S14053. http://dx.doi.org/10.4137/cin.s14053.

Full text
Abstract:
Medical imaging is becoming a vital component of war on cancer. Tremendous amounts of medical image data are captured and recorded in a digital format during cancer care and cancer research. Facing such an unprecedented volume of image data with heterogeneous image modalities, it is necessary to develop effective and efficient content-based medical image retrieval systems for cancer clinical practice and research. While substantial progress has been made in different areas of content-based image retrieval (CBIR) research, direct applications of existing CBIR techniques to the medical images pr
APA, Harvard, Vancouver, ISO, and other styles
21

Datta, Deepanwita, Shubham Varma, Ravindranath Chowdary C., and Sanjay K. Singh. "Multimodal Retrieval using Mutual Information based Textual Query Reformulation." Expert Systems with Applications 68 (February 2017): 81–92. http://dx.doi.org/10.1016/j.eswa.2016.09.039.

Full text
APA, Harvard, Vancouver, ISO, and other styles
22

Imhof, Melanie, and Martin Braschler. "A study of untrained models for multimodal information retrieval." Information Retrieval Journal 21, no. 1 (2017): 81–106. http://dx.doi.org/10.1007/s10791-017-9322-x.

Full text
APA, Harvard, Vancouver, ISO, and other styles
23

Soni, Ankita, and Richa Chouhan. "Multimodal Information Retrieval by using Visual and Textual Query." International Journal of Computer Applications 137, no. 1 (2016): 6–10. http://dx.doi.org/10.5120/ijca2016908637.

Full text
APA, Harvard, Vancouver, ISO, and other styles
24

Sattari, Saeid, and Adnan Yazici. "Multimodal query-level fusion for efficient multimedia information retrieval." International Journal of Intelligent Systems 33, no. 10 (2018): 2019–37. http://dx.doi.org/10.1002/int.21920.

Full text
APA, Harvard, Vancouver, ISO, and other styles
25

Hu, Peng, Dezhong Peng, Xu Wang, and Yong Xiang. "Multimodal adversarial network for cross-modal retrieval." Knowledge-Based Systems 180 (September 2019): 38–50. http://dx.doi.org/10.1016/j.knosys.2019.05.017.

Full text
APA, Harvard, Vancouver, ISO, and other styles
26

Vitay, Julien, and Fred H. Hamker. "Sustained Activities and Retrieval in a Computational Model of the Perirhinal Cortex." Journal of Cognitive Neuroscience 20, no. 11 (2008): 1993–2005. http://dx.doi.org/10.1162/jocn.2008.20147.

Full text
Abstract:
The perirhinal cortex is involved not only in object recognition and novelty detection but also in multimodal integration, reward association, and visual working memory. We propose a computational model that focuses on the role of the perirhinal cortex in working memory, particularly with respect to sustained activities and memory retrieval. This model describes how different partial informations are integrated into assemblies of neurons that represent the identity of an object. Through dopaminergic modulation, the resulting clusters can retrieve the global information with recurrent interacti
APA, Harvard, Vancouver, ISO, and other styles
27

Chávez, Ricardo Omar, Hugo Jair Escalante, Manuel Montes-y-Gómez, and Luis Enrique Sucar. "Multimodal Markov Random Field for Image Reranking Based on Relevance Feedback." ISRN Machine Vision 2013 (February 11, 2013): 1–16. http://dx.doi.org/10.1155/2013/428746.

Full text
Abstract:
This paper introduces a multimodal approach for reranking of image retrieval results based on relevance feedback. We consider the problem of reordering the ranked list of images returned by an image retrieval system, in such a way that relevant images to a query are moved to the first positions of the list. We propose a Markov random field (MRF) model that aims at classifying the images in the initial retrieval-result list as relevant or irrelevant; the output of the MRF is used to generate a new list of ranked images. The MRF takes into account (1) the rank information provided by the initial
APA, Harvard, Vancouver, ISO, and other styles
28

Wang, Xurui. "The application of NLP in information retrieval." Applied and Computational Engineering 42, no. 1 (2024): 290–97. http://dx.doi.org/10.54254/2755-2721/42/20230795.

Full text
Abstract:
The field of Natural Language Processing (NLP) has experienced impressive advancements and has found diverse applications. This paper presents a comprehensive review of the development of NLP in the field of information retrieval. It explores different stages of NLP techniques and methods, including keyword matching, rule-based approaches, statistical methods, and the utilization of machine learning and deep learning technologies. Furthermore, the paper provides detailed insights into the specific applications of NLP in domains such as academic information retrieval, medical information retrie
APA, Harvard, Vancouver, ISO, and other styles
29

Waykar, Sanjay B., and C. R. Bharathi. "Multimodal Features and Probability Extended Nearest Neighbor Classification for Content-Based Lecture Video Retrieval." Journal of Intelligent Systems 26, no. 3 (2017): 585–99. http://dx.doi.org/10.1515/jisys-2016-0041.

Full text
Abstract:
AbstractDue to the ever-increasing number of digital lecture libraries and lecture video portals, the challenge of retrieving lecture videos has become a very significant and demanding task in recent years. Accordingly, the literature presents different techniques for video retrieval by considering video contents as well as signal data. Here, we propose a lecture video retrieval system using multimodal features and probability extended nearest neighbor (PENN) classification. There are two modalities utilized for feature extraction. One is textual information, which is determined from the lectu
APA, Harvard, Vancouver, ISO, and other styles
30

Zou, Qiang, Shuli Cheng, Anyu Du, and Jiayi Chen. "Text-Enhanced Graph Attention Hashing for Cross-Modal Retrieval." Entropy 26, no. 11 (2024): 911. http://dx.doi.org/10.3390/e26110911.

Full text
Abstract:
Deep hashing technology, known for its low-cost storage and rapid retrieval, has become a focal point in cross-modal retrieval research as multimodal data continue to grow. However, existing supervised methods often overlook noisy labels and multiscale features in different modal datasets, leading to higher information entropy in the generated hash codes and features, which reduces retrieval performance. The variation in text annotation information across datasets further increases the information entropy during text feature extraction, resulting in suboptimal outcomes. Consequently, reducing
APA, Harvard, Vancouver, ISO, and other styles
31

Zhang, Guihao, and Jiangzhong Cao. "Feature Fusion Based on Transformer for Cross-modal Retrieval." Journal of Physics: Conference Series 2558, no. 1 (2023): 012012. http://dx.doi.org/10.1088/1742-6596/2558/1/012012.

Full text
Abstract:
Abstract With the popularity of the Internet and the rapid growth of multimodal data, multimodal retrieval has gradually become a hot area of research. As one of the important branches of multimodal retrieval, image-text retrieval aims to design a model to learn and align two modal data, image and text, in order to build a bridge of semantic association between the two heterogeneous data, so as to achieve unified alignment and retrieval. The current mainstream image-text cross-modal retrieval approaches have made good progress by designing a deep learning-based model to find potential associat
APA, Harvard, Vancouver, ISO, and other styles
32

Dong, Bin, Songlei Jian, and Kai Lu. "Learning Multimodal Representations by Symmetrically Transferring Local Structures." Symmetry 12, no. 9 (2020): 1504. http://dx.doi.org/10.3390/sym12091504.

Full text
Abstract:
Multimodal representations play an important role in multimodal learning tasks, including cross-modal retrieval and intra-modal clustering. However, existing multimodal representation learning approaches focus on building one common space by aligning different modalities and ignore the complementary information across the modalities, such as the intra-modal local structures. In other words, they only focus on the object-level alignment and ignore structure-level alignment. To tackle the problem, we propose a novel symmetric multimodal representation learning framework by transferring local str
APA, Harvard, Vancouver, ISO, and other styles
33

Zhang, Hongli. "Voice Keyword Retrieval Method Using Attention Mechanism and Multimodal Information Fusion." Scientific Programming 2021 (January 23, 2021): 1–11. http://dx.doi.org/10.1155/2021/6662841.

Full text
Abstract:
A cross-modal speech-text retrieval method using interactive learning convolution automatic encoder (CAE) is proposed. First, an interactive learning autoencoder structure is proposed, including two inputs of speech and text, as well as processing links such as encoding, hidden layer interaction, and decoding, to complete the modeling of cross-modal speech-text retrieval. Then, the original audio signal is preprocessed and the Mel frequency cepstrum coefficient (MFCC) feature is extracted. In addition, the word bag model is used to extract the text features, and then the attention mechanism is
APA, Harvard, Vancouver, ISO, and other styles
34

Escalante, Hugo Jair, Manuel Montes, and Enrique Sucar. "Multimodal indexing based on semantic cohesion for image retrieval." Information Retrieval 15, no. 1 (2011): 1–32. http://dx.doi.org/10.1007/s10791-011-9170-z.

Full text
APA, Harvard, Vancouver, ISO, and other styles
35

Demner-Fushman, Dina, Sameer Antani, Matthew Simpson, and George R. Thoma. "Design and Development of a Multimodal Biomedical Information Retrieval System." Journal of Computing Science and Engineering 6, no. 2 (2012): 168–77. http://dx.doi.org/10.5626/jcse.2012.6.2.168.

Full text
APA, Harvard, Vancouver, ISO, and other styles
36

Liang, Qi, Ning Xu, Weijie Wang, and Xingjian Long. "Multimodal information fusion based on LSTM for 3D model retrieval." Multimedia Tools and Applications 79, no. 45-46 (2020): 33943–56. http://dx.doi.org/10.1007/s11042-020-08817-6.

Full text
APA, Harvard, Vancouver, ISO, and other styles
37

He, Chao, Dalin Wang, Zefu Tan, Liming Xu, and Nina Dai. "Cross-Modal Discrimination Hashing Retrieval Using Variable Length." Security and Communication Networks 2022 (September 9, 2022): 1–12. http://dx.doi.org/10.1155/2022/9638683.

Full text
Abstract:
Fast cross-modal retrieval technology based on hash coding has become a hot topic for the rich multimodal data (text, image, audio, etc.), especially security and privacy challenges in the Internet of Things and mobile edge computing. However, most methods based on hash coding are only mapped to the common hash coding space, and it relaxes the two value constraints of hash coding. Therefore, the learning of the multimodal hash coding may not be sufficient and effective to express the original multimodal data and cause the hash encoding category to be less discriminatory. For the sake of solvin
APA, Harvard, Vancouver, ISO, and other styles
38

Gonsior, Barbara, Christian Landsiedel, Nicole Mirnig, et al. "Impacts of Multimodal Feedback on Efficiency of Proactive Information Retrieval from Task-Related HRI." Journal of Advanced Computational Intelligence and Intelligent Informatics 16, no. 2 (2012): 313–26. http://dx.doi.org/10.20965/jaciii.2012.p0313.

Full text
Abstract:
This work is a first step towards an integration ofmultimodality with the aim to make efficient use of both human-like, and non-human-like feedback modalities in order to optimize proactive information retrieval from task-related Human-Robot Interaction (HRI) in human environments. The presented approach combines the human-like modalities speech and emotional facial mimicry with non-human-like modalities. The proposed non-human-like modalities are a screen displaying retrieved knowledge of the robot to the human and a pointer mounted above the robot head for pointing directions and referring t
APA, Harvard, Vancouver, ISO, and other styles
39

Mithun, Niluthpol C., Juncheng Li, Florian Metze, and Amit K. Roy-Chowdhury. "Joint embeddings with multimodal cues for video-text retrieval." International Journal of Multimedia Information Retrieval 8, no. 1 (2019): 3–18. http://dx.doi.org/10.1007/s13735-018-00166-3.

Full text
APA, Harvard, Vancouver, ISO, and other styles
40

Antony Vigil M S. "Integration of Retrieval-Augmented Generation and Multimodal Technologies for Advanced Virtual Research Assistants." Journal of Information Systems Engineering and Management 10, no. 37s (2025): 716–29. https://doi.org/10.52783/jisem.v10i37s.6508.

Full text
Abstract:
Researchers now rely on AI-powered IVRAs for a wide range of tasks, including providing instantaneous access to research resources and general academic support. When it comes to complicated, multimodal data and providing personalised, context-sensitive replies, however, current technologies can be inadequate. This study investigates potential solutions to these problems by combining multimodal technology with Retrieval-Augmented Generation (RAG). The RAG assistant can find the right information and put it together in a logical way since it uses generative models in addition to information retr
APA, Harvard, Vancouver, ISO, and other styles
41

Figueroa, Cristhian, Hugo Ordoñez, Juan-Carlos Corrales, Carlos Cobos, Leandro Krug Wives, and Enrique Herrera-Viedma. "Improving business process retrieval using categorization and multimodal search." Knowledge-Based Systems 110 (October 2016): 49–59. http://dx.doi.org/10.1016/j.knosys.2016.07.014.

Full text
APA, Harvard, Vancouver, ISO, and other styles
42

Boughanem, M., C. Chrisment, and L. Tamine. "On using genetic algorithms for multimodal relevance optimization in information retrieval." Journal of the American Society for Information Science and Technology 53, no. 11 (2002): 934–42. http://dx.doi.org/10.1002/asi.10119.

Full text
APA, Harvard, Vancouver, ISO, and other styles
43

Sata, Ikumi, Motoki Amagasaki, and Masato Kiyama. "Multimodal Retrieval Method for Images and Diagnostic Reports Using Cross-Attention." AI 6, no. 2 (2025): 38. https://doi.org/10.3390/ai6020038.

Full text
Abstract:
Background: Conventional medical image retrieval methods treat images and text as independent embeddings, limiting their ability to fully utilize the complementary information from both modalities. This separation often results in suboptimal retrieval performance, as the intricate relationships between images and text remain underexplored. Methods: To address this limitation, we propose a novel retrieval method that integrates medical image and text embeddings using a cross-attention mechanism. Our approach creates a unified representation by directly modeling the interactions between the two
APA, Harvard, Vancouver, ISO, and other styles
44

Wang, Qiang, Wei Zheng, Fan Wu, et al. "Information Fusion for Spaceborne GNSS-R Sea Surface Height Retrieval Using Modified Residual Multimodal Deep Learning Method." Remote Sensing 15, no. 6 (2023): 1481. http://dx.doi.org/10.3390/rs15061481.

Full text
Abstract:
Traditional spaceborne Global Navigation Satellite Systems Reflectometry (GNSS-R) sea surface height (SSH) retrieval methods have the disadvantages of complicated error models, low retrieval accuracy, and difficulty using full DDM information. To compensate for these deficiencies while considering the heterogeneity of the input data, this paper proposes an end-to-end Modified Residual Multimodal Deep Learning (MRMDL) method that can utilize the entire range of DDM information. First, the MRMDL method is constructed based on the modified Residual Net (MResNet) and Multi-Hidden layer neural netw
APA, Harvard, Vancouver, ISO, and other styles
45

Rahman, Md Mahmudur, Daekeun You, Matthew S. Simpson, Sameer K. Antani, Dina Demner-Fushman, and George R. Thoma. "Multimodal biomedical image retrieval using hierarchical classification and modality fusion." International Journal of Multimedia Information Retrieval 2, no. 3 (2013): 159–73. http://dx.doi.org/10.1007/s13735-013-0038-4.

Full text
APA, Harvard, Vancouver, ISO, and other styles
46

Li, Ruxuan, Jingyi Wang, and Xuedong Tian. "A Multi-Modal Retrieval Model for Mathematical Expressions Based on ConvNeXt and Hesitant Fuzzy Set." Electronics 12, no. 20 (2023): 4363. http://dx.doi.org/10.3390/electronics12204363.

Full text
Abstract:
Mathematical expression retrieval is an essential component of mathematical information retrieval. Current mathematical expression retrieval research primarily targets single modalities, particularly text, which can lead to the loss of structural information. On the other hand, multimodal research has demonstrated promising outcomes across different domains, and mathematical expressions in image format are adept at preserving their structural characteristics. So we propose a multi-modal retrieval model for mathematical expressions based on ConvNeXt and HFS to address the limitations of single-
APA, Harvard, Vancouver, ISO, and other styles
47

Peng, Shenao, Zhongmei Wang, Jianhua Liu, Changfan Zhang, and Lin Jia. "Fine-Grained Local and Global Semantic Fusion for Multimodal Image–Text Retrieval." Big Data and Cognitive Computing 9, no. 3 (2025): 53. https://doi.org/10.3390/bdcc9030053.

Full text
Abstract:
An image–text retrieval method that integrates intramodal fine-grained local semantic information and intermodal global semantic information is proposed to address the weak fine-grained discrimination capabilities for the semantic features located between image and text modalities in cross-modal retrieval tasks. First, the original features of images and texts are extracted, and a graph attention network is employed for region relationship reasoning to obtain relation-enhanced local features. Then, an attention mechanism is used for different semantically interacting samples within the same mo
APA, Harvard, Vancouver, ISO, and other styles
48

Lin, Kaiyi, Xing Xu, Lianli Gao, Zheng Wang, and Heng Tao Shen. "Learning Cross-Aligned Latent Embeddings for Zero-Shot Cross-Modal Retrieval." Proceedings of the AAAI Conference on Artificial Intelligence 34, no. 07 (2020): 11515–22. http://dx.doi.org/10.1609/aaai.v34i07.6817.

Full text
Abstract:
Zero-Shot Cross-Modal Retrieval (ZS-CMR) is an emerging research hotspot that aims to retrieve data of new classes across different modality data. It is challenging for not only the heterogeneous distributions across different modalities, but also the inconsistent semantics across seen and unseen classes. A handful of recently proposed methods typically borrow the idea from zero-shot learning, i.e., exploiting word embeddings of class labels (i.e., class-embeddings) as common semantic space, and using generative adversarial network (GAN) to capture the underlying multimodal data structures, as
APA, Harvard, Vancouver, ISO, and other styles
49

Qian, Shengsheng, Dizhan Xue, Huaiwen Zhang, Quan Fang, and Changsheng Xu. "Dual Adversarial Graph Neural Networks for Multi-label Cross-modal Retrieval." Proceedings of the AAAI Conference on Artificial Intelligence 35, no. 3 (2021): 2440–48. http://dx.doi.org/10.1609/aaai.v35i3.16345.

Full text
Abstract:
Cross-modal retrieval has become an active study field with the expanding scale of multimodal data. To date, most existing methods transform multimodal data into a common representation space where semantic similarities between items can be directly measured across different modalities. However, these methods typically suffer from following limitations: 1) They usually attempt to bridge the modality gap by designing losses in the common representation space which may not be sufficient to eliminate potential heterogeneity of different modalities in the common space. 2) They typically treat labe
APA, Harvard, Vancouver, ISO, and other styles
50

Alamri, Mohammed. "Large Language Models and Information Retrieval in the Digital Environment: a theoretical analytical study." International Journal of Computers and Informatics 4, no. 6 (2025): 9–55. https://doi.org/10.59992/ijci.2025.v4n6p1.

Full text
Abstract:
This study explores the role of Large Language Models (LLMs) in information retrieval within the digital environment through a theoretical analysis of their concepts, operational mechanisms, and a comparison with traditional methods, alongside identifying key challenges and contemporary applications. The findings reveal that LLMs represent a qualitative shift in processing natural language due to their ability to understand context and generate precise responses. The study highlights their superiority in enhancing retrieval systems through integration with cognitive technologies such as Retrie
APA, Harvard, Vancouver, ISO, and other styles
We offer discounts on all premium plans for authors whose works are included in thematic literature selections. Contact us to get a unique promo code!