Vision-Augmented Retrieval and Generation (VARAG) - Vision first RAG Engine
-
Updated
Jul 23, 2025 - Python
Vision-Augmented Retrieval and Generation (VARAG) - Vision first RAG Engine
Multimodal RAG to search and interact locally with technical documents of any kind
Official Implementation of GENIUS: A Generative Framework for Universal Multimodal Search, CVPR 2025
Official code release for ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity (published at ICLR 2022)
ICLR 2026 Oral: WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
Evaluation code and datasets for the ACL 2024 paper, VISTA: Visualized Text Embedding for Universal Multi-Modal Retrieval. The original code and model can be accessed at FlagEmbedding.
Local-first OCR-VLM multimodal document retrieval and independent evaluation prototype.
[CVPR 2025] Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
The official code of "Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search"
The code used to train and run inference with MMDocIR
Embedding model prioritized towards Multimodal RAG, overall + VisDoc double top1 on MMEB benchmark
Explores early fusion and late fusion approaches for Multimodal medical Image Retrieval
This repository contains the dataset and source files to reproduce the results in the publication Müller-Budack et al. 2021: "Multimodal news analytics using measures of cross-modal entity and context consistency", In: International Journal on Multimedia Information Retrieval (IJMIR), Vol. 10, Art. no. 2, 2021.
A Survey of Multimodal Retrieval-Augmented Generation
Formalizing Multimedia Recommendation through Multimodal Deep Learning, accepted in ACM Transactions on Recommender Systems.
Multimodal retrieval in art with context embeddings.
Official Implementation of "Composed Object Retrieval: Object-level Retrieval via Composed Expressions"
[EMNLP 2025] M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
A list of research papers on knowledge-enhanced multimodal learning
The official code of "Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search" [SIGIR 2026]
To associate your repository with the multimodal-retrieval topic, visit your repo's landing page and select "manage topics."