conftrace
_
Papers
Trends
Conferences
Explore
More
Authors
Topics
Keywords
Insights
Papers
Trends
Conferences
Explore
Authors
Topics
Keywords
Insights
Achievements
Home
›
Keywords
›
vision-language model
vision-language model
2348 papers
Explore in graph
Also known as
VLM
VL MODEL
VLMS
Co-occurring keywords
multimodal learning
(4645)
zero-shot learning
(3650)
vision language model
(767)
contrastive learning
(4032)
large language model
(13587)
visual question answering
(1017)
transfer learning
(5449)
few-shot learning
(3398)
knowledge distillation
(3725)
semantic segmentation
(3186)
Papers
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
EMNLP 2024
MetaUAS: Universal Anomaly Segmentation with One-Prompt Meta-Learning
NIPS 2024
If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions
EMNLP 2024
MACAROON: Training Vision-Language Models To Be Your Engaged Partners
EMNLP 2024
Chitranuvad: Adapting Multi-lingual LLMs for Multimodal Translation
EMNLP 2024
Enhancing Advanced Visual Reasoning Ability of Large Language Models
EMNLP 2024
Decompose and Compare Consistency: Measuring VLMs’ Answer Reliability via Task-Decomposition Consistency Comparison
EMNLP 2024
Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
EMNLP 2024
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
EMNLP 2024
Multilingual Synopses of Movie Narratives: A Dataset for Vision-Language Story Understanding
EMNLP 2024
Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation
CVPR 2024
MeaCap: Memory-Augmented Zero-shot Image Captioning
CVPR 2024
Hierarchical Intra-modal Correlation Learning for Label-free 3D Semantic Segmentation
CVPR 2024
Hyperbolic Learning with Synthetic Captions for Open-World Detection
CVPR 2024
Learning Object State Changes in Videos: An Open-World Perspective
CVPR 2024
Bayesian Exploration of Pre-trained Models for Low-shot Image Classification
CVPR 2024
PeVL: Pose-Enhanced Vision-Language Model for Fine-Grained Human Action Recognition
CVPR 2024
ConCon-Chi: Concept-Context Chimera Benchmark for Personalized Vision-Language Tasks
CVPR 2024
SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation
CVPR 2024
Discovering Syntactic Interaction Clues for Human-Object Interaction Detection
CVPR 2024
FFF: Fixing Flawed Foundations in Contrastive Pre-Training Results in Very Strong Vision-Language Models
CVPR 2024
Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning
CVPR 2024
Voila-A: Aligning Vision-Language Models with User's Gaze Attention
NIPS 2024
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
NIPS 2024
Bayesian-guided Label Mapping for Visual Reprogramming
NIPS 2024
<
1
…
62
63
64
…
94
>