conftrace
_
Papers
Trends
Conferences
Explore
More
Authors
Topics
Keywords
Insights
Papers
Trends
Conferences
Explore
Authors
Topics
Keywords
Insights
Achievements
Home
›
Keywords
›
vision-language model
vision-language model
2348 papers
Explore in graph
Also known as
VLM
VL MODEL
VLMS
Co-occurring keywords
multimodal learning
(4645)
zero-shot learning
(3650)
vision language model
(767)
contrastive learning
(4032)
large language model
(13587)
visual question answering
(1017)
transfer learning
(5449)
few-shot learning
(3398)
knowledge distillation
(3725)
semantic segmentation
(3186)
Papers
Towards Better Vision-Inspired Vision-Language Models
CVPR 2024
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
CVPR 2024
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
CVPR 2024
Medical Vision-Language Pre-Training for Brain Abnormalities
COLING 2024
CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-spoofing
CVPR 2024
Tools Identification By On-Board Adaptation of Vision-and-Language Models
AAAI 2024
Open Vocabulary Semantic Scene Sketch Understanding
CVPR 2024
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
ACL 2024
VisDiaHalBench: A Visual Dialogue Benchmark For Diagnosing Hallucination in Large Vision-Language Models
ACL 2024
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models
ACL 2024
ChartAssistant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
ACL 2024
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
ACL 2024
Selective “Selective Prediction”: Reducing Unnecessary Abstention in Vision-Language Reasoning
ACL 2024
Question-Instructed Visual Descriptions for Zero-Shot Video Answering
ACL 2024
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
ACL 2024
USE: Universal Segment Embeddings for Open-Vocabulary Image Segmentation
CVPR 2024
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
ACL 2024
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
ACL 2024
ChartInstruct: Instruction Tuning for Chart Comprehension and Reasoning
ACL 2024
Improving Vision-Language Cross-Lingual Transfer with Scheduled Unfreezing
ACL 2024
DetCLIPv3: Towards Versatile Generative Open-vocabulary Object Detection
CVPR 2024
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
ACL 2024
Embodied Language Learning: Opportunities, Challenges, and Future Directions
ACL 2024
Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability
ACL 2024
ZONE: Zero-Shot Instruction-Guided Local Editing
CVPR 2024
<
1
…
60
61
62
…
94
>