conftrace
_
Papers
Trends
Conferences
Explore
More
Authors
Topics
Keywords
Insights
Papers
Trends
Conferences
Explore
Authors
Topics
Keywords
Insights
Achievements
Home
›
Keywords
›
vision-language model
vision-language model
2348 papers
Explore in graph
Also known as
VLM
VL MODEL
VLMS
Co-occurring keywords
multimodal learning
(4645)
zero-shot learning
(3650)
vision language model
(767)
contrastive learning
(4032)
large language model
(13587)
visual question answering
(1017)
transfer learning
(5449)
few-shot learning
(3398)
knowledge distillation
(3725)
semantic segmentation
(3186)
Papers
Enhancing Visual Classification using Comparative Descriptors
WACV 2025
Howard University-AI4PC at SemEval-2025 Task 1: Using GPT-4o and CLIP-ViLT to Decode Figurative Language Across Text and Images
ACL 2025
ReVision: A Dataset and Baseline VLM for Privacy-Preserving Task-Oriented Visual Instruction Rewriting
AACL 2025
RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
ICCV 2025
CL-Cross VQA: A Continual Learning Benchmark for Cross-Domain Visual Question Answering
WACV 2025
Unified Framework for Open-World Compositional Zero-Shot Learning
WACV 2025
Treble Counterfactual VLMs: A Causal Approach to Hallucination
EMNLP 2025
DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models
CVPR 2025
An Encoder-Agnostic Weakly Supervised Method for Describing Textures
WACV 2025
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
ACL 2025
CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition
EMNLP 2025
Test-Time Retrieval-Augmented Adaptation for Vision-Language Models
ICCV 2025
Multi-Modal Large Language Models are Effective Vision Learners
WACV 2025
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
CVPR 2025
MBQ: Modality-Balanced Quantization for Large Vision-Language Models
CVPR 2025
Dynamic Group Detection using VLM-augmented Temporal Groupness Graph
ICCV 2025
Concept-pedia: a Wide-coverage Semantically-annotated Multimodal Dataset
EMNLP 2025
Graph-guided Cross-composition Feature Disentanglement for Compositional Zero-shot Learning
ACL 2025
ProcVQA: Benchmarking the Effects of Structural Properties in Mined Process Visualizations on Vision–Language Model Performance
EMNLP 2025
VoCo-LLaMA: Towards Vision Compression with Large Language Models
CVPR 2025
RoboPearls: Editable Video Simulation for Robot Manipulation
ICCV 2025
Reproducible Vision-Language Models Meet Concepts Out of Pre-Training
CVPR 2025
Mixed Signals: Decoding VLMs’ Reasoning and Underlying Bias in Vision-Language Conflict
EMNLP 2025
On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free Approach
CVPR 2025
Making VLMs More Robot-Friendly: Self-Critical Distillation of Low-Level Procedural Reasoning
EMNLP 2025
<
1
…
31
32
33
…
94
>