Papers
Evaluating the Performance of RAG Methods for Conversational AI in the Airport Domain
Yuyang Li, Pjm Kerbusch, Rhr Pruim et al.
Evaluating the Prompt Steerability of Large Language Models
Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy et al.
Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
Elif Ecem Umutlu, Ayse Aysu Cengiz, Ahmet Kaan Sever et al.
Evaluating the Robustness and Accuracy of Text Watermarking Under Real-World Cross-Lingual Manipulations
Mansour Al Ghanim, Jiaqi Xue, Rochana Prih Hastuti et al.
Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks
Davide Romano, Jonathan Richard Schwarz, Daniele Giofrè
Evaluating Tokenizer Adaptation Methods for Large Language Models on Low-Resource Programming Languages
Georgy Andryushchenko, Vladimir V. Ivanov
Evaluating Transformers for OCR Post-Correction in Early Modern Dutch Theatre
Florian Debaene, Aaron Maladry, Els Lefever et al.
Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models
Kevin Zhou, Adam Dejl, Gabriel Freedman et al.
Evaluating Vision-Language Models as Evaluators in Path Planning
Mohamed Aghzal, Xiang Yue, Erion Plaku et al.
Evaluating Vision-Language Models for Emotion Recognition
Sree Bhattacharyya, James Z. Wang
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
ChaeHun Park, Yujin Baek, Jaeseok Kim et al.
Evaluating WMT 2025 Metrics Shared Task Submissions on the SSA-MTE African Challenge Set
Senyu Li, Felermino Dario Mario Ali, Jiayi Wang et al.
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
Fan Zhang, Shulin Tian, Ziqi Huang et al.
Evaluation and Facilitation of Online Discussions in the LLM Era: A Survey
Katerina Korre, Dimitris Tsirmpas, Nikos Gkoumas et al.
Evaluation and Incident Prevention in an Enterprise AI Assistant
Akash V. Maharaj, David Arbour, Daniel Lee et al.
Evaluation of Active Feature Acquisition Methods for Time-varying Feature Settings
Henrik von Kleist, Alireza Zamanian, Ilya Shpitser et al.
Evaluation of Attribution Bias in Generator-Aware Retrieval-Augmented Large Language Models
Amin Abolghasemi, Leif Azzopardi, Seyyed Hadi Hashemi et al.
Evaluation of Generated Poetry
David Mareček, Kateřina Motalík Hodková, Tomáš Musil et al.
Evaluation of Generated Poetry
David Mareček, Kateřina Motalík Hodková, Tomáš Musil et al.
Evaluation of Large Language Models on Arabic Punctuation Prediction
Asma Ali Al Wazrah, Afrah Altamimi, Hawra Aljasim et al.
Evaluation of LLM for English to Hindi Legal Domain Machine Translation Systems
Kshetrimayum Boynao Singh, Deepak Kumar, Asif Ekbal
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks
Nikita Soni, Pranav Chitale, Khushboo Singh et al.
Evaluation of LLMs in Medical Text Summarization: The Role of Vocabulary Adaptation in High OOV Settings
Gunjan Balde, Soumyadeep Roy, Mainack Mondal et al.
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
Aneta Zugecova, Dominik Macko, Ivan Srba et al.
Evaluation of Medical Large Language Models: Taxonomy, Review, and Directions
Anisio Lacerda, Gisele Pappa, Adriano César Machado Pereira et al.