Papers
19 papers found
Automated evaluation of written discourse coherence using GPT-4
Ben Naismith, Phoebe Mulcaire, Jill Burstein
Rating Short L2 Essays on the CEFR Scale with GPT-4
Kevin P. Yancey, Geoffrey Laflair, Anthony Verardi et al.
SummQA at MEDIQA-Chat 2023: In-Context Learning with GPT-4 for Medical Summarization
Yash Mathur, Sanketh Rangreji, Raghav Kapoor et al.
G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment
Yang Liu, Dan Iter, Yichong Xu et al.
Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4
Kellin Pelrine, Anne Imouza, Camille Thibault et al.
DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4
Yebowen Hu, Kaiqiang Song, Sangwoo Cho et al.
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks
Andrea Sottana, Bin Liang, Kai Zou et al.
Unnatural Error Correction: GPT-4 Can Almost Perfectly Handle Unnatural Scrambled Text
Qi Cao, Takeshi Kojima, Yutaka Matsuo et al.
Exploring the Boundaries of GPT-4 in Radiology
Qianchu Liu, Stephanie Hyland, Shruthi Bannur et al.
Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks
Xianzhi Li, Samuel Chan, Xiaodan Zhu et al.
FinePrompt: Unveiling the Role of Finetuned Inductive Bias on Compositional Reasoning in GPT-4
Jeonghwan Kim, Giwon Hong, Sung-Hyon Myaeng et al.
Information Extraction from Legal Wills: How Well Does GPT-4 Do?
Alice Kwak, Cheonkam Jeong, Gaetano Forte et al.
GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captions
Ting-Yao Hsu, Chieh-Yang Huang, Ryan Rossi et al.
Is GPT-4 a Good Data Analyst?
Liying Cheng, Xingxuan Li, Lidong Bing
Leveraging GPT-4 for Automatic Translation Post-Editing
Vikas Raunak, Amr Sharaf, Yiren Wang et al.
AraDetector at ArAIEval Shared Task: An Ensemble of Arabic-specific pre-trained BERT and GPT-4 for Arabic Disinformation Detection
Ahmed Bahaaulddin, Vian Sabeeh, Hanan Belhaj et al.
From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting
Griffin Adams, Alex Fabbri, Faisal Ladhak et al.
GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4
Tom Kocmi, Christian Federmann