2023

Judging LLM-as-a-Judge With MT-Bench and Chatbot arena

Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Ziyi Wu, Yonghao Zhuang, Zongyu Lin, Zhiyuan Li, Dustin Li, E. Xing

citations

Cite Score

85

AI summary

This paper explores using strong LLMs as judges to evaluate chat assistants on open-ended questions, introducing MT-bench and Chatbot Arena datasets, and demonstrating that GPT-4 can achieve over 80% agreement with human preferences, suggesting LLM-as-a-judge as a scalable evaluation method.

Main Contributions

  • Systematic study of LLM-as-a-judge, identifying and proposing solutions to mitigate biases like position, verbosity, and self-enhancement.
  • Introduction of two human preference datasets: MT-bench (multi-turn question set) and Chatbot Arena (crowdsourced battle platform).
  • Demonstration that strong LLM judges (e.g., GPT-4) can match human preferences with over 80% agreement on both controlled and crowdsourced benchmarks.
  • Proposition of LLM-as-a-judge as a scalable and explainable method for approximating human preferences, complementing traditional benchmarks.
  • Public release of MT-bench questions, 3K expert votes, and 30K conversations with human preferences for future research.

Abstract

Evaluating large language model (LLM) based chat assistants is challenging due to their broad capabilities and the inadequacy of existing benchmarks in measuring human preferences. To address this, we explore using strong LLMs as judges to evaluate these models on more open-ended questions. We examine the usage and limitations of LLM-as-a-judge, including position, verbosity, and self-enhancement biases, as well as limited reasoning ability, and propose solutions to mitigate some of them. We then verify the agreement between LLM judges and human preferences by introducing two benchmarks: MT-bench, a multi-turn question set; and Chatbot Arena, a crowdsourced battle platform. Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans. Hence, LLM-as-a-judge is a scalable and explainable way to approximate human preferences, which are otherwise very expensive to obtain. Additionally, we show our benchmark and traditional benchmarks complement each other by evaluating several variants of LLaMA and Vicuna. The MT-bench questions, 3K expert votes, and 30K conversations with human preferences are publicly available at https://github.com/lm-sys/FastChat/tree/main/fastchat/llm_judge.

Citation Graph

Loading graph...

References [52]

Sort:
Filter:

K. Papineni, S. Roukos, T. Ward, Wei Jing Zhu - 2002

20 papers in library cite

Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, C. Wainwright, Pamela Mishkin, Chiyuan Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, Ryan Lowe - 2022

15 papers in library cite

Chin Yew Lin - 2004

10 papers in library cite

Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, Wojciech Zaremba - 2021

16 papers in library cite

Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, Jacob Steinhardt - 2021

10 papers in library cite

Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, John Schulman - 2021

10 papers in library cite

Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord - 2018

7 papers in library cite

Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, Yejin Choi - 2019

9 papers in library cite

Missing author list

2022

6 papers in library cite

Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aman Gupta, Adria Garriga Alonso - 2022

11 papers in library cite

Siva Reddy, Deli Chen, Christopher D. Manning - 2018

7 papers in library cite

Hugo Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Roziere, N. Goyal, Eric Hambro, F. Azhar - 2023

3 papers in library cite

Jason Wei, Xinpeng Wang, Dale Schuurmans, Maarten Bosma, Fanyue Xia, E. Chi, Quoc V. Le, Denny Zhou - 2022

13 papers in library cite

Openai - 2023

14 papers in library cite

Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer - 2023

1 paper in library cites

S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Yiwei Li, S. Lundberg - 2023

5 papers in library cite

Jason Wei, Maarten Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, Quoc V. Le - 2021

5 papers in library cite

Stephen Lin, Jacob Hilton, Owain Evans - 2022

5 papers in library cite

Yuzhi Wang, Y. Kordi, Swaroop Mishra, A. Liu, Noah A. Smith, Daniel Khashabi, Hananneh Hajishirzi - 2023

2 papers in library cite

K. Sakaguchi, R. L. Bras, C. Bhagavatula, Yejin Choi - 2019

5 papers in library cite

Percy Liang, R. Bommasani, Teddy Lee, D. Tsipras, Dilara Soylu, Michihiro Yasunaga, Y. Z. Zhang, D. Narayanan, Yonghui Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, Chiyuan Zhang, C. Cosgrove, Christopher D. Manning, C. Re, D. A. Navas, D. A. Hudson, E. Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, H. Ren, H. Yao, J. Wang, K. Santhanam, L. Orr, Lianmin Zheng, M. Yuksekgonul, Mirac Suzgun, N. Kim, N. Guha, N. S. Chatterji, Omar Khattab, P. Henderson, Q. Huang, R. A. Chi, S. M. Xie, S. Santurkar, Surya Ganguli, Tatsunori Hashimoto, T. Icard, Tong Zhang, V. Chaudhary, Wenyi Wang, Xiang Lisa Li, Y. Mai, Y. Z. Zhang, Y. Koreeda - 2023

6 papers in library cite

Tri Dao, D. Fu, Stefano Ermon, A. Rudra, C. Re - 2022

3 papers in library cite

F. Gilardi, M. Alizadeh, M. Kubli - 2023

1 paper in library cites

C. H. Chiang, H. Y. Lee - 2023

3 papers in library cite

Peng Wang, Lei Li, L. C. Chen, D. Zhu, B. Lin, Yue Cao, Qian Liu, T. Liu, Zhifang Sui - 2023

3 papers in library cite

Swaroop Mishra, Daniel Khashabi, Chitta Baral, Hananneh Hajishirzi - 2021

7 papers in library cite

A. Gudibande, E. Wallace, C. Snell, X. Geng, Haozhe Liu, P. Abbeel, Sergey Levine, Dawn Song - 2023

1 paper in library cites

W. Zhong, R. Cui, Y. Guo, Yiqing Liang, S. Lu, Yuzhi Wang, A. Saied, Weizhu Chen, N. Duan - 2023

4 papers in library cite

Douwe Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Ziyi Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia - 2021

5 papers in library cite

R. Anil, A. M. Dai, O. Firat, M. J. Johnson, D. Lepikhin, A. Passos, Siamak Shakeri, E. Taropa, P. Bailey, Ziru Chen - 2023

3 papers in library cite

Wei-Lin Chiang, Zhiyuan Li, Zongyu Lin, Ying Sheng, Ziyi Wu, Haowei Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, Eric P. Xing - 2023

3 papers in library cite

Y. Dubois, Xiang Lisa Li, R. Taori, Tong Zhang, I. Gulrajani, Jimmy Lei Ba, C. Guestrin, Percy Liang, T. B. Hashimoto - 2023

2 papers in library cite

B. Peng, Chun-Liang Li, Pengcheng He, M. Galley, Jianfeng Gao - 2023

2 papers in library cite

X. Geng, A. Gudibande, Haozhe Liu, E. Wallace, P. Abbeel, Sergey Levine, Dawn Song - 2023

2 papers in library cite

A. Kopf, Y. Kilcher, D. V. Rutte, S. Anagnostidis, Z. R. Tam, K. Stevens, A. Barhoum, N. M. Duc, O. Stanley, R. Nagyfi - 2023

2 papers in library cite

R. Taori, I. Gulrajani, Tong Zhang, Y. Dubois, Xiang Lisa Li, C. Guestrin, Percy Liang, T. B. Hashimoto - 2023

2 papers in library cite

P. Raghubir, A. Valenzuela - 2006

1 paper in library cites

J. D. Brown - 1986

1 paper in library cites

Yuzhi Wang, H. Ivison, P. Dasigi, J. Hessel, Tushar Khot, Khyathi Raghavi Chandu, D. Wadden, K. Macmillan, Noah A. Smith, I. Beltagy - 2023

1 paper in library cites

Chang Zhou, P. Liu, P. Xu, S. Iyer, Jian Sun, Y. Mao, Xueguang Ma, Avia Efrat, P. Yu, Longhui Yu - 2023

1 paper in library cites

S. Diao, R. Pan, H. Dong, K. S. Shum, J. Zhang, W. Xiong, Tong Zhang - 2023

1 paper in library cites

M. Ko, Jaehoon Lee, Hannah Kim, G. Kim, J. Kang - 2020

1 paper in library cites

J. Feng, Q. Sun, Chenfeng Xu, P. Zhao, Yining Yang, C. Tao, D. Zhao, Q. Lin - 2022

1 paper in library cites

Yuzhi Wang, Z. Yu, Z. Zeng, L. Yang, Caitlin Wang, H. Chen, C. Jiang, Ruobing Xie, J. Wang, X. Xie, W. Ye, S. Zhang, Y. Z. Zhang - 2023

1 paper in library cites

Xinpeng Wang, N. Golbandi, M. Bendersky, D. Metzler, M. Najork - 2018

1 paper in library cites

N. J. Blunch - 1984

1 paper in library cites

Zhilin Yang, Ziyi Wu, M. Luo, Wei-Lin Chiang, R. Bhardwaj, W. Kwon, Siyuan Zhuang, F. S. Luan, G. Mittal, S. Shenker, Ion Stoica - 2023

1 paper in library cites

Yuzhi Wang, Swaroop Mishra, Pegah Alipoormolabashi, Y. Kordi, A. Mirzaei, A. Arunkumar, A. Ashok, A. S. Dhanasekaran, A. Naik, D. Stap - 2022

1 paper in library cites

S. Longpre, Le Hou, T. Vu, A. Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei - 2023

1 paper in library cites

Chenfeng Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, D. Jiang - 2023

1 paper in library cites

Cited by

4

papers in your library

Cites

28

papers in your library

Read

on June 6, 2026

Your review

Tags

Benchmark

Paper Aliases

No aliases