2016

The Mythos of Model Interpretability

Zachary C. Lipton

citations

Cite Score

81

AI summary

This paper analyzes the concept of interpretability in machine learning models, highlighting the diverse motivations and notions associated with it, and clarifying that interpretability is not a monolithic concept but reflects several distinct ideas, seeking to bring focus to the dialogue.

Main Contributions

  • Examines the motivations underlying interest in interpretability, finding them to be diverse and occasionally discordant.
  • Addresses model properties and techniques thought to confer interpretability, identifying transparency to humans and post-hoc explanations as competing notions.
  • Discusses the feasibility and desirability of different notions of interpretability.
  • Questions the assertions that linear models are interpretable and that deep neural networks are not.
  • Provides a comprehensive taxonomy of both the desiderata and methods in interpretability research.

Abstract

Supervised machine learning models boast remarkable predictive capabilities. But can you trust your model? Will it work in deployment? What else can it tell you about the world? We want models to be not only good, but interpretable. And yet the task of interpretation appears underspecified. Papers provide diverse and sometimes non-overlapping motivations for interpretability, and offer myriad notions of what attributes render models interpretable. Despite this ambiguity, many papers proclaim interpretability axiomatically, absent further explanation. In this paper, we seek to refine the discourse on interpretability. First, we examine the motivations underlying interest in interpretability, finding them to be diverse and occasionally discordant. Then, we address model properties and techniques thought to confer interpretability, identifying transparency to humans and post-hoc explanations as competing notions. Throughout, we discuss the feasibility and desirability of different notions, and question the oft-made assertions that linear models are interpretable and that deep neural networks are not.

Citation Graph

Loading graph...

References [31]

Sort:
Filter:

Tomas Mikolov, Ilya Sutskever, K. Chen, G. S. Corrado, Jeffrey Dean - 2013

32 papers in library cite

Geoffrey Hinton - 2008

7 papers in library cite

Rob Fergus - 2014

7 papers in library cite

M. T. Ribeiro, Shivalika Singh, C. Guestrin - 2016

1 paper in library cites

A. Mordvintsev, Christopher Olah, M. Tyka - 2015

2 papers in library cite

R. Tibshirani - 1996

4 papers in library cite

J. Mcauley, J. Leskovec - 2013

2 papers in library cite

Y. Lou, Rich Caruana, J. Gehrke, G. Hooker - 2013

1 paper in library cites

J. Huysmans, K. Dejaeger, C. Mues, J. Vanthienen, B. Baesens - 2011

1 paper in library cites

C. L. Liu, P. Rani, N. Sarkar - 2005

1 paper in library cites

Rich Caruana, H. Kangarloo, J. Dionisio, U. Sinha, D. Johnson - 1999

1 paper in library cites

J. Pearl - 2009

1 paper in library cites

K. Simonyan, A. Vedaldi, Andrew Zisserman - 2013

1 paper in library cites

Zhengtao Wang, N. D. Freitas, M. Lanctot - 2015

1 paper in library cites

B. Goodman, S. Flaxman - 2016

1 paper in library cites

A. Chouldechova - 2016

1 paper in library cites

F. D. Velez, B. Wallace, R. A. Adams - 2015

1 paper in library cites

Y. Lou, Rich Caruana, J. Gehrke - 2012

1 paper in library cites

Rich Caruana, Y. Lou, J. Gehrke, P. Koch, M. Sturm, N. Elhadad - 2015

1 paper in library cites

B. Kim, E. Glassman, B. Johnson, J. Shah - 2015

1 paper in library cites

G. Ridgeway, D. Madigan, T. Richardson, J. O'kane - 1998

1 paper in library cites

F. I. Corporation - 2011

1 paper in library cites

S. Krening, B. Harrison, K. Feigh, C. Isbell, M. Riedl, A. Thomaz - 2016

1 paper in library cites

A. D. Dragan, K. C. Lee, S. S. Srinivasa - 2013

1 paper in library cites

S. Athey, G. W. Imbens - 2015

1 paper in library cites

Zachary C. Lipton, D. C. Kale, R. Wetzel - 2016

1 paper in library cites

J. Chang, S. Gerrish, Caitlin Wang, J. L. B. Graber, D. M. Blei - 2009

1 paper in library cites

H. X. Wang, L. Fratiglioni, G. B. Frisoni, M. Viitanen, B. Winblad - 1999

1 paper in library cites

B. Kim, C. Rudin, J. A. Shah - 2014

1 paper in library cites

A. Mahendran, A. Vedaldi - 2015

1 paper in library cites

Cited by

1

papers in your library

Cites

5

papers in your library

Read

on November 22, 2025

Your review

Tags

Paper Aliases

No aliases