Synthese. forthcomingMachine learning models based on artificial neural networks have been increasingly used in scientific research and the public domain. These models are notoriously opaque and may conceal caveats in important tasks. Explainable artificial intelligence (XAI) develops network interpretation strategies that reveal how these models work. However, the situation of XAI does not meet its bright expectations: algorithms proliferate, but there are no agreed-upon standards. I suggest that the notion of the experimenters’ regress (ER) in scientific instruments is useful to understand this situation and yield sobering lessons for XAI. I identify a new variant of ER in the design and application of XAI algorithms: the validity of XAI algorithms and the workings of black-box models cannot be secured without assuming the other. This variant shows two distinct features: strategies for breaking the regress are context-specific, and metrics for evaluating XAI algorithms face a second-order regress. They defy XAI’s current approaches of self-justification and search for universal evaluation standards. ER provides two lessons for XAI. First, as various practical, conventional and communal factors can play a role in decision-making, one should be aware of the considerations involved in designing and applying XAI algorithms. Philosophers can contribute to this by analyzing XAI’s “living standards”, the non-universal and unjustified rationales practitioners use. Second, instead of searching for universal standards, those using XAI should reach pragmatic agreements in each specific context of its application based on empirical test results. Therefore, XAI’s support in the trust of AI through transparency also requires empirical validation.
