The Shapley value is a method of wealth distribution defined in the 1950s as the only method that satisfies a set of desirable axioms. In recent years, it has been applied to databases to define measures of responsibility that quantify the contribution of each fact to a given response. These measures—and other similar applications of the Shapley value to different fields—generally have two drawbacks, one conceptual and the other more practical: (1) Shapley’s axioms are invoked to justify them, but these properties—however desirable they may be in economics—are not meaningful in all contexts, and as a result, the resulting measures sometimes exhibit unexpected behavior; (2) Measures based on the Shapley value are often difficult to compute, since in many cases there are very simple queries for which the problem is nonetheless #P-hard, even in terms of data complexity. This thesis extends these accountability measures to queries involving ontologies and addresses these two shortcomings of existing Shapley-value-based accountability measures, both in the context of databases and ontologies.
For the first problem, we reexamine the question of what constitutes a good measure of responsibility for responses to queries; to this end, we identify properties inspired by Shapley’s axioms that are truly desirable in the specific context we are studying.
For the second, we define new measures, which are still based on the Shapley value but are much easier to compute. To do this, we use other “resource functions” to model the instance under study as a cooperative game to which the Shapley value applies. In addition to these conceptual considerations, we study in detail the complexity of all the measures under consideration, focusing primarily on various forms of conjunctive queries—sometimes enriched with unions and negative atoms—as well as ontologies expressed in lightweight description logics from the DL-Lite and EL families. Finally, we draw on the insights into the Shapley value derived from our study of such accountability measures to take a fresh look at other applications of the Shapley value, in particular the SHAP score, a measure widely used in artificial intelligence to explain the results of classifiers.