The NoTape project hypothesizes that machine learning can be more efficient and easier to interpret if the underlying distance measure used for comparing data is allowed to locally adapt. For example, if data consist of observations of human body shape, the NoTape position is that it may be beneficial to allow for different distance measures when comparing thin people than when comparing heavy ones. This may sound as being a predominantly theoretical exercise, but the view has direct benefits to both society and applied branches of science.
One of the key tasks in machine learning is to learn representations of observational data that are suitable for a given task. For instance, if the machine learning system should learn to differentiate between ill and well patients, it may seek a representation in which these groups are well apart. For knowledge discovery tasks it is, however, less clear what constitutes a good representation. For instance, when analyzing biological data to design new drugs, we often do not know precisely what we are looking for, and therefore resort to learning compact representations of data in the hope that this will discard insignificant parts of data. With modern techniques based on neural networks and deep learning, this approach has become applied across disciplines, but at a cost. When compressing data into a compact representation we often observe that large distortions in our data; we see groupings of data that are purely compression artifacts, and we see that almost identical data become dissimilar in the compressed representation. To make matters worse, we often observe that if we re-run an algorithm on the same data it may recover significantly different representations. This can lead to misinterpretations of data and to phrasings of incorrect scientific hypotheses.
One of the key contributions of the NoTape project is a mathematical solution to this problem that can be easily incorporated into existing models. By allowing the distance measure of the learned representation to locally adapt it can be designed such as to compensate for compression artifacts in the representation. Statistically, this can be seen as a partial solution to the decades-old "identifiability problem" that has plagued latent variable models. We have shown that under this approach, distances between compressed representations become identical across algorithmic runs, and retain the key information of the observed data. In this view, we avoid drawing conclusions from the new representation that isn't grounded in the data, thereby limiting the risk of misinterpreting the data.
One of the biggest risks with automated decision-making by artificially intelligent agents is that they may misinterpret the data on which they are trained. By systematically removing compression artifacts from data in learned representations, we drastically limit this risk.
Mathematically, NoTape is concerned with the study of random Riemannian metrics; that is distance measures that not only change throughout space but also come with a stochastic aspect. We may think of this data as drawn on a rubber sheet that stretches and wobbles as we try to measure which observations are similar. The NoTape project has developed elementary foundations for this largely unstudied mathematical topic.