CONVISE focused on the improvement of appearance-based gaze estimation methods, based on deep learning. Appearance-based gaze estimation is the regression problem of mapping facial images of (human) subjects to specific gaze directions (or points to the screen). Deep learning methods for gaze estimation work by training a neural network using examples of facial images and gaze direction vectors.
Perhaps the most important challenge in appearance-based gaze estimation is how to deal with the multiple sources of input variance, which can significantly affect the precision of the solution. The first group of variances is related to the external environment. Those can be controlled by defining stringent experimental settings, such as well-defined camera specifications, experiment locations, illumination conditions, subject distances/angles from the sensor, and/or even employing head/chin rests wherever possible. The second group of variances are related to the visual appearance of individual subjects, such as their physical characteristics (e.g. age, gender, skin/eye/face color/dimensions, and ophthalmic health conditions), which require meticulous effort to control.
The standard method for controlling the latter sources of variance is the so-called “calibration process”, which occurs at the start of each eye-tracking data collection process. This process includes the following steps. Crafted images of some calibration patterns are shown to the subjects, e.g. crosses placed at specific, known points of the screen. Subjects are asked to look at those points, while their facial images are collected. Then, those images are used to finetune/re-train the model on the new subjects, thus improving the overall accuracy of the solution for each specific subject. The disadvantages of this approach are that a) it is time-consuming, b) it requires honest cooperation of the subject, and c) it only works when the person collecting the data (e.g. a researcher) is present to verify that the calibration process was successful. To address those challenges, CONVISE developed a novel methodology that addresses the above limitations.
Second, CONVISE conducted a thorough literature review of advertising research that featured eye-tracking experiments, in order to collect eye-tracking data to develop neural networks for eye-tracking prediction. From the beginning of this work, it became evident that standards seem lacking in advertising studies using eye-tracking data. The eye-tracking data collection procedure is affected by several factors starting from the eye-tracking sensor and its configuration, the scene geometry, the data acquisition quality assurance, and the software.
Modern eye-tracking sensors compute a multi-dimensional signal consisting of geometric variables captured every specific time intervals according to the sensor sampling frequency, that includes a) the estimated relative eye positions, in the format of screen coordinates, the pupil dilation data over time. Processing of these raw data results in the eye-tracking variables normally seen in advertsing research papers, e.g. fixation time/counts, saccade duration/count, and their variants or aggregations. Having access to this raw data, along with appropriate reporting of the remainder factors, allows for research reproducibility. However, providing access to raw data is rarely the case in advertising research. The remainder factors include:
• Eye-tracking Equipment. The type of sensor used and its configuration.
• Eye-tracking software and algorithms used to extract eye-tracking measures.
• The geometry of the scene, such as the distance of the subject/participant to the screen, the screen size and its resolution and lighting conditions.
• Data quality assurance practises, such as calibration.
• The precise visual stimuli definition, in terms of AOIs, their size their correspondance with the actual stimuli.
• The eye-tracking measures used and how they have been calculated.
Based on our study featuring 34 recently published articles including eye-tracking experiments in advertising and business research, it has been found that many of the above factors are underreported and significantly vary across eye-tracking experiments.