Operation is the hardest teacher: estimating DNN accuracy looking for mispredictions (ICSE 2021 - Technical Track)

Who

Antonio Guerriero, Roberto Pietrantuono, Stefano Russo

Track

ICSE 2021 Technical Track

Time Zone

The program is currently displayed in (GMT+02:00) Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna.

Use conference time zone: (GMT+02:00) Amsterdam, Berlin, Bern, Rome, Stockholm, ViennaSelect other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Tue 25 May 2021 10:30 - 10:50 at Blended Sessions Room 3 - 1.1.3. Deep Neural Networks: Validation #1 Chair(s): Oscar Dieste
Tue 25 May 2021 22:30 - 22:50 at Blended Sessions Room 3 - 1.1.3. Deep Neural Networks: Validation #1

Abstract

Deep Neural Networks (DNN) are typically tested for accuracy relying on a set of unlabelled real world data (operational dataset), from which a subset is selected, to be labelled and used as test suite. This subset is required to be small (due to manual labelling) yet faithfully represent the operational context, with the resulting test suite containing roughly the same proportion of examples causing misprediction (i.e., failing test cases) as the operational dataset. However, while testing to estimate accuracy, it is desirable to also learn as much as possible from the failing tests in the operational dataset, since they inform about possible bugs of the DNN. A smart sampling strategy may allow to intentionally include in the test suite many examples causing misprediction, thus providing this way more valuable inputs for DNN improvement while preserving the ability to get trustworthy unbiased estimates. This paper presents a test selection technique (DeepEST) to actively look for failing test cases in the operational dataset of a DNN, with the goal of assessing the DNN expected accuracy by building a small and “informative” test suite, namely with a high number of mispredictions, for subsequent DNN improvement. Experiments with five subjects, combining four DNN models and three datasets, are described. The results show that DeepEST provides DNN accuracy estimates with precision close to (and often better than) those of existing sampling-based DNN testing techniques, while detecting from 5 to 30 times more mispredictions, with the same test suite size.

Link to Preprint

https://arxiv.org/abs/2102.04287

Antonio Guerriero

Università di Napoli Federico II

Roberto Pietrantuono

Università di Napoli Federico II

Stefano Russo

Università di Napoli Federico II

Operation is the hardest teacher: estimating DNN accuracy looking for mispredictions - video

Time Zone

The program is currently displayed in (GMT+02:00) Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna.

Use conference time zone: (GMT+02:00) Amsterdam, Berlin, Bern, Rome, Stockholm, ViennaSelect other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

Display full programSpecify a time band

Save

Session Program

Tue 25 May
Displayed time zone: Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna change

10:30 - 11:30	1.1.3. Deep Neural Networks: Validation #1Technical Track at Blended Sessions Room 3 +12h Chair(s): Oscar Dieste Universidad Politécnica de Madrid

10:30 20m Paper		Operation is the hardest teacher: estimating DNN accuracy looking for mispredictionsTechnical Track Technical Track Antonio Guerriero Università di Napoli Federico II, Roberto Pietrantuono Università di Napoli Federico II, Stefano Russo Università di Napoli Federico II Pre-print Media Attached
10:50 20m Paper		AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemTechnical Track Technical Track Xiaoyu Zhang Xi'an Jiaotong University, Juan Zhai Rutgers University, Shiqing Ma Rutgers University, Chao Shen Xi'an Jiaotong University Pre-print Media Attached
11:10 20m Paper		Self-Checking Deep Neural Networks in DeploymentTechnical Track Technical Track Yan Xiao National University of Singapore, Ivan Beschastnikh University of British Columbia, David Rosenblum George Mason University, Changsheng Sun National University of Singapore, Sebastian Elbaum University of Virginia, Yun Lin National University of Singapore, Jin Song Dong National University of Singapore Pre-print Media Attached

22:30 - 23:30	1.1.3. Deep Neural Networks: Validation #1Technical Track at Blended Sessions Room 3

22:30 20m Paper		Operation is the hardest teacher: estimating DNN accuracy looking for mispredictionsTechnical Track Technical Track Antonio Guerriero Università di Napoli Federico II, Roberto Pietrantuono Università di Napoli Federico II, Stefano Russo Università di Napoli Federico II Pre-print Media Attached
22:50 20m Paper		AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemTechnical Track Technical Track Xiaoyu Zhang Xi'an Jiaotong University, Juan Zhai Rutgers University, Shiqing Ma Rutgers University, Chao Shen Xi'an Jiaotong University Pre-print Media Attached
23:10 20m Paper		Self-Checking Deep Neural Networks in DeploymentTechnical Track Technical Track Yan Xiao National University of Singapore, Ivan Beschastnikh University of British Columbia, David Rosenblum George Mason University, Changsheng Sun National University of Singapore, Sebastian Elbaum University of Virginia, Yun Lin National University of Singapore, Jin Song Dong National University of Singapore Pre-print Media Attached