TY - GEN
T1 - An evaluation of feature sets and sampling techniques for de-identification of medical records
AU - Gardner, James
AU - Xiong, Li
AU - Wang, Fusheng
AU - Post, Andrew
AU - Saltz, Joel
AU - Grandison, Tyrone
PY - 2010
Y1 - 2010
N2 - De-identification of text medical records is of critical importance in any health informatics system in order to facilitate research and sharing of medical records. While statistical learning based techniques have shown promising results for de-identification purposes, few such systems are publicly available. It remains a challenge for practitioners to build an accurate and efficient system as it involves a significant amount of feature engineering, i.e. creation and examination of new features used in the system. A comprehensive evaluation is needed to thoroughly understand the effects of different feature sets and potential impacts of sampling and their trade-offs between the often conflicting goals of precision (or positive predictive value), recall (or sensitivity), and efficiency. In this paper, we present the Health Information DE-identification (HIDE) framework and evaluate the open- source software. We present an evaluation of various types of features used in HIDE, and introduce a window sampling technique (only the terms within a specified distance from personal health information are used to train the classifier) and evaluate its effect on both quality and efficiency. Our results show that the context features (previous and next terms) are particularly important and the sampling technique can be used to increase recall with minimal impact on precision. We obtained token-level label precision of 0.967, recall of 0.986 and F-Score of 0.977 when not including true negatives. The overall HIDE system achieves token-level precision of .998, recall of .999, and f-score of .999 on the previous i2b2 challenge task.
AB - De-identification of text medical records is of critical importance in any health informatics system in order to facilitate research and sharing of medical records. While statistical learning based techniques have shown promising results for de-identification purposes, few such systems are publicly available. It remains a challenge for practitioners to build an accurate and efficient system as it involves a significant amount of feature engineering, i.e. creation and examination of new features used in the system. A comprehensive evaluation is needed to thoroughly understand the effects of different feature sets and potential impacts of sampling and their trade-offs between the often conflicting goals of precision (or positive predictive value), recall (or sensitivity), and efficiency. In this paper, we present the Health Information DE-identification (HIDE) framework and evaluate the open- source software. We present an evaluation of various types of features used in HIDE, and introduce a window sampling technique (only the terms within a specified distance from personal health information are used to train the classifier) and evaluate its effect on both quality and efficiency. Our results show that the context features (previous and next terms) are particularly important and the sampling technique can be used to increase recall with minimal impact on precision. We obtained token-level label precision of 0.967, recall of 0.986 and F-Score of 0.977 when not including true negatives. The overall HIDE system achieves token-level precision of .998, recall of .999, and f-score of .999 on the previous i2b2 challenge task.
KW - conditional random fields
KW - de-identification
KW - medical text
UR - https://www.scopus.com/pages/publications/78650948123
U2 - 10.1145/1882992.1883019
DO - 10.1145/1882992.1883019
M3 - Conference contribution
AN - SCOPUS:78650948123
SN - 9781450300308
T3 - IHI'10 - Proceedings of the 1st ACM International Health Informatics Symposium
SP - 183
EP - 190
BT - IHI'10 - Proceedings of the 1st ACM International Health Informatics Symposium
T2 - 1st ACM International Health Informatics Symposium, IHI'10
Y2 - 11 November 2010 through 12 November 2010
ER -