Construction and Evaluation of a Public LiverTox Benchmark for Herbal- and Drug-Induced Liver Injury with Duplication-Controlled Synthetic Data Generation
Keywords:
Liver Diseases, Drug-Induced, Phytotherapy, Machine Learning, Data MiningAbstract
Purpose: We constructed a parser-audited public benchmark of herbal-induced liver injury (HILI) and non-herbal drug-induced liver injury (DILI) from LiverTox portable document format (PDF) chapters and assessed duplication-controlled synthesis. Methods: A PDF-based pipeline converted LiverTox narratives and tables into structured patient-episode records, followed by manual audit and correction. Under an agent-level split, three generators were compared: bootstrap with jitter, an unconstrained Bayesian Gaussian mixture model (BGMM), and a quantile-bounded distance-to-closest-record (DCR) rejection BGMM (QBound+DCR BGMM). Evaluation used parser audit metrics, feature-wise Wasserstein distance, correlation preservation, near-duplicate rate, and train-on-synthetic, test-on-real (TSTR) classification. Results: The final benchmark included 64 cases from 31 agents (20 HILI; 44 DILI). The QBound+DCR BGMM model reduced the pooled near-duplicate rate to 35.6% ± 6.3%, compared with 60.4% ± 4.9% for BGMM and 100.0% ± 0.0% for bootstrap with jitter, and eliminated class-conditional near-duplicate generation. However, fidelity worsened. The TSTR AUROC was 0.233 ± 0.119, compared with 0.183 ± 0.083 for BGMM, 0.539 ± 0.280 for bootstrap, and 0.222 for the train-on-real reference. Conclusion: The benchmark supports reproducible evaluation and shows that duplication control improves novelty but introduces a fidelity-utility trade-off.
Downloads
Additional Files
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Tae-Yoon KIM, Jung-Hyun KIM

All papers published in Applied Medical Informatics are licensed under a Creative Commons Attribution (CC BY 4.0) International License.