Assessment of Creditworthiness Models Privacy-Preserving Training with Synthetic Data
Keywords: Credit scoring; Generative adversarial networks; Synthetic data; Variational autoencoders
Abstract
Credit scoring models are the primary instrument used by financial institutions to manage credit risk. The scarcity of research on behavioral scoring is due to the difficult data access. Financial institutions have to maintain the privacy and security of borrowersâ information refrain them from collaborating in research initiatives. In this work, we present a methodology that allows us to evaluate the performance of models trained with synthetic data when they are applied to real-world data. Our results show that synthetic data quality is increasingly poor when the number of attributes increases. However, creditworthiness assessment models trained with synthetic data show a reduction of 3% of AUC and 6% of KS when compared with models trained with real data. These results have a significant impact since they encourage credit risk investigation from synthetic data, making it possible to maintain borrowersâ privacy and to address problems that until now have been hampered by the availability of information.
Más información
| Título según WOS: | Assessment of Creditworthiness Models Privacy-Preserving Training with Synthetic Data |
| Título según SCOPUS: | Assessment of Creditworthiness Models Privacy-Preserving Training with Synthetic Data |
| Título de la Revista: | Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) |
| Volumen: | 13469 |
| Editorial: | Springer Science and Business Media Deutschland GmbH |
| Fecha de publicación: | 2022 |
| Página final: | 384 |
| Idioma: | English |
| DOI: |
10.1007/978-3-031-15471-3_32 |
| Notas: | ISI, SCOPUS |