Item request has been placed! ×
Item request cannot be made. ×
loading  Processing Request

Handling Imbalanced Data Classification With Variational Autoencoding And Random Under-Sampling Boosting

Item request has been placed! ×
Item request cannot be made. ×
loading   Processing Request
  • Additional Information
    • Publication Information:
      Uppsala universitet, Statistiska institutionen
    • Publication Date:
      2020
    • Collection:
      Uppsala University: Publications (DiVA)
    • Abstract:
      In this thesis, a comparison of three different pre-processing methods for imbalanced classification data, is conducted. Variational Autoencoder, Random Under-Sampling Boosting and a hybrid approach of the two, are applied to three imbalanced classification data sets with different class imbalances. A logistic regression (LR) model is fitted to each pre-processed data set and based on its classification performance, the pre-processing methods are evaluated. All three methods shows indications of different advantages when handling class imbalances. For each pre-processed data, the LR-model has is better at correctly classifying minority class observations, compared to a LR-model fitted to the original class imbalanced data sets. Evaluating the overall classification performance, both VAE and RUSBoost shows improving classification results while the hybrid method performs worse for the moderate class imbalanced data and best for the highly imbalanced data.
    • File Description:
      application/pdf
    • Online Access:
      http://urn.kb.se/resolve?urn=urn:nbn:se:uu:diva-412838
    • Rights:
      info:eu-repo/semantics/openAccess
    • Accession Number:
      edsbas.67E342E7