https://www.mdu.se/

mdu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Smart Data Driven Decision Trees Ensemble Methodology for Imbalanced Big Data
Univ Granada, Andalusian Res Inst Data Sci & Computat Intelligen, Dept Software Engn, Granada 18071, Spain..ORCID iD: 0000-0002-1927-8673
Univ Granada, Andalusian Res Inst Data Sci & Computat Intelligen, Dept Comp Sci & Artificial Intelligence, Granada 18071, Spain..
Mälardalen University, School of Innovation, Design and Engineering, Embedded Systems.ORCID iD: 0000-0001-9857-4317
Univ Granada, Andalusian Res Inst Data Sci & Computat Intelligen, Dept Comp Sci & Artificial Intelligence, Granada 18071, Spain..
2024 (English)In: Cognitive Computation, ISSN 1866-9956, E-ISSN 1866-9964, Vol. 16, no 4, p. 1572-1588Article in journal (Refereed) Published
Abstract [en]

Differences in data size per class, also known as imbalanced data distribution, have become a common problem affecting data quality. Big Data scenarios pose a new challenge to traditional imbalanced classification algorithms, since they are not prepared to work with such amount of data. Split data strategies and lack of data in the minority class due to the use of MapReduce paradigm have posed new challenges for tackling the imbalance between classes in Big Data scenarios. Ensembles have been shown to be able to successfully address imbalanced data problems. Smart Data refers to data of enough quality to achieve high-performance models. The combination of ensembles and Smart Data, achieved through Big Data preprocessing, should be a great synergy. In this paper, we propose a novel Smart Data driven Decision Trees Ensemble methodology for addressing the imbalanced classification problem in Big Data domains, namely SD_DeTE methodology. This methodology is based on the learning of different decision trees using distributed quality data for the ensemble process. This quality data is achieved by fusing random discretization, principal components analysis, and clustering-based random oversampling for obtaining different Smart Data versions of the original data. Experiments carried out in 21 binary adapted datasets have shown that our methodology outperforms random forest.

Place, publisher, year, edition, pages
SPRINGER , 2024. Vol. 16, no 4, p. 1572-1588
Keywords [en]
Big data, Smart data, Classification, Ensemble, Imbalanced data, Decision tree
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:mdh:diva-69422DOI: 10.1007/s12559-024-10295-zISI: 001236089200002Scopus ID: 2-s2.0-85194861723OAI: oai:DiVA.org:mdh-69422DiVA, id: diva2:1920395
Available from: 2024-12-11 Created: 2024-12-11 Last updated: 2025-10-10Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Xiong, Ning

Search in DiVA

By author/editor
Garcia-Gil, DiegoXiong, Ning
By organisation
Embedded Systems
In the same journal
Cognitive Computation
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 88 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf