Approaching Early Dyslexia Prediction with Ensemble Machine Learning: A Secondary Empirical Analysis of the Rello Et Al. Public Behavioural Dataset
Authors
(PhD Researcher on Management Information Systems) National Open University of Nigeria (Nigeria)
(PhD Researcher on Management Information Systems) National Open University of Nigeria (Nigeria)
(PhD Researcher on Management Information Systems) National Open University of Nigeria (Nigeria)
(PhD Researcher on Management Information Systems) National Open University of Nigeria (Nigeria)
Article Information
DOI: 10.51583/IJLTEMAS.2026.150800116
Subject Category: Machine Learning
Volume/Issue: 15/8 | Page No: 1602-1619
Publication Timeline
Submitted: 2026-09-03
Accepted: 2026-09-08
Published: 2026-09-19
Abstract
Dyslexia is a specific learning disorder associated with persistent difficulties in reading accuracy, decoding, spelling and fluent word recognition. Early identification is important because timely educational support can reduce the educational consequences of undetected reading difficulties. The increasing availability of digital behavioural data provides an opportunity to complement conventional assessment with machine-learning-based screening approaches. This study presents a secondary empirical investigation of ensemble machine learning for early dyslexia prediction using the publicly available behavioural dataset originally reported by Rello et al. The study uses the Dyt-desktop dataset containing 3,644 observations and 196 predictive variables derived from demographic characteristics and performance across 32 linguistic tasks. The target distribution consists of 392 dyslexia cases and 3,252 non-dyslexia cases, producing a substantial class imbalance. Three tree-based ensemble learners—Random Forest (RF), Extreme Gradient Boosting (XGBoost) and Extra Trees (ET)—were evaluated. Random-Forest Recursive Feature Elimination (RF-RFE) was additionally investigated as a feature-selection strategy, while probability-level stacking combined predictions from RF, XGBoost and ET. Models were evaluated using accuracy, precision, recall, specificity, F1-score, ROC-AUC and Matthews Correlation Coefficient (MCC). XGBoost produced the strongest individual baseline performance, achieving 90.64% accuracy, 64.41% precision, 29.08% recall, 98.06% specificity, 40.07% F1-score, 0.8845 ROC-AUC and 0.3912 MCC. RF-RFE + XGBoost improved performance to 90.81% accuracy, 31.89% recall, 42.74% F1-score, 0.8858 ROC-AUC and 0.4122 MCC. The heterogeneous stacking model achieved substantially higher recall at 71.17%, together with an F1-score of 51.71% and MCC of 0.4644, although accuracy and specificity declined to 85.70% and 87.45%, respectively. The findings demonstrate that ensemble architecture materially affects the operating characteristics of dyslexia prediction. In particular, accuracy-oriented models and sensitivity-oriented screening models represent different deployment objectives. The study contributes a secondary empirical evaluation of the Rello dataset and demonstrates that heterogeneous ensemble learning can provide a potentially useful sensitivity-oriented screening strategy. Nevertheless, the findings should be interpreted as predictive screening evidence rather than clinical diagnosis, and independent external validation, threshold optimization, calibration and explainable artificial intelligence are required before practical deployment.
Keywords
Dyslexia; Ensemble Machine Learning; Secondary Empirical Study; Random Forest; XGBoost; Extra Trees; RF-RFE; Stacking; Behavioural Data; Early Prediction; Educational Artificial Intelligence.
Downloads
References
1. G. R. Lyon, S. E. Shaywitz, and B. A. Shaywitz, “A definition of dyslexia,” Annals of Dyslexia, vol. 53, pp. 1–14, 2003. [Google Scholar] [Crossref]
2. L. Rello, R. Baeza-Yates, A. Ali, J. P. Bigham, and M. Serra, “Predicting risk of dyslexia with an online gamified test,” PLOS ONE, vol. 15, no. 12, e0241687, 2020. [Google Scholar] [Crossref]
3. L. Breiman, “Random forests,” Machine Learning, vol. 45, pp. 5–32, 2001. [Google Scholar] [Crossref]
4. T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794, 2016. [Google Scholar] [Crossref]
5. P. Geurts, D. Ernst, and L. Wehenkel, “Extremely randomized trees,” Machine Learning, vol. 63, pp. 3–42, 2006. [Google Scholar] [Crossref]
6. Y. Alkhurayyif and A. R. W. Sait, “A review of artificial intelligence-based dyslexia detection techniques,” Diagnostics, vol. 14, no. 21, 2362, 2024. [Google Scholar] [Crossref]
7. D. H. Wolpert, “Stacked generalization,” Neural Networks, vol. 5, no. 2, pp. 241–259, 1992. [Google Scholar] [Crossref]
8. S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems, vol. 30, pp. 4765–4774, 2017. [Google Scholar] [Crossref]
9. J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,” Journal of Machine Learning Research, vol. 13, pp. 281–305, 2012. [Google Scholar] [Crossref]
10. J. Snoek, H. Larochelle, and R. P. Adams, “Practical Bayesian optimization of machine learning algorithms,” in Advances in Neural Information Processing Systems, vol. 25, 2012. [Google Scholar] [Crossref]
11. F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011. [Google Scholar] [Crossref]