<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN"
  "https://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink"
         xmlns:mml="http://www.w3.org/1998/Math/MathML"
         article-type="research-article"
         dtd-version="1.2">

  <!-- ============================================================ FRONT -->
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJLTEMAS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Latest Technology in Engineering, Management &amp; Applied Science (IJLTEMAS)</journal-title>
        <abbrev-journal-title abbrev-type="publisher">IJLTEMAS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2278-2540</issn>
      <publisher>
        <publisher-name>IJLTEMAS</publisher-name>
      </publisher>
    </journal-meta>

    <article-meta>
      <!-- IDs -->
      <article-id pub-id-type="publisher-id">310</article-id>
            <article-id pub-id-type="doi">10.51583/IJLTEMAS.2026.150800116</article-id>
      
      <!-- Categories -->
            <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Machine Learning</subject>
        </subj-group>
      </article-categories>
      
      <!-- Title -->
      <title-group>
        <article-title>Approaching Early Dyslexia Prediction with Ensemble Machine Learning: A Secondary Empirical Analysis of the Rello Et Al. Public Behavioural Dataset</article-title>
      </title-group>

      <!-- Authors -->
      <contrib-group>
                <contrib contrib-type="author">
                    <name>
            <surname>Ufomba Nwogu</surname>
            <given-names>Patrick</given-names>
          </name>
                              <aff>
            (PhD Researcher on Management Information Systems) National Open University of Nigeria                        <country>Nigeria</country>
                      </aff>
                    
        </contrib>
                <contrib contrib-type="author">
                    <name>
            <surname>Chigozirim Ajaegbu</surname>
            <given-names>Prof.</given-names>
          </name>
                              <aff>
            (PhD Researcher on Management Information Systems) National Open University of Nigeria                        <country>Nigeria</country>
                      </aff>
                    
        </contrib>
                <contrib contrib-type="author">
                    <name>
            <surname>Faruk Umar Ambursa</surname>
            <given-names>Dr.</given-names>
          </name>
                              <aff>
            (PhD Researcher on Management Information Systems) National Open University of Nigeria                        <country>Nigeria</country>
                      </aff>
                    
        </contrib>
                <contrib contrib-type="author">
                    <name>
            <surname>Femi Adeluyi</surname>
            <given-names>Dr.</given-names>
          </name>
                              <aff>
            (PhD Researcher on Management Information Systems) National Open University of Nigeria                        <country>Nigeria</country>
                      </aff>
                    
        </contrib>
              </contrib-group>

      <!-- Volume / Issue / Pages -->
            <volume>15</volume>
                  <issue>8</issue>
                        <fpage>1602</fpage>
            <lpage>1619</lpage>
            
      <!-- Dates -->
      <history>
                <date date-type="received">
          <day>03</day>
          <month>09</month>
          <year>2026</year>
        </date>
                        <date date-type="accepted">
          <day>08</day>
          <month>09</month>
          <year>2026</year>
        </date>
              </history>

            <pub-date pub-type="epub">
        <day>19</day>
        <month>09</month>
        <year>2026</year>
      </pub-date>
      
      <!-- DOI Self-URI -->
            <self-uri xlink:href="https://doi.org/10.51583/IJLTEMAS.2026.150800116"/>
      
      <!-- Keywords -->
            <kwd-group kwd-group-type="author">
                <kwd>Dyslexia; Ensemble Machine Learning; Secondary Empirical Study; Random Forest; XGBoost; Extra Trees; RF-RFE; Stacking; Behavioural Data; Early Prediction; Educational Artificial Intelligence.</kwd>
              </kwd-group>
      
    </article-meta>
  </front>

  <!-- ============================================================ BODY (Abstract) -->
  <body>
        <sec>
      <title>Abstract</title>
      <p>Dyslexia is a specific learning disorder associated with persistent difficulties in reading accuracy, decoding, spelling and fluent word recognition. Early identification is important because timely educational support can reduce the educational consequences of undetected reading difficulties. The increasing availability of digital behavioural data provides an opportunity to complement conventional assessment with machine-learning-based screening approaches. This study presents a secondary empirical investigation of ensemble machine learning for early dyslexia prediction using the publicly available behavioural dataset originally reported by Rello et al. The study uses the Dyt-desktop dataset containing 3,644 observations and 196 predictive variables derived from demographic characteristics and performance across 32 linguistic tasks. The target distribution consists of 392 dyslexia cases and 3,252 non-dyslexia cases, producing a substantial class imbalance. Three tree-based ensemble learners—Random Forest (RF), Extreme Gradient Boosting (XGBoost) and Extra Trees (ET)—were evaluated. Random-Forest Recursive Feature Elimination (RF-RFE) was additionally investigated as a feature-selection strategy, while probability-level stacking combined predictions from RF, XGBoost and ET. Models were evaluated using accuracy, precision, recall, specificity, F1-score, ROC-AUC and Matthews Correlation Coefficient (MCC). XGBoost produced the strongest individual baseline performance, achieving 90.64% accuracy, 64.41% precision, 29.08% recall, 98.06% specificity, 40.07% F1-score, 0.8845 ROC-AUC and 0.3912 MCC. RF-RFE + XGBoost improved performance to 90.81% accuracy, 31.89% recall, 42.74% F1-score, 0.8858 ROC-AUC and 0.4122 MCC. The heterogeneous stacking model achieved substantially higher recall at 71.17%, together with an F1-score of 51.71% and MCC of 0.4644, although accuracy and specificity declined to 85.70% and 87.45%, respectively. The findings demonstrate that ensemble architecture materially affects the operating characteristics of dyslexia prediction. In particular, accuracy-oriented models and sensitivity-oriented screening models represent different deployment objectives. The study contributes a secondary empirical evaluation of the Rello dataset and demonstrates that heterogeneous ensemble learning can provide a potentially useful sensitivity-oriented screening strategy. Nevertheless, the findings should be interpreted as predictive screening evidence rather than clinical diagnosis, and independent external validation, threshold optimization, calibration and explainable artificial intelligence are required before practical deployment.</p>
    </sec>
      </body>

  <!-- ============================================================ BACK (References) -->
    <back>
    <ref-list>
      <title>References</title>
            <ref id="ref1">
        <label>1</label>
        <mixed-citation>G. R. Lyon, S. E. Shaywitz, and B. A. Shaywitz, “A definition of dyslexia,” Annals of Dyslexia, vol. 53, pp. 1–14, 2003.</mixed-citation>
      </ref>
            <ref id="ref2">
        <label>2</label>
        <mixed-citation>L. Rello, R. Baeza-Yates, A. Ali, J. P. Bigham, and M. Serra, “Predicting risk of dyslexia with an online gamified test,” PLOS ONE, vol. 15, no. 12, e0241687, 2020.</mixed-citation>
      </ref>
            <ref id="ref3">
        <label>3</label>
        <mixed-citation>L. Breiman, “Random forests,” Machine Learning, vol. 45, pp. 5–32, 2001.</mixed-citation>
      </ref>
            <ref id="ref4">
        <label>4</label>
        <mixed-citation>T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794, 2016.</mixed-citation>
      </ref>
            <ref id="ref5">
        <label>5</label>
        <mixed-citation>P. Geurts, D. Ernst, and L. Wehenkel, “Extremely randomized trees,” Machine Learning, vol. 63, pp. 3–42, 2006.</mixed-citation>
      </ref>
            <ref id="ref6">
        <label>6</label>
        <mixed-citation>Y. Alkhurayyif and A. R. W. Sait, “A review of artificial intelligence-based dyslexia detection techniques,” Diagnostics, vol. 14, no. 21, 2362, 2024.</mixed-citation>
      </ref>
            <ref id="ref7">
        <label>7</label>
        <mixed-citation>D. H. Wolpert, “Stacked generalization,” Neural Networks, vol. 5, no. 2, pp. 241–259, 1992.</mixed-citation>
      </ref>
            <ref id="ref8">
        <label>8</label>
        <mixed-citation>S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems, vol. 30, pp. 4765–4774, 2017.</mixed-citation>
      </ref>
            <ref id="ref9">
        <label>9</label>
        <mixed-citation>J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,” Journal of Machine Learning Research, vol. 13, pp. 281–305, 2012.</mixed-citation>
      </ref>
            <ref id="ref10">
        <label>10</label>
        <mixed-citation>J. Snoek, H. Larochelle, and R. P. Adams, “Practical Bayesian optimization of machine learning algorithms,” in Advances in Neural Information Processing Systems, vol. 25, 2012.</mixed-citation>
      </ref>
            <ref id="ref11">
        <label>11</label>
        <mixed-citation>F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.</mixed-citation>
      </ref>
          </ref-list>
  </back>
  
</article>
