<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN"
  "https://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink"
         xmlns:mml="http://www.w3.org/1998/Math/MathML"
         article-type="research-article"
         dtd-version="1.2">

  <!-- ============================================================ FRONT -->
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJLTEMAS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Latest Technology in Engineering, Management &amp; Applied Science (IJLTEMAS)</journal-title>
        <abbrev-journal-title abbrev-type="publisher">IJLTEMAS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2278-2540</issn>
      <publisher>
        <publisher-name>IJLTEMAS</publisher-name>
      </publisher>
    </journal-meta>

    <article-meta>
      <!-- IDs -->
      <article-id pub-id-type="publisher-id">335</article-id>
            <article-id pub-id-type="doi">10.51583/IJLTEMAS.2026.150800141</article-id>
      
      <!-- Categories -->
            <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Machine Learning</subject>
        </subj-group>
      </article-categories>
      
      <!-- Title -->
      <title-group>
        <article-title>A Comparative Study of CNN and Vision Transformer Models in Alzheimer's Detection</article-title>
      </title-group>

      <!-- Authors -->
      <contrib-group>
                <contrib contrib-type="author">
                    <name>
            <surname>Stephen</surname>
            <given-names>D</given-names>
          </name>
                              <aff>
            Research Scholar, Department of Computer Science and Engineering, Saveetha School of Engineering, Saveetha Institute of Medical and Technical Sciences (SIMATS), Chennai, India                        <country>India</country>
                      </aff>
                    
        </contrib>
                <contrib contrib-type="author">
                    <name>
            <surname>R Samuel Rajesh Babu</surname>
            <given-names>Dr.</given-names>
          </name>
                              <aff>
            Associate Professor, Department of Electronics and Communication Engineering, Saveetha School of Engineering, Saveetha Institute of Medical and Technical Sciences (SIMATS), Chennai, India                        <country>India</country>
                      </aff>
                    
        </contrib>
              </contrib-group>

      <!-- Volume / Issue / Pages -->
            <volume>15</volume>
                  <issue>8</issue>
                        <fpage>1952</fpage>
            <lpage>1960</lpage>
            
      <!-- Dates -->
      <history>
                <date date-type="received">
          <day>03</day>
          <month>09</month>
          <year>2026</year>
        </date>
                        <date date-type="accepted">
          <day>08</day>
          <month>09</month>
          <year>2026</year>
        </date>
              </history>

            <pub-date pub-type="epub">
        <day>26</day>
        <month>09</month>
        <year>2026</year>
      </pub-date>
      
      <!-- DOI Self-URI -->
            <self-uri xlink:href="https://doi.org/10.51583/IJLTEMAS.2026.150800141"/>
      
      <!-- Keywords -->
            <kwd-group kwd-group-type="author">
                <kwd>Index Terms Alzheimer's disease</kwd>
                <kwd>MRI classification</kwd>
                <kwd>hierarchical attention</kwd>
                <kwd>vision transformers</kwd>
                <kwd>deep learning</kwd>
                <kwd>HierAttNet</kwd>
                <kwd>convolutional neural networks.</kwd>
              </kwd-group>
      
    </article-meta>
  </front>

  <!-- ============================================================ BODY (Abstract) -->
  <body>
        <sec>
      <title>Abstract</title>
      <p>In this work, a novel deep learning framework the Hierarchical Attention Network (HierAttNet) is proposed to replace the previously used Vision Transformer (ViT-B/16) and all CNN baselines in the literature to stage Alzheimer's disease (AD) from MRI brain images. The ConvNeXt-Small feature extractor and the two-scale cross-attention module alternately attend to the local patch level and the region level in a hierarchical fashion, which is not reported in the AD neuroimaging literature. We compare HierAttNet with six additional state-of-the-art algorithms, which were not presented in the original paper: EfficientNetV2-S, Swin-T, MobileViT-S, CrossViT-15, DeiT-S and ConvNeXt-S. The following experiments are performed on the Kaggle Alzheimer's MRI dataset (6,400+ images, 4 stages of Alzheimer's). HierAttNet obtains 99.91% accuracy and F1=0.9993 outperforming all the evaluated models and even the previous state-of-the-art ViT-B/16 (99.87%). The performance of all models consistently dropped with data augmentation. In this work, we introduce HierAttNet as a new state-of-the-art in Alzheimer's MRI classification, and a stringent benchmark for subsequent multimodal diagnostics studies.</p>
    </sec>
      </body>

  <!-- ============================================================ BACK (References) -->
    <back>
    <ref-list>
      <title>References</title>
            <ref id="ref1">
        <label>1</label>
        <mixed-citation>M. Tan and Q. V. Le, "EfficientNetV2: Smaller Models and Faster Training," in Proc. ICML, 2021.</mixed-citation>
      </ref>
            <ref id="ref2">
        <label>2</label>
        <mixed-citation>Z. Liu et al., "A ConvNet for the 2020s," in Proc. IEEE CVPR, 2022.</mixed-citation>
      </ref>
            <ref id="ref3">
        <label>3</label>
        <mixed-citation>Z. Liu et al., "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows," in Proc. IEEE ICCV, 2021.</mixed-citation>
      </ref>
            <ref id="ref4">
        <label>4</label>
        <mixed-citation>S. Mehta and M. Rastegari, "MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer," arXiv:2110.02178, 2021.</mixed-citation>
      </ref>
            <ref id="ref5">
        <label>5</label>
        <mixed-citation>C.-F. R. Chen et al., "CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification," in Proc. IEEE ICCV, 2021.</mixed-citation>
      </ref>
            <ref id="ref6">
        <label>6</label>
        <mixed-citation>H. Touvron et al., "Training Data-Efficient Image Transformers &amp; Distillation through Attention," in Proc. ICML, 2021.</mixed-citation>
      </ref>
            <ref id="ref7">
        <label>7</label>
        <mixed-citation>Dosovitskiy et al., "An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale," in Proc. ICLR, 2021.</mixed-citation>
      </ref>
            <ref id="ref8">
        <label>8</label>
        <mixed-citation>Radosavovic et al., "Designing Network Design Spaces," in Proc. IEEE CVPR, 2020.</mixed-citation>
      </ref>
            <ref id="ref9">
        <label>9</label>
        <mixed-citation>J. Hu et al., "Squeeze-and-Excitation Networks," IEEE CVPR, 2018.</mixed-citation>
      </ref>
            <ref id="ref10">
        <label>10</label>
        <mixed-citation>G. Huang et al., "Densely Connected Convolutional Networks," in Proc. IEEE CVPR, 2017.</mixed-citation>
      </ref>
            <ref id="ref11">
        <label>11</label>
        <mixed-citation>P. LaMontagne et al., "OASIS-3: Longitudinal Neuroimaging, Clinical and Cognitive Dataset," medRxiv, 2019.</mixed-citation>
      </ref>
            <ref id="ref12">
        <label>12</label>
        <mixed-citation>H. Touvron et al., "ResMLP: Feedforward networks for image classification with data-efficient training," IEEE TPAMI, 2022.</mixed-citation>
      </ref>
          </ref-list>
  </back>
  
</article>
