<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN"
  "https://jats.nlm.nih.gov/publishing/1.2/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink"
         xmlns:mml="http://www.w3.org/1998/Math/MathML"
         article-type="research-article"
         dtd-version="1.2">

  <!-- ============================================================ FRONT -->
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJLTEMAS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Latest Technology in Engineering, Management &amp; Applied Science (IJLTEMAS)</journal-title>
        <abbrev-journal-title abbrev-type="publisher">IJLTEMAS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2278-2540</issn>
      <publisher>
        <publisher-name>IJLTEMAS</publisher-name>
      </publisher>
    </journal-meta>

    <article-meta>
      <!-- IDs -->
      <article-id pub-id-type="publisher-id">64</article-id>
            <article-id pub-id-type="doi">10.51583/IJLTEMAS.2026.150700059</article-id>
      
      <!-- Categories -->
            <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Computer Science</subject>
        </subj-group>
      </article-categories>
      
      <!-- Title -->
      <title-group>
        <article-title>Explainable 3D Convolutional Neural Networks Spatiotemporal Learning for Human Handshake Interaction Recognition</article-title>
      </title-group>

      <!-- Authors -->
      <contrib-group>
                <contrib contrib-type="author">
                    <name>
            <surname>S. Kumaravel</surname>
            <given-names>Dr.</given-names>
          </name>
                              <aff>
            Associate Professor &amp; Head (Retired)Department of Computer Science (Aided)Sri Ramakrishna Mision Vidyalaya College of Arts and Science Coimbatore – 641 020 .Tamilnadu, INDIA.                        <country>India</country>
                      </aff>
                    
        </contrib>
                <contrib contrib-type="author">
                    <name>
            <surname>S. Veni</surname>
            <given-names>Dr.</given-names>
          </name>
                              <aff>
            Professor Department of Computer Science Karpagam Academy of Higher Education (KAHE) Coimbatore-641021. Tamilnadu, INDIA.                        <country>India</country>
                      </aff>
                    
        </contrib>
              </contrib-group>

      <!-- Volume / Issue / Pages -->
            <volume>15</volume>
                  <issue>7</issue>
                        <fpage>722</fpage>
            <lpage>732</lpage>
            
      <!-- Dates -->
      <history>
                <date date-type="received">
          <day>27</day>
          <month>07</month>
          <year>2026</year>
        </date>
                        <date date-type="accepted">
          <day>01</day>
          <month>08</month>
          <year>2026</year>
        </date>
              </history>

            <pub-date pub-type="epub">
        <day>12</day>
        <month>08</month>
        <year>2026</year>
      </pub-date>
      
      <!-- DOI Self-URI -->
            <self-uri xlink:href="https://doi.org/10.51583/IJLTEMAS.2026.150700059"/>
      
      <!-- Keywords -->
            <kwd-group kwd-group-type="author">
                <kwd>Explainable AI</kwd>
                <kwd>3D Convolutional Neural Network (3D CNN)</kwd>
                <kwd>Spatiotemporal Learning</kwd>
                <kwd>Human Activity Recognition (HAR)</kwd>
                <kwd>Handshake Interaction Recognition</kwd>
              </kwd-group>
      
    </article-meta>
  </front>

  <!-- ============================================================ BODY (Abstract) -->
  <body>
        <sec>
      <title>Abstract</title>
      <p>Human Activity Recognition (HAR) has gained significant attention in computer vision due to its wide range of applications in surveillance, social behaviour analysis, and human–computer interaction. Among various human-to-human interactions, handshake recognition is particularly important as it represents social intention and cooperative behaviour. This study presents an efficient and interpretable deep learning framework for automatic handshake recognition from video sequences. The proposed approach employs a pretrained 3D Convolutional Neural Network (3D CNN) to directly learn spatiotemporal features, enabling effective modelling of both motion dynamics and spatial relationships between interacting individuals. The experiments are conducted using two dataset namely UT-Interaction Human Interaction Dataset and SBU Kinect Interaction dataset, focusing exclusively on the handshake interaction as the target class. The dataset provides accurate ground-truth annotations, including temporal intervals and bounding boxes, which support precise localization and reliable recognition of handshake actions. Each dataset is split into 80% for training, 10% for validation, and 10% for testing to ensure robust performance evaluation. The experimental results demonstrated that the proposed 3D CNN-based framework achieved a handshake recognition accuracy of 98.92% on the UT-Interaction dataset, representing performance improvements of 9.72%, 6.32%, and 7.12% compared to CNN, BiLSTM, and RNN models, respectively.</p>
    </sec>
      </body>

  <!-- ============================================================ BACK (References) -->
    <back>
    <ref-list>
      <title>References</title>
            <ref id="ref1">
        <label>1</label>
        <mixed-citation>K. Yaseen, O.-J. Kwon, J. Kim, S. Jamil, J. Lee, and F. Ullah, Next-Gen Dynamic Hand Gesture Recognition: MediaPipe, Inception-v3 and LSTM-Based Enhanced Deep Learning Model, vol. 12, pp. 117233–117245, 2024, doi: 10.1109/ACCESS.2024.3432197.</mixed-citation>
      </ref>
            <ref id="ref2">
        <label>2</label>
        <mixed-citation>L. I. B. López, J. A. V. Paredes, and R. M. Delgado, CNN-LSTM and Post-Processing for EMG-Based Hand Gesture Recognition, Intelligent Systems with Applications, vol. 22, pp. 1–12 2024. doi: 10.1016/j.iswa.2024.200352.</mixed-citation>
      </ref>
            <ref id="ref3">
        <label>3</label>
        <mixed-citation>Y. Song, M. Liu, F. Wang, J. Zhu, A. Hu, and N. Sun, Gesture Recognition Based on CNN-BiLSTM for Wearable Wrist Sensors, IEEE Sensors Journal, vol. 24, no. 4, pp. 3561–3572, Feb. 2024, doi: 10.1109/JSEN.2023.3349214.</mixed-citation>
      </ref>
            <ref id="ref4">
        <label>4</label>
        <mixed-citation>R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization,    IEEE Transactions on Computer Vision and Pattern Recognition,  vol. 128, no. 2, pp. 336–359, 2023.</mixed-citation>
      </ref>
            <ref id="ref5">
        <label>5</label>
        <mixed-citation>S. A. Mahmoudi, M. R. Keyvanpour, and A. Ghorbani, Explainable Deep Learning for Human Action Recognition: A Survey and Experimental Analysis, vol. 11, pp. 124901–124920, 2023,  doi: 10.1109/ACCESS.2023.3324186.</mixed-citation>
      </ref>
            <ref id="ref6">
        <label>6</label>
        <mixed-citation>Haroon, Umair and Ullah, Amin and Hussain, Tanveer and Ullah, Waseem and Sajjad,</mixed-citation>
      </ref>
            <ref id="ref7">
        <label>7</label>
        <mixed-citation>Muhammad and Muhammad, Khan and Lee, Mi Young and Baik, Sung Wook, "A Multi-Stream Sequence Learning Framework for Human Interaction Recognition," in IEEE Transactions on Human-Machine Systems, vol. 52, no. 3, pp. 435-444, June 2022, doi: 10.1109/THMS.2021.3138708.</mixed-citation>
      </ref>
            <ref id="ref8">
        <label>8</label>
        <mixed-citation>Shah M, Nawaz T, Nawaz R, Rashid N, Ali MO (2025) InterAcT: A generic keypoints-</mixed-citation>
      </ref>
            <ref id="ref9">
        <label>9</label>
        <mixed-citation>based lightweight transformer model for recognition of human solo actions and interactions in aerial videos. PLoS One 20(5): e0323314. https://doi.org/10.1371/journal.pone.0323314</mixed-citation>
      </ref>
            <ref id="ref10">
        <label>10</label>
        <mixed-citation>Hoangcong Le, Cheng-Kai Lu, A low-latency deep learning approach for human action</mixed-citation>
      </ref>
            <ref id="ref11">
        <label>11</label>
        <mixed-citation>recognition in medical internet of things applications, Computers and Electrical Engineering, Volume 132, 2026, 111005, ISSN 0045-7906, https://doi.org/10.1016/j.compeleceng.2026.111005.</mixed-citation>
      </ref>
            <ref id="ref12">
        <label>12</label>
        <mixed-citation>Liu M, Li W, He B, Wang C, Qu L. Human Action Recognition Based on 3D Convolution and Multi-Attention Transformer. Applied Sciences. 2025; 15(5):2695. https://doi.org/10.3390/app15052695</mixed-citation>
      </ref>
            <ref id="ref13">
        <label>13</label>
        <mixed-citation>Jin L, Fan R, Han X and Cui X (2025) Convolutional spatio-temporal sequential</mixed-citation>
      </ref>
            <ref id="ref14">
        <label>14</label>
        <mixed-citation>inference model for human interaction behavior recognition. Front. Comput. Sci. 7:1576775. doi: 10.3389/fcomp.2025.1576775.</mixed-citation>
      </ref>
          </ref-list>
  </back>
  
</article>
