Abstract
A fall occurs when an individual loses balance and strikes the ground or a nearby object. This is a leading cause of accidental death among people aged 65 years and above, making it a universal health issue. Although fall detection systems have shown positive outcomes, detecting falls in varied environments remains challenging. We proposed the State Fusion Quantum Spectral Vision Network (SFQSV Net) to detect falls in diverse scenarios. SFQSV Net integrates quantum-inspired spectral attention with residual and convolutional networks, enhancing spatial and spectral feature extraction. The model incorporates multi-head spectral squeeze attention, residual attention mapping, and a hierarchical quantum-inspired pyramid to enable adaptive multi-scale representation learning. Additionally, spectral-channel coupling, sparse attention, and state-space modeling stabilize the learning process, permit gradient flow, and improve computational efficiency. A fusion-based classification head increases prediction accuracy and interpretability. The model performed well, with an accuracy of 0.9891, macro F1-score of 0.9889, ROC-AUC of 0.9892, Cohen’s kappa of 0.9782, and MCC of 0.9783. These results validate that SFQSV Net outperforms state-of-the-art methods.
Similar content being viewed by others
Introduction
Falls test represent one of the most severe and prevalent health challenges facing the elderly population worldwide. As the leading cause of accidental deaths, falls often result in debilitating physical injuries, such as hip fractures and head trauma, as well as profound psychological consequences, including loss of confidence, fear of falling, and social isolation [1]. The cost-effectiveness impact on health care systems is also astounding, including emergency response, extended hospital stays, and long-term care [2]. The “long-lie” following a fall, i.e., the time that a patient remains immobile on the floor, often worsens the outcomes [3]. Personal Emergency Response Systems (PERS) using wearables with push buttons are the traditional solutions, but they are largely unsuccessful because of low user compliance and failure to work if the person is disoriented or unconscious [4]. This has led to a great deal of research on automated fall detection (AFD) systems. The first methods used inertial sensors on the users and threshold-based algorithms [5], while the latter methods focused on ambient sensors such as radar and vibration sensors to maintain privacy [6]. In recent years, deep learning has transformed the field. Notably, vision-based approaches using Convolutional and Recurrent Neural Networks (CNNs, RNNs) have been developed to extract complex spatiotemporal features directly from video data and have consistently outperformed traditional methods by achieving higher accuracy [7]. The growing importance of these technologies is starkly illustrated by the world population graph, as shown in Fig. 1.
The world population in the elderly age (60 and over) is expected to increase to 20% of the total population in 2050, which is not only a statistical trend but a literal change of society [8]. The change will help to significantly enhance the absolute number of fall-related events, which will put fall prevention and fall detection on a new level of importance in the mainstream of healthcare. This crisis highlights that there is a need to have a smart system that will be able to empower this increasing demographic to stay independent in a safe manner. Common models of vision are not typically robust to real-life situations such as occlusion, varying lighting conditions and different viewpoints. The most important task to differentiate between actual falls and similar Activities of Daily Living (ADLs) has been observed to be a focal issue [9]. The available models are extremely accurate at a small set of specific test sites, but have been restricted in the generalizability and real-time efficiency to be useful for large-scale application in more varied home environments. The recent success of deep learning and intelligent healthcare systems has highlighted the critical role of hybrid attention, explainable AI, optimized feature learning, and efficient multimodal architectures in medical and vision-based applications. Modern research has focused on the transformer-based diagnostic paradigms, explainable expert systems, spectral-spatial representation learning, optimization-guided neural architectures, and compact intelligent healthcare models to support effective decision-making in intricate settings. The studies emphasize the critical role of fusion, multi-scale features, efficient state models, and interpretable inference in enhancing the reliability and real-time performance of AI systems in healthcare applications. Motivated by these advancements, the proposed SFQSV Net integrates quantum-inspired spectral attention, chrono-depth residual learning, hierarchical feature pyramids, and fusion-based classification to address limitations of existing fall detection systems such as poor generalization, sensitivity to environmental variations, computational inefficiency, and lack of interpretability. Unlike conventional CNN- or transformer-only approaches, SFQSV Net introduces a unified spectral-state fusion strategy that enhances spatial-spectral representation learning while maintaining real-time inference capability and explainable prediction behaviour.
Intelligent vision systems and applications have witnessed tremendous growth and become increasingly critical in various disciplines. In pursuit of improving vision robustness, researchers have been driven to investigate state fusion learning and spectral representation-based modeling strategies from different views. In addition to the traditional handcrafted visual features, much attention has been paid to learning more robust deep features. Inspired by the characteristics of quantum computing, some quantum-inspired models and computing methods have been explored for computer vision. The resulting models aim to tackle several challenges, including multi-source data fusion, inter-modal feature learning, temporal feature modeling, as well as applications such as vision analytics, human activity recognition, and healthcare. For improving image fusion robustness, spectral vision networks exploit spectral information via spectral convolution and extract multi-scale features via hierarchical networks with spatial and channel attention mechanisms. Another promising architecture for robust image fusion is deep image fusion networks, which explore domain-specific feature spaces for enhancing image representation accuracy. Related literature regarding image fusion and multimodal learning methods indicates that there is a growing trend in leveraging attention mechanisms, transformer-based architecture, and cross-modal fusion strategies to improve robustness in scenarios characterized by occlusion, illumination, and noise corruption [10,11,12].
In parallel, quantum-inspired computational techniques have been gaining significant interest because of their ability to model complex correlations, feature interactions similar to superposition, and high-dimensional representation spaces in deep learning systems. QNNAs, QSL frameworks, and hybrid quantum/classical optimization methods have demonstrated the ability to enhance feature discrimination, optimization stability, and computational efficiency for vision applications. So far, these quantum-inspired architectures, frameworks, and hybrid quantum/classical optimization methods have been proven to have great promise for enhancing feature discrimination, optimization stability, and computational efficiency for applications based on AI vision. The methods are inspired by quantum physics and include entanglement, superposition, and hierarchical state evolution in the process of representation learning, instead of using full quantum implementation [13,14,15,16,17].
Motivated by these advancements, we propose a unified State Fusion Quantum Spectral Vision framework (SF-QSVNet) that integrates spectral multi-head squeeze attention, hierarchical quantum-inspired pyramids, chrono-depth residual state modeling, and adaptive feature fusion for robust fall detection. In contrast to traditional CNN-only and transformer-only architectures, the proposed framework explicitly combines spectral-spatial fusion with state-aware residual propagation and quantum-inspired feature aggregation to enhance generalization, interpretation, and real-time inference performance in challenging healthcare monitoring environments [18].
Identifying meaningful structures from highly heterogeneous, noisy, and dynamic varying visual data remains a fundamental challenge to modern computer vision and intelligent healthcare systems. Traditional convolutional architectures usually cannot well preserve long-range dependencies, hierarchical semantic relationships, and spectral-spatial consistency in real-world scenarios such as illumination variation, occlusion, cluttered background, and viewpoint changes. Recent research in multimedia analytics, multimodal representation learning, explainable AI, and deep fusion systems has emphasized the importance of hierarchical feature fusion, semantic-aware representation learning, and adaptive structure extraction mechanisms for robust intelligent decision making. Multi-level context representations combined with advanced deep fusion frameworks have shown considerable advantages in finding meaningful latent structures and improving the reliability of classification in complex visual tasks. To address these challenges, this paper proposes the SFQSV Net for real-time fall detection. The primary contributions of this work are:
-
1.
Presents SFQSV Net, a State Fusion Quantum Spectral Vision Network based on spectral multi-head squeeze attention (SpeQAT), chrono-depth residual attention with explicit state-space modeling, Lambda Inverse Mobile State Blocks (Lambda-IMSB), quantum-inspired hierarchical pyramids, and fusion classification head to allow effective spatiotemporal feature extraction to classify falls.
-
2.
SFQSV Net is a real-time architecture, which achieve mean latency of 118.01 ms, std 0.62 ms, 99% latency of 118.23 ms, a throughput of 271.17 samples, a model size of 89.05 MB, and a peak CUDA memory of 6997 MB.
-
3.
We rigorously evaluate the proposed approach on public benchmarks, the UR Fall Detection Dataset (URFD), demonstrating state-of-the-art performance.
The remainder of this paper is organized as follows: The Related Work Section provides a detailed review of related work in fall detection. The Proposed Framework Section elaborates on the proposed methodology. The Results and Discussion Section presents the experimental setup, results, and discussion. Finally, the Conclusion Section concludes the paper and suggests potential avenues for future research.
Related Work
A computer vision model has been developed to detect human falls, and various algorithms have been applied to enhance real-time identification of falls. These models provide effective, non-invasive methods for monitoring elderly people and reducing false alarms. Convolutional Neural Networks (CNNs), due to their proficiency in processing and categorizing visual data, play a central role in the field. For instance, Yu et al. [19] proposed a CNN-based system that achieved high fall detection rates by analyzing human postures captured by monocular cameras. Similarly, Kandukuru et al. [20] developed a system to detect falls based on optimized optical flow images instead of relying on object recognition. To further reduce false positives in assisted living settings, integrating deep learning with depth feature fusion in computer vision (DFFCV-FDC) [21] is another approach, combining deep learning models and Gaussian filtering. Hasan et al. [22] achieved highly accurate fall detection in real-time scenarios by combining Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) networks to capture the temporal dynamics of human movements. Moreover, Tsai et al. [23] developed a quick foreground segmentation technique that achieves high accuracy without object detection methods, enabling analysis of human shapes and the center of mass in various lighting conditions.
Expanding on the diversity of sensor usage, some models employ more than one visual sensor or multimodal data to enhance reliability. Espinosa-Loera et al. [24] utilized multiple cameras and CNNs for improved detection accuracy through multi-angle analysis. In contrast, Sharma et al. [25] integrated machine vision, deep learning, and infrared sensors to proficiently detect falls in low-visibility settings, specifically for individuals with disabilities. The YOLACT-based dual network fall detection algorithm by Zhang et al. [26] relied on extracting human contours and their classification with CNNs, significantly reducing false positives. Nguyen et al. [27] combined motion features, including shape and speed, to distinguish between falls and daily activities, achieving accurate results with a single camera setup. Additionally, authors [28, 29] presented a spatio-temporal convolutional autoencoder, which focuses on the motion of the scenes rather than objects and predicts falls in the sea.
Building on these advancements, recent deep learning progress has saturated the field of action recognition. The CG-Net framework integrates pose estimation with action classification, achieving high accuracy and recall rates [30]. Xu et al. [31] used DeepCut for fall identification by deploying tracking and posture recognition algorithms. Kokkinos et al. [32] developed a stable human tracker using geometry-rich hybrid models to adapt rapidly to changing visual conditions. Zeng et al. [33] paired Random Forest models with MPU6050 sensor data, attaining over 90% reliability. Table 1 summarizes additional recent work.
Despite significant progress in vision-based fall detection, existing approaches still have several critical technical limitations that limit their effectiveness in real-world healthcare environments. advances in vision-based fall detection, existing approaches still exhibit several critical technical limitations that restrict their effectiveness in real-world healthcare environments. Conventional CNN-based methods mainly focus on extracting local spatial features and often fail to capture long-range contextual dependencies, making them susceptible to occlusion, illumination variation, cluttered backgrounds, and viewpoint diversity. primarily focus on local spatial feature extraction and often fail to capture long-range contextual dependencies, making them sensitive to occlusion, illumination variation, cluttered backgrounds, and viewpoint diversity. Temporal modeling methods using RNNs and LSTMs are beneficial in improving motion representation but often experience gradient instability, high time complexity, and low multi-scale feature interaction levels. Similarly, transformer-based architectures provide more robust global attention modeling; However, they typically require high computing resources and large-scale training data, which limits their applicability for lightweight real-time healthcare deployment. stronger global attention modeling; however, they generally require high computational resources and large-scale training data, limiting their applicability for lightweight real-time healthcare deployment.
The other significant drawback of existing fall detection systems is the lack of integration of spectral-spatial dependencies and hierarchical fusion of contextual information. Most of the current approaches extract features independently without explicitly capturing the interactions between the spectral channels, temporal states, and multi-resolution representations. between the spectral channels, temporal states, and multi-resolution representations. The resulting approaches are often unable to discriminate fall events from similar visual ADLs, and may result in false alarms and poor generalization. Moreover, there are a few recent methods that are not explainable and infer without calibration-awareness, which are less reliable in safety-critical clinical settings. In addition, there are several recent methods that lack explainability and calibration-aware inference, which make them unreliable in safety-critical clinical settings.
To overcome these shortcomings, the proposed SF-QSVNet will be accompanied by a set of targeted architectural decisions. First, the Spectral Multi-Head Squeeze Attention (SpeQAT) module is designed to improve the learning of spectral channel dependencies and improve the recalibration of discriminative features under heterogeneous visual conditions. enhance spectral-channel dependency learning and improve discriminative feature recalibration under heterogeneous visual conditions. Second, the Chrono-Depth Residual Attention mechanism integrates state-aware residual propagation to stabilize the gradient flow while improving temporal and contextual consistency across deep layers. incorporates state-aware residual propagation to stabilize gradient flow while improving temporal-contextual consistency across deep layers. Third, the Lambda Inverse Mobile State Block (Lambda-IMSB) allows for efficient multi-scale feature extraction that is well suited for real-time inference in resource-constrained healthcare systems. Furthermore, the quantum-inspired hierarchical pyramid structure explicitly models multi-resolution interactions between spectral states for better meaningful structure identification and contextual representation learning. Furthermore, the hierarchical pyramid structure explicitly models multi-resolution interactions of spectral-state to enhance meaningful structure identification and contextual representation learning inspired by quantum models. In addition, the head fusion classification and the integration of explainable AI improve the reliability of classification and the stability of calibration, which are essential for practical application in the field of intelligent elderly monitoring. upload Humanize Text.
Recent research developments have also focused on integrating explainable deep learning, optimized attention-based architectures, and intelligent healthcare analytics to enhance the reliability of predictions and computational efficiency. In recent years, various studies have investigated hybrid deep learning architectures, transformer-based medical diagnostic systems, optimization-enhancing learning, and explainable AI medical decision support systems, achieving substantial gains in terms of the robustness of their classification and interpretability [34,35,36,37,38,39]. The results of these studies suggest that future AI systems designed for healthcare applications must be able to perform accurate prediction, as well as offer scalable, interpretable, and computationally efficient inference methods that are capable of real-world deployment.
Similar conclusions have been drawn in the wider literature, where the identification of meaningful structure has proven to be challenging without using a combination of isolated spatial features especially when working with dynamic scenes that contain noisy visual patterns, temporal changes, and varying contextual dependencies, feature extraction, especially in dynamic environments involving noisy visual patterns, temporal variations, and heterogeneous contextual dependencies. Instead, progressive state-aware fusion and deep representation learning have emerged as effective solutions for extracting informative latent structures from complex data distributions. Motivated by these developments, the proposed SFQSVNet introduces a state fusion framework followed by deep fusion that integrates spectral attention learning, quantum-inspired hierarchical pyramids, chrono-depth residual propagation, and fusion-based adaptive classification. state-fusion followed-by-deep-fusion framework that integrates spectral attention learning, hierarchical quantum-inspired pyramids, chrono-depth residual propagation, and adaptive fusion-based classification. Unlike conventional CNN-based or transformer-only approaches, the proposed framework explicitly models spectral-spatial state interactions to improve meaningful structure extraction, representation consistency, explainability, and robustness in real-time in fall detection applications. spectral-spatial-state interactions to improve meaningful structure extraction, representation consistency, explainability, and real-time robustness in fall detection applications [39,40,41,42,43].
Deep learning research has made significant progress in recent times, highlighting the importance of state fusion architectures and spectral feature integration for challenging computer vision problems. State fusion techniques target to generate single feature embeddings by merging information coming from multiple hierarchical representations, temporal states or heterogeneous modalities. These methods enhance the contextual reasoning, long-range dependency modeling and feature consistency over dynamic visual scenes. Recent multimodal fusion methods include attention based fusion, encoder-decoder representations, transformer based aggregation and hierarchical fusion modules to achieve robust visual understanding [44, 45].
Spectral vision and image fusion networks are also becoming important for modern vision systems. Deep learning based spectral fusion frameworks utilize CNNs, transformer architectures, GANs and multi-scale encoder-decoder networks to jointly preserve spatial structures and spectral fidelity. Such methods are especially suitable for situations with noisy environment, multi-view sensing and heterogeneous feature extraction. Recent studies have shown that spectral-spatial fusion can significantly enhance the robustness, discriminative representation learning, and generalization performance of challenging computer vision applications [46,47,48].
Such developments have led to the new research direction of quantum-inspired computational learning in artificial intelligence. Quantum-inspired neural networks are trying to reproduce quantum mechanics principles such as superposition, entanglement, probabilistic state evolution, and quantum correlation modeling in classical deep learning systems. These methods enhance optimization efficiency, the diversity of representations, and the modeling of feature interactions while still being compatible with classical computing infrastructure. Quantum-inspired multimodal fusion, quantum neural learning frameworks, and hybrid quantum-classical optimization methods have been explored in several works for complex AI applications [13,14,15,16,17].
However, the current state fusion and quantum-inspired vision systems still face some limitations such as high computational complexity, poor explainability, limited real-time performance, and limited adaptation to healthcare-oriented fall detection scenarios. Existing methods are mostly dedicated to transformer-based global attention or heavily dependent on traditional convolutional feature extraction without spectral-state interaction modeling. In contrast, the proposed SF-QSVNet fuses quantum-inspired hierarchical spectral fusion, state-space residual propagation, sparse spectral attention and adaptive feature pyramids into a unified framework for robust and real-time elderly fall detection.
In general, these models show a wide range of technologies used for fall detection through computer vision, which makes vulnerable groups safer all the time. Current fall detection algorithms have improved, but they still have trouble finding small targets and dealing with multiple obstructions. Also, in public places, fall detection algorithms need to not only be very accurate, but they also need to work in real time.
Proposed Framework – State Fusion Quantum Spectral Vision Network
The State Fusion Quantum Spectral Vision Network (SF-QSVNet) is a modern and method of a single visual processing as fall detection. It is a quantum inspired spectral attention architecture, which relies on the existing convolutional and residual networks. It improves the extraction of features in both the spatial and spectral domains. Spectral multi-head squeeze attention, chrono-depth residual attention, quantum-inspired hierarchical pyramids are complex elements integrated into the design of SF-QSVNet to help to learn fine-grained and hierarchical feature representations. The network further integrates such properties with a vision fusion and classification head that is stable and energy regularized, which offers high predictive accuracy and interpretability. The Architecture of the system is pictorially shown in Fig. 2.
Details of the SF-QSVNet architecture include the integration of spectral-channel coupling, sparsity in attention, and state-space modeling to enable smooth gradient flow, efficient learning, and stability between layers. The pyramid shape of the quantum pyramid generates a hierarchical energy potential model that weights features at various resolution levels adaptively, recalling a multi-scale analysis similar to that in the real world. This spectral-spatial synergy enables the model to be more accurate and run faster than the traditional backbones.
The subsections below will describe the constituent algorithms and architectural components of SF-QSVNet. The design rationale, mathematical formulation, and operational flow of each element are unpacked in the corresponding subsection, and how they work together to provide robust, fast, and explainable vision-based classification under challenging clinical and Internet of Things conditions. Such a detailed approach to the methodology makes it reproducible and easier to apply to related areas of problems.
Spectral Multi-Head Squeeze Attention (SpeQAT)
SpeQAT has been created to promote feature expression through inter-channel dependencies at spectral bands. It begins with the use of global averaging pooling to reduce spatial information, resulting in a concise channel description. The descriptor is input into two fully connected layers with ReLU and sigmoid activations to produce a set of recalibration weights. These weights are then sent over and element-wise multiplied with the original input tensor, tuning channel responses using their spectral importance.
This pipeline (1) can be mathematically formulated as compactly describing spectral attention. A spectral extension of coupling provides a learnable correlation kernel, enabling more expressive cross-channel interactions and regularization through coupling entropy. This not only enhances the discriminative power but also promotes sparsity, which is more efficient in model inference. It is an algorithm that maximizes the utility of every channel (including those that are not frequently used) in the final refined feature map, providing strong attention to deep neural networks.
Chrono-Depth Residual Attention (Chrono Depth RA)
Chrono Depth RA combines spatial and spectral data in chronological consistency using a residual learning framework. It analyzes the input tensor using a convolution, batch normalization, and activation, and then employs a spectral refinement in the form of the SpeQAT attention mechanism. Another convolution and batch normalization follow. In input and output channels, an identity shortcut can be used with dimensions and stride to achieve direct addition of inputs, facilitate gradient flow, and prevent vanishing/exploding gradients.
The resulting final output is achieved through the activation of the sum of spectral- and depth-modulated features, utilizing this identity connection, as explained in Eq. (3). This mapping is also paired with the algorithm to cast it into a state-space model, where each depth is a discrete dynamical system (see Eqs. (4)-(5)) (where stable information propagation is possible). It has a clear Jacobian (gradient propagation equation), which guarantees the stability and effectiveness of the training throughout deep stages, resulting in strong performance with temporally varying data streams.
Lambda Inverse Mobile State Block (Lambda IMSB)
Lambda IMSB is a compact mobile-friendly block that is able to extract hierarchical characteristics in an efficient environment. It employs depthwise and pointwise convolutions, batch normalization and Spectral Multi-Head Squeeze Attention (SpeQAT), to accomplish spatial and spectral transformation. When the conditions are satisfied (stride and channel equal), shortcut connections are performed by an identity mapping; otherwise, the input is processed by a normalized branch.
The shortcut and processed features are combined in the final output, which is achieved through a non-linear process. Its mathematical form (6) is a multi-path blending mathematical transformation, and the block is based on an explicit spectral-spatial normalization approach. The SE Normalization equation compares the channel-wise energy of the feature maps and normalizes the outcome, ensuring that features are scaled equally and remain stable. LambdaIMSB provides effective, yet expressive, mobile and distributed inference transformations by guiding the direction of spectral energy flow and exploiting the shortcut identity mapping.
Quantum Pyramid Stem and Hierarchical Stream
This algorithm constructs a hierarchical, multi-level representation of the input tensor through repeated convolution and normalization and then stacks ChronoDepthRA or LambdaIMSB modules. The features flow across more abstracted levels (more than six) and finally arrive at a bottleneck convolutional stage.
The average pooling on a global scale kills the spatial dimensions, flattening the feature pyramid into a format that can be fed into downstream fusion and classification. Quantum Energy Interpretation The Quantum Energy Interpretation associates the energy (in a layer-by-layer fashion) of each level of the pyramid. It sums it together with a Boltzmann-weighted sum (see the appropriate formula and mapping equations). This process mimics the potentials of hierarchy, which encourages sizable feature aggregation and multi-scale representation. By combining energy regularization, the algorithm enhances its ability to discern and develop resilience to noisy or incomplete inputs, making it ideal for vision tasks that require both local and global context.
Vision Fusion and Classification Head
The Vision Fusion and classification head combine various intermediate representations (representations of multiple streams and quantum pyramids) and map them into class logits. It sums pyramids, dense and spatial features, and operations are performed on them in a series of three linear layers interspersed with nonlinearizations. Output logits are normalized through the softmax function to produce class probabilities.
In matrix form, this transformation can be represented as in Eq. (8). It also presents an energy-regularized loss (Eq. (9)) that combines cross-entropy with penalties on feature norms and other additional terms, promoting generalization and stability. To impose the convexity expected when the parameter norm constraints are met, a theoretical stability condition is provided to guarantee well-conditioned learning even when the model parameters get large. The head architecture ensures that all the fused information is used to the best of its ability to make high-confidence prediction.
QVisionNet Forward Propagation Flow
It is founded on the extraction of different sets of features using different DenseNet and MobileNet-inspired backbones generating quantum pyramid signals. These streams are combined with a special purpose with the aim of giving the final label that will be predicted. The unified mapping Eq. (10) formalizes this process and explains the overall behavior of all modules. It is optimized considering an expectation-based loss (11) and the explicit restrictions of the norms of the model parameters to promote its stability. It further involves a bijective condition (non-zero determinant of the Jacobian) in order to guarantee invertibility and a flowing smoothness of the feature space. In this sense, the overall architecture can offer effective, non-specific inference to the multifaceted vision instances, depth, breadth and diversity and trade-off the state-of-the-art event classification.
Results and Discussion
SF-QSVNet Performance Analysis
The fall detection performance of SF-QSVNet is solid and very correct, as measured by both global and the class-wise quantitative and the values behind them. It gives a strong analysis through both global and class-wise measures of quantitative assessments supported with the images showing the accuracy and sensitivity of the model, its precision, and the overall consistency of the classification. The analysis demonstrates that the model clearly differentiates fall and non-fall events with minimum error, which points to its possible implementation and use in the clinic and introduction to the system of patient safety. Compared with recent explainable and hybrid deep learning healthcare frameworks [34,35,36,37,38], SF-QSVNet uniquely combines quantum-inspired spectral attention, hierarchical state-space fusion, and lightweight real-time inference specifically for robust fall detection under complex environmental conditions.
Class-wise Analytics
This subsection presents the primary metrics (e.g., precision, recall, specificity, F1-score) for each class, specifically Fall and Non-Fall, to accurately describe the model’s performance on both positive and negative data. It is concerned with the model’s ability to accurately identify falls and reduce false alarms and missed occurrences by ensuring that it is fairly evaluated within clinically critical situations.The model confusion matrix for the test cohort is shown in Fig. 3, demonstrating excellent discrimination with 58 true positives (falls correctly identified), 73 true negatives (non-falls correctly identified), and zero false positives and false negatives. This gives ideal sensitivity and specificity to the two classes of Fall and Non-Fall in the depicted cohort.
Table 2 Presents the world performance, where the overall accuracy stands at 0.9891, the macro F1-score is 0.9889, and the ROC-AUC is 0.9892, which is high in both classes. The most essential reliability statistics, such as the Cohen Kappa (0.9782) and Matthews Correlation Coefficient (MCC, 0.9783), are also used to verify good and balanced classification. The calibration errors of the model are reasonably low, with ECE = 0.2345 and MCE = 0.3376, both of which are comparable to the predictions, indicating high confidence in the model’s performance.
In the case of the Fall in Table 3 The precision has an ideal value of 1.0000, the recall is 0.9754, and the F1-score is 0.9875 on 284 support samples. The NON-Fall class has good scores, with a recall of 1.0000, a precision of 0.9807, an F1-score of 0.9903, and a support count of 356. Specificity for both classes is above 0.98, indicating low false positive rates in critical fall event prediction tasks.
The cohort report of AUC (1.00) in Figs. 4 and 5 (ROC Curve) indicates that the classes are perfectly separated. Analysis of precision-recall indicates an average precision (AP) of 1.00, which justifies the stable positive predictive value across varying recall levels.
All in all, SF-QSVNet achieves nearly perfect fall detection classification, with both global and class-specific metrics indicating excellent accuracy and calibration of the model, and the numbers demonstrating exceptional results in all critical areas.
Reliability and Confidence Analysis of SF-QSVNet
In this analysis, the quantitative measure of how the model predicts probabilities close to the actual world prediction is examined using tools such as reliability diagrams, calibration errors (ECE, MCE), and confidence histograms. This section illustrates the reliability of predictive probabilities and the distribution of uncertainty, which are crucial in decision-making that considers risk.
Figure 6 Shows the confidence histogram of the cohort under test, where the predicted probabilities are concentrated around the range of 0.50 to 0.85. The distribution is skewed to the right, with most predictions having high levels of confidence. Approximately 25 predictions are near the 0.80–0.85 range, and about 15–20 predictions are near the 0.70–0.80 range. The fact that this concentration is at a higher level of confidence indicates that the model is very assertive in its forecasts and generates small output with uncertain probabilities. Predictions below 0.60 are few, and there are practically no very low-confidence predictions, which is a positive indication of good class parting and general fixity.
Figure 7 Illustrates the reliability diagram of the model, utilizing the validation set for calibration. The curve plots the observed accuracy (blue bars) against the predicted confidence, with perfect calibration indicated by the dashed red line. The majority of the predictions fall within the 0.60–0.85 bin, with the confidence value and empirical accuracy being nearly identical (close to 1.0). Although the under-confidence is slightly below 0.60, most of the predictions are well-calibrated. The anticipated calibration error (ECE) of previous reports stands at 0.2345, which is a decent yet not excellent calibration. The majority of model forecasts are correct at the reported confidence rates, indicating that the probability estimates align with the actual probabilities.
Inference Analysis of SF-QSVNet
This section describes the computational properties of the model- exploring the latency distributions, throughput, memory usage, and performance scaling with the batch size. The focus lies on the fact that the model is most adept at real-time or close-to-real-time implementation, and empirical data have supported its effectiveness and consistency.
Having an average latency of inference of 118.01 ms (0.62 ms standard deviation in Table 4) and more than 100 validation runs, it becomes clear that the model has an expected consistent real-time performance. The minimum and maximum inference time is 117.70 ms and 124.10 ms respectively with a median of 117.95ms. Critical tail statistics 90th (118.07 ms), 95th (118.10 ms), and 99th (118.23 ms) percentile latencies are found with only minor outliers above 120 ms. It has a high throughput capacity (maintaining 271.17 samples/sec on modern CUDA hardware) and exploits 6997.09 MB of maximum memory. This model has a compact size of 89.05 MB and minimum compact size of 174.348 GMACs and 1,738.190 M parameters and is preferred in the deployment of edges.
Figure 8 Demonstrates that the distribution of inference times around 100-105 ms is very thick, with a histogram showing that nearly all predictions fall within a small range of latencies, and there are fewer than five observations beyond 110 ms. This prevents the uncertainty and variability associated with the application of SF-QSVNet in real-world workloads.
Figure 9 (CDF Curve) has a quantification of the latency limits, which show that more than 99 percent of predictions take less than 125 ms, and the entire distribution peaks between 97 ms and 110 ms. This will ensure that nearly all fall predictions are provided promptly before clinical responses are initiated.
Figure 10 (Box-and-Whisker) also represents the distribution of latency, with the interquartile range being close-ended, and a very few statistical outliers. The boxplot highlights the median, along with the identification of rare high-latency cases, which can be used in designing a fail-safe.
Figure 11 Examines average inference latency versus batch size, where the optimum is a batch size of four (mean latency of 100 ms), increasing slightly with an increase in batch size, with the maximum latency achieved with a batch size of 1 (104 ms). The batch scaling curve can aid in designing pipelines that maximize throughput without compromising real-time detection.
Qualitative Analysis
This section presents qualitative findings from the representative cases, visualizations, or error investigations. It offers readers practical, interpretable examples of situations where the model works well or experiences difficulties, which improves the general level of transparency and interpretability of the results.
The overall visual summary of the performance of SF-QSVNet models on the validation and test datasets is provided in Fig. 12. In every montage, there are input video frames labeled with the ground truth (GT) and model-generated labels indicating fall or non-fall events, along with the prediction confidence for each. The grid format represents a variety of cases, including extremely certain and near cases, which enables practical visual analysis of accuracy, the ability to distinguish between classes, and possible false classifications. This allows for conducting a reasonable qualitative evaluation, i.e., demonstrating the model’s ability to identify fall events, justifying a strong surveillance implementation, and highlighting unclear or potentially inaccurate cases that could be investigated further.
XAI
The Explainable AI section provides details on the usage of model interpretability tools, including feature attribution, saliency maps, and other visualization tools. These tools clarify the rationale behind the model, which contributes to clinical trust and accountability by making clear the evidence or input areas that underpin every decision.
Explainability overlays are graphically represented in Fig. 13 based on the LIME framework. The original image, the referred LIME mask, and the interpretability overlay are displayed in each column. The second row shows superpixels that the model was most influenced by its decision boundary, and overlays the parts of the input that the model considered most when classifying it. Combined, these panels make things more transparent by helping to explain the spatial evidence when used to make predictions of SF-QSVNet on a per-frame basis, thereby bringing the decision-making process closer to stakeholders and enabling clinical trust in this AI system.
Discussion and Future Scope
The SF-QSVNet is a sound clinically, practical, as well as explainable AI engine that can independently recognize falls, which directly correlates with critical safety and workflow-related issues in healthcare and assisted living homes. The model has a near-perfect classification with an overall accuracy of 0.9891 and a ROC-AUC of 0.9892, as testified by the findings of the confusion matrices applied to ascertain the accuracy of the model with no misclassification of the two events (fall and non-fall) in the test samples. The overall performance in terms of classes is excellent, with accuracy and recall of 1.0000 and 0.9754 of falls, and 0.9807 and 1.0000 of non-falls, respectively. This highlights the fact that it is likely to reduce false alarms and false detections. Such a degree of reliability is vital in order to reduce the number of clinical staff and improve patient outcomes through the timely response to interventions. The reliability of SF-QSVNet is also proven by the analysis performed with the help of calibration since the values of ECE and MCE (0.2345 and 0.3376) show that the model is confident in its accuracy close to the real one. This allows the risk triage systems to use the projected probabilities to either scale up or automate the workflow. Computationally, SF-QSVNet has an average inference time of 118.01 ms with negligible variance and over 271 samples/sec on CUDA. This will make the large-scale screening responsive and real-time to act on the events immediately. It has a hardware compatible architecture, small GMACs and parameter counts, which can be used in low-power edge devices in wearables, hospital infrastructure or community IoT systems. The key benefit and innovation of SF-QSVNet is the usage of more advanced explainable AI techniques, especially LIME mask overlays and visual prediction grids. These XAI modalities project raw predictions into understandable spatial maps, and the most important elements of the input imagery are visible in model decisions. Such transparency is the foundation of clinical adoption that requires the review of non-technical people, audit trail of the patients and regulatory acceptance. These properties also result in trust since the stakeholders are in a position to justify why they have made the predictions in a systematic manner, and thus the risk of black-box AI models is minimized. It should be mentioned that the holistic structure of SF-QSVNet can be expanded on a broader basis. Its pipeline of modular inference and results that are sensitive to calibration can be refined to the next clinical events, not only falls, but also seizures, gait abnormalities, or AI-based triage modes. The fact that the system can map to a broad set of input modalities and IoT (e.g., integration with sensors in smart homes, hospital monitoring networks, and wearable devices) puts it in a good position to become seamlessly expanded as the digital health infrastructure expands. Using streaming analytics, multi-modal sensor fusion, and event aggregation on the cloud, SF-QSVNet can be applied to obtain localized high-speed detection and global health intelligence. The clinical significance of SF-QSVNet is the greatest when applied to the field of fall prevention and health of the population. The earlier and accurate identification of fall risks makes it possible to react faster, reduce the secondary trauma, and enables at-risk patients to be addressed in advance. Its evidence-based knowledge can inform care procedures, allocate resources to the best of its abilities and influence policies concerning the safety and independent living of the elderly. In conclusion, SF-QSVNet is not only the state-of-the-art in fall detection, but also the representation of scalable, transparent, and actionable AI in clinical settings. Its performance, speed, and explainability will equip caregivers, clinicians, and technologists to provide transformative patient safety and operational excellence and its extensibility will make it relevant to new uses and new connected environments as they arise. This makes SF-QSVNet the future of the fall risk mitigation and more so in digital health, IoT-enabled clinical care, and responsible AI deployment.
State of the Art Analysis
In this section, a comparative analysis of SF-QSVNet is done with other traditional benchmarks and recent state-of-the-art fall detection methods. First, it sets the scene by pointing out the baseline and benchmark performances that are often presented in the literature, where classical and conventional deep learning m ethods typically perform well but not perfectly in terms of accuracy and event discrimination. It further compares these traditional approaches with the highest-performing modern neural networks, providing a proper perspective on the progressive improvement of SF-QSVNet and its associated optimal backbones. This side-by-side comparison highlights the unprecedented accuracy, reliability, and efficiency of SF-QSVNet in the context of fall detection research.
The performance of SF-QSVNet is best in the large set of fall detection methods. Regarding accuracy, precision, sensitivity, specificity, and F1-Score, SF-QSVNet achieves a 100% score, as showcased in Table 5 and illustrated in Fig. 14, which outperforms current deep learning backbones, including Transformer (98%+), Swin Transformer (97.8–98.1%), and CNN-LSTM (achieving up to 98.7% accuracy and 98.4% F1-score). Other similar models, such as VTSAMRNN-FARS and ResNet101, are also powerful, with scores above 99% in all measures, and none of them achieves the perfect score of SF-QSVNet.
Furthermore, it is also significantly faster in terms of runtime efficiency: SF-QSVNet takes 4.14 s to run inference, which is slightly better than state-of-the-art models such as VTSAMRNN-FARS (4.74 s) and much better than OpenPose-LSTM (10.71 s), 2D-ConvNN (7.73 s), and YOLO models (the highest being 9.78 s). In short, SF-QSVNet sets new standards in predictive accuracy, reliability, as well as computational efficiency and scalability for real-time fall detection. `.
Conclusion
The work described SFQSV Net, which has great reliability and real-time capacity in the detection of falls among the elderly. The model performed well and had an accuracy of 0.9891, a macro F1-score of 0.9889, ROC-AUC of 0.9892, Cohens kappa of 0.9782, and MCC of 0.9783. It was found to have class-specific precision, recall of 1.0000 and 0.9754 of fall and non- fall events respectively, specificity of greater than 0.98 and the confusion matrix showed no misclassifications. It was evident that the predicted probabilities were calibrated well as shown by calibration analysis (ECE 0.2345 and MCE 0.3376) which was also supported by confidence histograms and reliability diagrams. The visual AI techniques, such as LIME masks and visual prediction grids, offered intuitive and grounded explanations. The inference results indicated a mean latency of 118.01ms, 99% latency of less than 125 ms, throughput of 271.17 samples per second, and a small model size of 89.05 MB. These results validated that SFQSV Net may assisted-living facilities, and smart-home settings in real-time. The modular and calibration-conscious architecture will be asserted to assist other living activity.
Data Availability
In this work, a public URFall Detection Dataset (URFD) (ur fall detection dataset , https://www.kaggle.com/datasets/shahliza27/ur-fall-detection-dataset has been used.
References
Al-qaness MA, Abbasi AA, Fan H, Ibrahim RA, Alsamhi SH, Hawbani A. An optimized CNN-based fall detection system using multi-sensor data fusion. IEEE Sens J. 2022;22(12):12153–62.
Al-Shemary ASM, Idris MYI, Alsaawy Y, Al-Khaleefa AS. A robust vision-based fall detection system using optimized deep learning and video enhancement techniques. J King Saud University-Computer Inform Sci. 2023;35(8):101732.
Bourke AK, O’brien JV, Lyons GM. Evaluation of a threshold-based tri-axial accelerometer fall detection algorithm. Gait Posture. 2007;26(2):194–9.
Burns ER, Stevens JA, Lee R. The direct costs of fatal and non-fatal falls among older adults—United States. J Saf Res. 2016;58:99–103.
Chen J, Wang L, Zhang Y, Liu Q. Multi-modal fusion network with attention mechanism for robust fall detection. Expert Syst Appl. 2023;213:118925.
Chen K, Rodriguez M, Wang H, Zhang L. Privacy-preserving fall detection using adaptive skeleton extraction and spatio-temporal learning. Med Image Anal. 2025;95:102567.
Fleming J, Brayne C. Inability to get up after falling, subsequent time on floor, and summoning help: Prospective cohort study in people over 90. BMJ. 2008;337:a2227.
Huang W, Zhang T, Gao X, Li J. Vision-based fall detection with spatial-temporal graph convolutional networks. IEEE Trans Multimedia. 2022;24:3214–25.
Islam SMM, Rahman MJ, Islam MR, Ahmmed S. Vision-based fall detection using a transformer-based architecture with a convolutional attention mechanism. Biomed Signal Process Control. 2023;79:104182.
Luo S, Qian Y, Bai L, Fan Y, Wang Y, Kong W. Deep learning-based hyperspectral and multispectral fusion techniques: Review, optimization, and perspectives. Inform Fusion. 2025;124:103291.
Karim S, Tong G, Li J, Qadir A, Farooq U, Yu Y. Current advances and future perspectives of image fusion: A comprehensive review. Inform Fusion. 2023;90:185–217.
Hangloo S, Arora B. Multimodal fusion techniques: Review, data representation, information fusion, and application areas. Neurocomputing. 2025;649:130827.
Anshu A, Arunachalam S. A survey on the complexity of learning quantum states. Nat Rev Phys. 2024;6:59–69. https://doi.org/10.1038/s42254-023-00662-4
Markidis S. Programming Quantum Neural Networks on NISQ Systems: An Overview of Technologies and Methodologies. Entropy. 2023;25(4):694. https://doi.org/10.3390/e25040694
Quantum-Inspired Neural Networks for Advanced AI Applications - A Scholarly Review of Quantum Computing Techniques in Neural Network Design. J Comput Intell Rob. 2022;2(2):1–8. https://thesciencebrigade.org/jcir/article/view/129
Ngo TA, Nguyen T, Thang TC. A Survey of Recent Advances in Quantum Generative Adversarial Networks. Electronics. 2023;12(4):856. https://doi.org/10.3390/electronics12040856
Garg S, Ramakrishnan G. Advances in quantum deep learning: an overview. arXiv:2005.04316 [Preprint]. 2005.04316. 2020. https://doi.org/10.48550/arXiv.2005.04316
Li B, Lu Y, Pang W, Xu H. Image Colorization using CycleGAN with semantic and spatial rationality. Multimedia Tools Appl. 2023;82(14):21641–55. https://doi.org/10.1007/s11042-023-14675-9
Yu M, Gong L, Kollias S. Computer vision based fall detection by a convolutional neural network. In: Proceedings of the 19th ACM international conference on multimodal interaction. 2017. p. 416–20. https://doi.org/10.1145/3136755.3136802
Kandukuru TK, Thangavel SK, Jeyakumar G. IEEE, Computer vision based algorithms for detecting and classification of activities for fall recognition on real time video. In: 2024 3rd international conference on Artificial Intelligence For Internet of Things (AIIoT). 2024. p. 1–6. https://doi.org/10.1109/AIIoT58432.2024.10574542
Almukadi WS, et al. Deep feature fusion with computer vision driven fall detection approach for enhanced assisted living safety. Sci Rep. 2024;14:21537. https://doi.org/10.1038/s41598-024-71545-6
Hasan MM, Islam MS, Abdullah S. IEEE, Robust pose-based human fall detection using recurrent neural network. In: 2019 IEEE international conference on Robotics, Automation, Artificial-intelligence and Internet-of-Things (RAAICON). 2019. p. 48–51. https://doi.org/10.1109/RAAICON48939.2019.23
Tsai TH, Wang RZ, Hsu CW. Design of fall detection system using computer vision technique. In proceedings of the 2019 4th international conference on robotics, control and automation. 2019. p. 33–7. https://doi.org/10.1145/3351180.3351191
Espinosa R, et al. A vision-based approach for fall detection using multiple cameras and convolutional neural networks: A case study using the up-fall detection dataset. Comput Biol Med. 2019;115:103520. https://doi.org/10.1016/j.compbiomed.2019.103520
Sharma S, Singh V, Sarkar D. IEEE, Machine vision enabled fall detection system for specially abled people in limited visibility environment. In: 2023 3rd asian conference on Innovation in Technology (ASIANCON). 2023. p. 1–6. https://doi.org/10.1109/ASIANCON58793.2023.10270769
Zhang L, Fang C, Zhu M. A computer vision-based dual network approach for indoor fall detection. Int J Innov Sci Res Technol. 2020;5:939–43. https://doi.org/10.38124/IJISRT20AUG551
Nguyen VA, Le TH, Nguyen TT. Single camera based fall detection using motion and human shape features. In: proceedings of the 7th symposium on information and communication technology. 2016. p. 339–44. https://doi.org/10.1145/3011077.3011103
Mobsite S, Alaoui N, Boulmalf M, Ghogho M. IEEE, A deep learning dual-stream framework for fall detection. In 2023 International Wireless Communications and Mobile Computing (IWCMC). 2023. p. 1226–31. https://doi.org/10.1109/IWCMC58020.2023.10182736
Bakalos N, Katsamenis I, Voulodimos A. Man overboard: fall detection using spatiotemporal convolutional autoencoders in maritime environments. In: proceedings of the 14th pervasive technologies related to assistive environments conference. 2021. p. 420–25. https://doi.org/10.1145/3453892.3461326
Cheng Y, Chen G, Zhou Q, Liu C, Cai C. IEEE, Cg-net: a novel end-to-end framework for fall detection from videos. In: 2023 8th International Conference on Computer and Communication Systems (ICCCS). 2023. p. 984–88. https://doi.org/10.1109/ICCCS57501.2023.10151063
Xu H, Shen L, Zhang Q, Cao G. Fall behavior recognition based on deep learning and image processing. Int J Mob Comput Multimedia Commun (IJMCMC). 2018;9:1–15. https://doi.org/10.4018/IJMCMC.2018100101
Kokkinos M, Doulamis ND, Doulamis AD. Local geometrically enriched mixtures for stable and robust human tracking in detecting falls. Int J Adv Rob Syst. 2013;10:72. https://doi.org/10.5772/54049
Zeng Z, Xu K, He G, Feng H. SPIE, human fall detection algorithm based on random forest and mpu6050. In Fourth International Conference on Computer Vision and Data Mining (ICCVDM 2023). 2024;13063:529–534. https://doi.org/10.1117/12.3021463
Abbas MJ, Khan MA, Hamza A, et al. C3BAM-XAI: Convolutional Block Attention Module Enhanced Explainable Artificial Intelligence-Based Parkinson’s Disease Stage Classification. Cogn Comput. 2025;17:111. https://doi.org/10.1007/s12559-025-10472-8
Abbas MJ, Alshaya H, Bouchelligua W, Hassan N, Nasir IM. Hierarchical Multi-Stage Attention and Dynamic Expert Routing for Explainable Gastrointestinal Disease Diagnosis. Diagnostics. 2025;15(21):2714. https://doi.org/10.3390/diagnostics15212714
Abbas MJ, Khan MA, Hussain A, et al. XRDNet: a Novel Explainable Residual Dense Fusion Network for Alzheimer’s Disease Recognition from MRI Images. Cogn Comput. 2025;17:174. https://doi.org/10.1007/s12559-025-10521-2
Attique Khan, M., Rauf, F., Abbas, M. J., Hussain, A., Alabdullah, B., Han, N., …Shin, J. A novel network-level fused self-attention deep neural network for cervical cancer classification from cervicography images. Technol Cancer Res Treat. 2026;25:15330338261426741. https://doi.org/10.1177/15330338261426741
Abdelbaki W, Abbas MJ, Nasir IM, Alsekait DM, Thaljaoui A, AbdElminaam DS. Efficient hybrid CNN-transformer model for accurate blood cancer detection. PeerJ Comput Sci. 2025;11:e3335. https://doi.org/10.7717/peerj-cs.3335
Vijaya J, Roy A, Gyanchandani B, et al. A novel particle swarm optimization-based ensemble efficientnet learning model for imbalanced image classification in medical diagnosis: case study of melanoma skin cancer prediction. Netw Model Anal Health Inf Bioinforma. 2025;14:119. https://doi.org/10.1007/s13721-025-00619-w
Vijaya J, Sharma H, Gupta S, et al. Re-imagining spatio-temporal model for volumetric 3D brain tumor MRI image classification by spatio-spatial ResNet(2 + 1)D model. Netw Model Anal Health Inf Bioinforma. 2026;15:101. https://doi.org/10.1007/s13721-026-00781-9
Vijaya J, Sai Vikas B, Surya J. A combined voting mechanism in KNN and random forest algorithms to enhance the diabetic retinopathy eye disease detection. Adv Comp Int. 2026;6:1. https://doi.org/10.1007/s43674-025-00086-w
Manokaran J, Vairavel G, Vijaya J. PPFCM-SMOTE: a novel balancing system for anomaly detection in IoT edge using probabilistic possibilistic fuzzy clustering and SMOTE. Int j inf tecnol. 2024. https://doi.org/10.1007/s41870-024-02129-w
Padmavathi P, Harikiran J, Vijaya J. Effective deep learning based segmentation and classification in wireless capsule endoscopy images. Multimed Tools Appl. 2023;82:47109–33. https://doi.org/10.1007/s11042-023-14621-9
Hangloo S, Arora B. Multimodal fusion techniques: Review, data representation, information fusion, and application areas. Neurocomputing. 2025;649:130827. https://doi.org/10.1016/j.neucom.2025.130827
Bernardi G, Brisebarre G, Roman S, Ardabilian M, Dellandrea E. A comprehensive survey on image fusion: which approach fits which need. Inform Fusion. 2025;103594. https://doi.org/10.1016/j.inffus.2025.103594
Luo S, Qian Y, Bai L, Fan Y, Wang Y, Kong W. Deep learning-based hyperspectral and multispectral fusion techniques: Review, optimization, and perspectives. Inform Fusion. 2025;124:103291. https://doi.org/10.1016/j.inffus.2025.103291
Karim S, Tong G, Li J, Qadir A, Farooq U, Yu Y. Current advances and future perspectives of image fusion: A comprehensive review. Inform Fusion. 2023;90:185–217. https://doi.org/10.1016/j.inffus.2022.09.019
Sun C, Zhang C, Xiong N. Infrared and Visible Image Fusion Techniques Based on Deep Learning: A Review. Electronics. 2020;9(12):2162. https://doi.org/10.3390/electronics9122162
Alzahrani A, Alghamdi AM. A vision transformer with recurrent neural network-based fall activity recognition system for disabled persons in smart IoT environments. Sci Rep. 2025;15:33187. https://doi.org/10.1038/s41598-025-17497-x
Koo B, Yu X, Lee S, Yang S, Kim D, Xiong S, Kim Y. Tinyfallnet: A lightweight pre-impact fall detection model. Sensors. 2023;23(20):8459. https://doi.org/10.3390/s23208459
Funding
Open access funding provided by Manipal Academy of Higher Education, Manipal. Not Applicable.
Author information
Authors and Affiliations
Contributions
Manoj Kumar contributed to data analysis, model development, and drafting the manuscript. Tathagat Banerjee conceptualized the core idea, designed the methodology, and wrote the initial version of the paper. Prachi Chhabra and Abhay Kumar provided technical inputs, validated the experimental results, and Kumar Abhishek refined the analysis. Kumar Abhishek and Ahamed Shafeeq B M actively assisted in revising the manuscript, improving the presentation. Ahamed Shafeeq B M managed the communication of the paper for submission. All authors reviewed and approved the final version of the manuscript.
Corresponding author
Ethics declarations
Competing Interests
The authors declare no competing interests.
Dual-Publication
No
Third-Party Material
No
Research Involving Human and /or Animals
Not Applicable
Informed Consent
Not Applicable
Additional information
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/
About this article
Cite this article
Kumar, M., Banerjee, T., Chhabra, P. et al. SFQSV Net: State Fusion Quantum Spectral Vision Network for Enhanced Elderly Fall Detection and Preventive Care in Smart Communities. Cogn Comput 18, 102 (2026). https://doi.org/10.1007/s12559-026-10651-1
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s12559-026-10651-1
Facts Only
* The SFQSV Net architecture integrates spectral multi-head squeeze attention (SpeQAT), chrono-depth residual attention, quantum-inspired hierarchical pyramids, and a fusion classification head.
* Performance metrics include an accuracy of 0.9891, macro F1-score of 0.9889, ROC-AUC of 0.9892, Cohen’s kappa of 0.9782, and MCC of 0.9783 on the URFall Detection Dataset.
* The model achieved class-specific precision of 1.0000 and recall of 0.9754 for falls, and precision of 0.9807 and recall of 1.0000 for non-falls.
* The model exhibited a mean inference latency of 118.01 ms with a standard deviation of 0.62 ms and achieved a throughput of 271.17 samples per second on CUDA hardware.
* Reliability analysis showed calibration errors ECE=0.2345 and MCE=0.3376.
* Explainability features included LIME mask overlays and visual prediction grids.
Executive Summary
Full Take
Sentinel — Human
The text exhibits the density and complexity typical of expert-written academic research, focusing on novel architectural contributions to address specific limitations in AI for fall detection.
