Abstract
Building energy forecasting plays a crucial role in improving energy efficiency and sustainability, yet large-scale deployment faces critical challenges from privacy regulations, data heterogeneity, and communication constraints. Traditional Federated Learning (FL) approaches struggle with two fundamental limitations: inability to efficiently process extended temporal sequences and prohibitive communication overhead on resource-constrained edge devices. This paper introduces SWIFT-KD, a federated learning framework that addresses these challenges through hierarchical sliding window transformers and federated knowledge distillation. The hierarchical architecture decomposes long energy sequences into overlapping segments, capturing both fine-grained local patterns and long-term dependencies while maintaining computational efficiency for edge deployment. Knowledge distillation (KD) compresses model updates by transmitting soft predictions instead of full weights, achieving a 300-fold communication reduction. Evaluation on the ASHRAE dataset with 100 heterogeneous buildings demonstrates superior performance (Coefficient of Determination (R\(^{2}\)) = 0.9708, Root Mean Squared Error (RMSE) = 92.24 kWh, Mean Absolute Error (MAE) = 44.58 kWh), outperforming the standard federated averaging by 21% and exceeding the centralized training by 13.3%. The framework reaches peak performance within two communication rounds and maintains stable accuracy throughout the remaining rounds of training. These results establish that privacy-preserving distributed learning can achieve the accuracy requirements of practical building management without sacrificing efficiency.
Data Availability
The dataset used in this study is the ASHRAE Great Energy Predictor III dataset, which is publicly available and widely used for building energy forecasting research. The dataset can be accessed from the following repository: https://www.kaggle.com/ competitions/ashrae-energy-prediction/overview. The processed data and scripts used for data preprocessing and model implementation are available from the corresponding author upon reasonable request. The source code, hyperparameter configuration files, random seed settings, and the exact data partitions used to produce the results in this manuscript are available at https://github.com/Git-Jess-Hub/swift-kd
References
Belfeki, Z., Krichen, M., & Zidi, S. (2026). A systematic survey on clustering in federated learning. Multimedia Tools and Applications,85, 429. https://doi.org/10.1007/s11042-026-21541-x
Chen, S., Long, G., Jiang, J., & Zhang, C. (2025). Federated foundation models on heterogeneous time series. Proceedings of the AAAI Conference on Artificial Intelligence,39, 15839–15847.
Doriguzzi-Corin, R., & Siracusa, D. (2024). Flad: Adaptive federated learning for ddos attack detection. Computers & Security,137, Article 103597. https://doi.org/10.1016/j.cose.2023.103597
Faheem, M., Al-Khasawneh, M. A., Khan, A. A., & Madni, S. H. H. (2024). Cyberattack patterns in blockchain-based communication networks for distributed renewable energy systems: A study on big datasets. Data in Brief,53, Article 110212.
Fekri, M. N., Grolinger, K., & Mir, S. (2022). Distributed load forecasting using smart meter data: Federated learning with recurrent neural networks. International Journal of Electrical Power & Energy Systems,137, Article 107669. https://doi.org/10.1016/j.ijepes.2021.107669
General Data Protection Regulation (GDPR). https://gdpr.eu/ Accessed: 30 September 2022 (2022)
Gholizadeh, N., & Musilek, P. (2022). Federated learning with hyperparameter-based clustering for electrical load forecasting. Internet of Things,17, Article 100470. https://doi.org/10.1016/j.iot.2021.100470
Hamdi, A., Noura, H. N., & Azar, J. (2025). A multi-teacher knowledge distillation framework with aggregation techniques for lightweight deep models. Applied System Innovation,8(5), 146. https://doi.org/10.3390/asi8050146
Harb, H., & Makhoul, A. (2019). Energy-efficient scheduling strategies for minimizing big data collection in cluster-based sensor networks. Peer-to-Peer Networking and Applications,12(3), 620–634.
Harb, H., Makhoul, A., Jaber, A., & Tawbi, S. (2019). Energy efficient data collection in periodic sensor networks using spatio-temporal node correlation. International Journal of Sensor Networks,29(1), 1–15.
Howard, A., Balbach, C., Miller, C., Haberl, J., Gowri, K., & Dane, S. (2019). Ashrae - great energy predictor iii. https://www.kaggle.com/competitions/ashrae-energy-prediction/overview
IEA: Buildings. https://www.iea.org/reports/buildings Paris. Accessed: 30 September 2022 (2022)
Jithish, J., Alangot, B., Mahalingam, N., & Yeo, K. S. (2023). Distributed anomaly detection in smart grids: A federated learning-based approach. IEEE Access,11, 7157–7179. https://doi.org/10.1109/ACCESS.2023.3237554
Kawoosa, A. I., Prashar, D., Faheem, M., Jha, N., & Khan, A. A. (2023). Using machine learning ensemble method for detection of energy theft in smart meters. IET Generation, Transmission & Distribution,17(21), 4794–4809.
Khan, A. A., Driss, M., Boulila, W., Sampedro, G. A., Abbas, S., & Wechtaisong, C. (2023). Privacy preserved and decentralized smartphone recommendation system. IEEE Transactions on Consumer Electronics,70(1), 4617–4624.
Kim, J., Kim, H., Kim, H., Lee, D., & Yoon, S. (2025). A comprehensive survey of deep learning for time series forecasting: Architectural diversity and open challenges. Artificial Intelligence Review,58(7), 1–95. https://doi.org/10.48550/arXiv.2411.05793
Li, Y., Hu, F., Ryan, M., Wang, R., & Liu, Y. (2022). Knowledge distillation for energy consumption prediction in additive manufacturing. IFAC-PapersOnLine,55(2), 390–395. https://doi.org/10.1016/j.ifacol.2022.04.225
Li, Y., Mamouei, M., Salimi-Khorshidi, G., Rao, S., Hassaine, A., Canoy, D., Lukasiewicz, T., & Rahimi, K. (2022). Hi-behrt: Hierarchical transformer-based model for accurate prediction of clinical events using multimodal longitudinal electronic health records. IEEE journal of biomedical and health informatics,27(2), 1106–1117. https://doi.org/10.1109/JBHI.2022.3224727
Li, Z., Yao, W., Luo, J., & Huang, Z. (2025). Flow-based iot intrusion detection via improved generative federated distillation learning. IEEE Internet of Things Journal. https://doi.org/10.1109/JIOT.2025.3526874
Liu, Y., Zhang, L., Ge, N., & Li, G. (2020). A systematic literature review on federated learning: From a model quality perspective. arXiv preprint https://doi.org/10.48550/arXiv.2012.01973arXiv:2012.01973
Makhoul, A., Laiymani, D., Harb, H., & Bahi, J. M. (2015). An adaptive scheme for data collection and aggregation in periodic sensor networks. International journal of sensor networks,18(1–2), 62–74.
McMahan, B., Moore, E., Ramage, D., Hampson, S., & Arcas, B.A. (2017). Communication-efficient learning of deep networks from decentralized data. In: Artificial Intelligence and Statistics, pp. 1273–1282 https://doi.org/10.48550/arXiv.1602.05629 . PMLR
Nishtar, Z., Wang, F., Jaskani, F. H., & Afzaal, H. (2025). Real-time fault detection and isolation in power systems for improved digital grid stability using an intelligent neuro-fuzzy logic. Computer Modeling in Engineering & Sciences,143(3), 2919–2956. https://doi.org/10.32604/cmes.2025.065098
Nishter, Z., & Wang, F. (2024). Implementation of fuzzy logic scheme for assessment of power transformer oil deterioration using imprecise information. Energies,17(21), 5412. https://doi.org/10.3390/en17215412
Oliveira, H. S., & Oliveira, H. P. (2023). Transformers for energy forecast. Sensors,23(15), 6840. https://doi.org/10.3390/s23156840
Qin, L., Zhu, T., Zhou, W., & Yu, P. S. (2025). Knowledge distillation in federated learning: A survey on long lasting challenges and new solutions. International Journal of Intelligent Systems,2025(1), 7406934. https://doi.org/10.1155/int/7406934
Rafi, S. H., Deeba, S. R., & Hossain, E. (2021). A short-term load forecasting method using integrated cnn and lstm network. IEEE access,9, 32436–32448. https://doi.org/10.1109/ACCESS.2021.3060654
Rao, S., Li, Y., Ramakrishnan, R., Hassaine, A., Canoy, D., Cleland, J., Lukasiewicz, T., Salimi-Khorshidi, G., & Rahimi, K. (2022). An explainable transformer-based deep learning model for the prediction of incident heart failure. IEEE Journal of Biomedical and Health Informatics,26(7), 3362–3372. https://doi.org/10.1109/JBHI.2022.3148820
Thein, T. T., Shiraishi, Y., & Morii, M. (2024). Personalized federated learning-based intrusion detection system: Poisoning attack and defense. Future Generation Computer Systems,153, 182–192. https://doi.org/10.1016/j.future.2023.10.005
Ullah, F., Asmat, H., Khan, A. A., Mohmand, M. I., Ali, F., Alsisi, R. H., Aldhyani, T. H., & Kwak, D. (2025). Lightweight multimedia anomaly and integrity detection for consumer iot using knowledge distillation. IEEE Transactions on Consumer Electronics. https://doi.org/10.1109/tce.2025.3644297
Ullah, F., Pun, C.-M., Mohmand, M. I., Mahendran, R. K., Khan, A. A., Alhammad, S. M., Rodrigues, J. J., & Farouk, A. (2025). Privacy-aware secure data auditing for cloud-based intelligence of things environment. IEEE Internet of Things Journal,12(11), 15288–15303.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems. https://doi.org/10.48550/arXiv.1706.03762
Wang, F., & Nishter, Z. (2024a). Real-time load forecasting and adaptive control in smart grids using a hybrid neuro-fuzzy approach. Energies,17(11), 2539. https://doi.org/10.3390/en17112539
Wang, F., & Nishter, Z. (2024b). Innovative load forecasting models and intelligent control strategy for enhancing distributed load levelling techniques in resilient smart grids. Electronics,13(17), 3552. https://doi.org/10.3390/electronics13173552
Wang, R. (2025). Buildings energy data analytics with multi-task and federated learning. University of British Columbia. https://doi.org/10.14288/1.0448054
Wang, R., Bai, L., Rayhana, R., & Liu, Z. (2024). Personalized federated learning for buildings energy consumption forecasting. Energy and Buildings,323, Article 114762. https://doi.org/10.1016/j.enbuild.2024.114762
Wen, Q., Zhou, T., Zhang, C., Chen, W., Ma, Z., Yan, J., & Sun, L. (2022). Transformers in time series: A survey. arXiv preprint https://doi.org/10.48550/arXiv.2202.07125arXiv:2202.07125
Wu, C., Wu, F., Lyu, L., Huang, Y., & Xie, X. (2022). Communication-efficient federated learning via knowledge distillation. Nature Communications,13(1), 1–8.
Wu, Z., Zhang, H., Wang, P., & Sun, Z. (2022). Rtids: A robust transformer-based approach for intrusion detection system. IEEE Access,10, 64375–64387. https://doi.org/10.1109/ACCESS.2022.3182333
Xiao, J.-W., Cao, M., Fang, H., Wang, J., & Wang, Y.-W. (2023). Joint load prediction of multiple buildings using multi-task learning with selected-shared-private mechanism. Energy and Buildings,293, Article 113178. https://doi.org/10.1016/j.enbuild.2023.113178
Zhao, Y., Li, M., Lai, L., Suda, N., Civin, D., & Chandra, V. (2018). Federated learning with non-iid data. arXiv preprint https://doi.org/10.48550/arXiv.1806.00582arXiv:1806.00582
Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., & Zhang, W. (2021). Informer: Beyond efficient transformer for long sequence time-series forecasting. Proceedings of the AAAI Conference on Artificial Intelligence,35, 11106–11115. https://doi.org/10.1609/aaai.v35i12.17325
Acknowledgements
This work has been achieved in the frame of the EIPHI Graduate School (contract "ANR-17-EURE-0002").
Author information
Authors and Affiliations
Contributions
Conceptualization, J.A.A. and H.H.; methodology, H.H.; software, J.A.A.; validation, A.M.; formal analysis, A.M.; investigation, J.A.A.; resources, H.H.; data curation, J.A.A.; writing-original draft preparation, J.A.A. and H.H.; writing-review and editing, A.M.; visualization, J.A.A.; supervision, H.H.; project administration, A.M. All authors reviewed and approved the final manuscript.
Corresponding author
Ethics declarations
Conflict of interest
The authors declare that they have no competing interests.
Ethical Approval
This study does not involve human participants or animals. All data used in this research are either publicly available or synthetically generated. Therefore, ethical approval and informed consent were not required.
Informed Consent
This study does not involve human participants or animals. All data used in this research are either publicly available or synthetically generated. Therefore, ethical approval and informed consent were not required.
Additional information
Editors: Bruno Casella, Linara Adilova, Michael Kamp.
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Al Achy, J., Harb, H. & Makhoul, A. SWIFT-KD: Sliding Window Intelligent Federated Transformer Learning with Knowledge Distillation for Building Energy Prediction. Mach Learn 115, 216 (2026). https://doi.org/10.1007/s10994-026-07153-4
Received:
Revised:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s10994-026-07153-4
Facts Only
* J. Al Achy, H. Harb, and A. Makhoul developed the SWIFT-KD framework.
* The framework is designed for building energy forecasting.
* SWIFT-KD utilizes hierarchical sliding window transformers and federated knowledge distillation.
* The system was evaluated using the ASHRAE Great Energy Predictor III dataset.
* The evaluation involved 100 heterogeneous buildings.
* Measured results include a Coefficient of Determination (R²) of 0.9708, a Root Mean Squared Error (RMSE) of 92.24 kWh, and a Mean Absolute Error (MAE) of 44.58 kWh.
* SWIFT-KD outperformed standard federated averaging by 21%.
* SWIFT-KD exceeded centralized training performance by 13.3%.
* The framework achieved peak performance within two communication rounds.
* Communication overhead was reduced 300-fold through the transmission of soft predictions instead of full model weights.
* Source code and data partitions are hosted at https://github.com/Git-Jess-Hub/swift-kd
* The work was funded by the EIPHI Graduate School under contract ANR-17-EURE-0002.
Executive Summary
Building energy forecasting is often hindered by privacy regulations, data heterogeneity across different buildings, and the limited communication capacity of edge devices. To address these constraints, the SWIFT-KD framework employs a hierarchical architecture that breaks long energy sequences into overlapping segments. This allows the model to capture both immediate local patterns and long-term dependencies without overloading the computational resources of the devices where the data resides.
The framework integrates federated knowledge distillation, which significantly lowers communication costs by sharing soft predictions rather than entire model weights. Testing on a dataset of 100 heterogeneous buildings indicates that this distributed approach not only maintains high accuracy—surpassing both standard federated averaging and centralized training—but also converges rapidly. These results suggest that high-precision energy management is possible while adhering to strict privacy and resource constraints.
Full Take
The study presents a sophisticated integration of Transformer architectures and Federated Learning (FL), specifically targeting the "communication-accuracy" trade-off. From a methodology standpoint, the use of a 300-fold communication reduction via knowledge distillation is a significant claim. However, a peer reviewer would likely question the stability of "soft predictions" across highly heterogeneous building types; if building energy profiles vary wildly, the "teacher" model's soft targets may not always translate effectively to "student" models in vastly different environments.
The claim that SWIFT-KD exceeds centralized training by 13.3% is provocative. Typically, centralized models act as the upper bound because they have access to the entire global dataset. This suggests that the federated approach may be acting as a form of regularization, preventing the overfitting that often plagues centralized energy models. The framing of novelty is justified through the combination of hierarchical windowing and distillation, extending the work of Vaswani et al. and McMahan et al. into the specific domain of time-series energy data.
For these findings to matter outside the lab, the "soft predictions" must remain robust against adversarial noise or data drift over time. The rapid convergence in two rounds is an impressive efficiency gain, but long-term stability across different seasons (which the ASHRAE dataset covers, but the study's specific temporal partitions may not fully exhaust) remains a critical variable.
Bridge Questions:
1. Does the performance advantage over centralized training persist when the number of heterogeneous buildings increases from 100 to 1,000?
2. How does the framework handle "outlier" buildings whose energy patterns deviate fundamentally from the global soft-prediction trend?
Counterstrike Scan: This is a standard scholarly submission following academic norms; it does not match the structural patterns of an influence campaign.
Sentinel — Human
This text appears to be a scientifically rigorous summary of novel machine learning methodology applied to energy forecasting, exhibiting the structure and detail characteristic of published academic research.
